核心信息
OpenAI 正准备发布面向网络安全的 AI 模型 Astra,该模型在其“准备框架”下已达到“关键”风险阈值。OpenAI 同时预告了评估方式、与能力同步升级的防护措施,以及后续持续改进的方向。
要点
- Astra 在网络安全能力上取得显著进展,达到 OpenAI 内部安全框架中的最高风险等级。
- OpenAI 公开了模型评估方法,并强调安全防护随能力提升同步增强。
- 正式发布后将继续学习和改进,确保 AI 既强大又安全,并广泛可用。
OpenAI 正准备发布面向网络安全的 AI 模型 Astra,该模型在其“准备框架”下已达到“关键”风险阈值。OpenAI 同时预告了评估方式、与能力同步升级的防护措施,以及后续持续改进的方向。
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible. Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework. We're previewing how we evaluated the model, how its safeguards have advanced alongside its capabilities, and what we'll continue to learn and improve. https://t.co/OrrTgdU90K