核心信息
OpenAI 发布了一套新的框架,用于追踪、调查并公开披露模型失对齐(misalignment)事件,并明确了披露标准与时间表,即使相关行为尚未完全解释或缓解。
要点
- 该框架规定了 OpenAI 披露失对齐案例的条件和时间表。
- 更复杂的案例可能需要更长时间调查,或与第三方协调处理。
- OpenAI 表示将优先披露能揭示新型失对齐机制、已知行为发生重大变化,或挑战现有安全假设的发现。
OpenAI 发布了一套新的框架,用于追踪、调查并公开披露模型失对齐(misalignment)事件,并明确了披露标准与时间表,即使相关行为尚未完全解释或缓解。
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. https://t.co/ismCCkeE0L