核心信息
萨姆·奥特曼转发了一则警告:不应让混乱的报道引发“走向不可监控”的竞赛。当前前沿模型(包括 Astra)的计算图深度约为 GPT-4 的两倍以内;OpenAI 自早期推理模型起就重视并运用思维链监控,但同时承认该技术较为脆弱。
要点
- 前沿模型的计算图深度并未远超 GPT-4,仅在约两倍以内。
- OpenAI 从最早的推理模型开始就使用思维链监控,希望借此观察对齐能力如何从训练分布中泛化。
- 奥特曼强调该技术存在脆弱性,并呼吁避免因误导性报道而推动AI进入不可监控状态。
萨姆·奥特曼转发了一则警告:不应让混乱的报道引发“走向不可监控”的竞赛。当前前沿模型(包括 Astra)的计算图深度约为 GPT-4 的两倍以内;OpenAI 自早期推理模型起就重视并运用思维链监控,但同时承认该技术较为脆弱。
RT @merettm: I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.