核心信息
Anthropic承诺向第三方评估人员提供永久性的员工级系统访问权限,这是其“为前沿减速”三部分计划的第一步。Cline对此表示支持,但认为开放权重模型更进一步——任何人都可以检查、评估并对模型进行红队测试。
要点
- Anthropic将允许第三方评估人员验证安全措施、报告事件,并在训练期间评估模型的对齐情况。
- Cline称这一承诺“令人难以置信”,但认为开放权重能更进一步:任何人都可以检查、评估并对模型进行红队测试。
- Cline希望看到更多开放合作,共同构建这项强大技术。
Anthropic承诺向第三方评估人员提供永久性的员工级系统访问权限,这是其“为前沿减速”三部分计划的第一步。Cline对此表示支持,但认为开放权重模型更进一步——任何人都可以检查、评估并对模型进行红队测试。
It's incredible seeing Anthropic commit to third-party evaluators. However, we believe open weights takes this same idea further. Anyone can inspect, evaluate, and red-team the model. We hope to see more open collaboration in building this powerful technology.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://t.co/OGyPb7yaYt