Key Info
OpenAI has released updated versions of the BrokenArXiv and ArXivMath benchmarks, now centered on conjectures refuted on ArXiv within the past month and evaluated inside a harness rather than through direct API calls. GPT-6 Astra leads the performance results.
Highlights
- The refreshed benchmarks focus on recently refuted ArXiv conjectures, making them more current and more challenging.
- Models are executed inside a harness for a more controlled, reliable evaluation setup instead of direct API access.
- GPT-6 Astra sits at the top of the results, with the release promoted as fast, frontier-level, efficient, and for everyone.