Key Info

The team is releasing ExplorationBench, a new benchmark focused on measuring how AI systems explore beyond known problems.

Highlights

  • Scientific discovery requires framing hypotheses, designing experiments, and learning from results
  • Evaluating genuine novelty is difficult because truly new answers lack straightforward checks
  • The benchmark targets the exploration capability of AI systems rather than standard task performance