Key Info
The team is releasing ExplorationBench, a new benchmark focused on measuring how AI systems explore beyond known problems.
Highlights
- Scientific discovery requires framing hypotheses, designing experiments, and learning from results
- Evaluating genuine novelty is difficult because truly new answers lack straightforward checks
- The benchmark targets the exploration capability of AI systems rather than standard task performance