Key Info

Qwen unveiled RecreationBench, a benchmark for hybrid computer-use agents with 250 application-recreation tasks spanning Ubuntu, macOS, Windows, Android, and Web, covering domains like productivity, development, graphics, multimedia, and science. Agents must explore a running reference app, recreate it in code, and pass both programmatic tests and VLM-based visual evaluation.

Highlights

  • RecreationBench goes beyond GUI-only or terminal-only benchmarks by forcing agents to switch between using applications and writing code.
  • The 250 tasks span five major platforms and multiple domains, making it a broad testbed for real-world agent capability.
  • Evaluation combines automated programmatic tests with VLM-based visual checks for deeper verification.
  • The playground and dataset are already available for researchers and developers to build on.