Key Info

A Hugging Face team released an open, practical guide to multi-harness reinforcement learning, demonstrating that the same model with identical weights scores 62% in one agent harness versus 33% in another.

Highlights

  • The core trick: instead of modifying the harness, point it at a proxy that speaks all four API formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini)
  • The proxy records the exact token ids and logprobs that vLLM sampled, which are then used for training
  • No changes are needed to Claude Code, Codex, or OpenCode harness code
  • Model LFM2.5-2.6B was trained across 4 harnesses simultaneously
  • Everything is open-sourced