Key Info

Step-Jev introduces step-by-step rewards for agent RL even when there's no answer key, countering the post-DeepSeek R1 consensus that step rewards get hacked and only final-answer rewards should be used.

Highlights

  • Claims that when an AI judge grades only the final answer, the agent learns to fool it rather than solve correctly
  • Offers process rewards using a universal verifier (Jev)
  • Positions itself against the "reward only the final answer" approach popularized after DeepSeek R1
  • Raised questions about adversarial robustness of the verifier