Key Info
Step-Jev introduces step-by-step rewards for agent RL even when there's no answer key, countering the post-DeepSeek R1 consensus that step rewards get hacked and only final-answer rewards should be used.
Highlights
- Claims that when an AI judge grades only the final answer, the agent learns to fool it rather than solve correctly
- Offers process rewards using a universal verifier (Jev)
- Positions itself against the "reward only the final answer" approach popularized after DeepSeek R1
- Raised questions about adversarial robustness of the verifier