Key Info
Xiaomi identified a tool-call repetition issue in MiMo-V2.6 models that wasted context and stalled tasks, traced to a reward blind spot in RL training. They fixed it with a lightweight repetition-specialized RL teacher trained on just 12 steps and ~7k examples.
Highlights
- The issue caused models to repeat identical or highly similar tool calls, hurting performance in MiMo Desktop, MiMo Code, and OpenCode.
- Root cause: flooding penalty only activated beyond 32 tool calls per turn, leaving inefficient behaviors unpunished.
- Fix: a lightweight RL teacher trained on ~7k examples in 12 steps without extensive retraining.