Key Info

Tencent's research revisits critical-batch-size theory for online LLM RL, demonstrating that scaling batch size improves efficiency on fixed hardware.

Highlights

  • Learning-rate retuning can preserve learning per response over a bounded range of batch sizes, across GRPO and PPO.
  • Scaling up batch size improves PPO generation-stage throughput by up to 2.29x.
  • Best measured GRPO configuration reaches the same validation target in 29% less time.