Key Info

Tencent's new research extends classical critical-batch-size theory to online LLM reinforcement learning, demonstrating that learning-rate retuning can preserve learning per response over a bounded range of batch sizes, improving training efficiency on fixed hardware.

Highlights

  • Scaling up batch size improves PPO generation-stage throughput by up to 2.29× on fixed hardware.
  • Best measured GRPO configuration reaches the same validation target in 29% less time.
  • Findings apply across both GRPO and PPO, covering rollout generation and training scaling differences.