Key Info
Tencent's new research extends classical critical-batch-size theory to online LLM reinforcement learning, demonstrating that learning-rate retuning can preserve learning per response over a bounded range of batch sizes, improving training efficiency on fixed hardware.
Highlights
- Scaling up batch size improves PPO generation-stage throughput by up to 2.29× on fixed hardware.
- Best measured GRPO configuration reaches the same validation target in 29% less time.
- Findings apply across both GRPO and PPO, covering rollout generation and training scaling differences.