Key Info
Tencent's research revisits critical-batch-size theory for online LLM RL, demonstrating that scaling batch size improves efficiency on fixed hardware.
Highlights
- Learning-rate retuning can preserve learning per response over a bounded range of batch sizes, across GRPO and PPO.
- Scaling up batch size improves PPO generation-stage throughput by up to 2.29x.
- Best measured GRPO configuration reaches the same validation target in 29% less time.