Key Info

TRL v1.15 has been released as the project's biggest optimization yet, with a fused LM head enabled by default.

Highlights

  • Up to 82% less peak VRAM usage
  • Sequences up to 7x longer
  • DPO capacity grows from 10k to 59k tokens
  • GRPO capacity grows from 29k to 115k tokens