Key Info
TRL v1.15 has been released as the project's biggest optimization yet, with a fused LM head enabled by default.
Highlights
- Up to 82% less peak VRAM usage
- Sequences up to 7x longer
- DPO capacity grows from 10k to 59k tokens
- GRPO capacity grows from 29k to 115k tokens