Key Info
MachGen has open-sourced its variant of WC Attention for NVIDIA Blackwell, claiming roughly 2x the speed of BF16 FlashAttention on B200 for models like MiniMax H3. The team also made improvements beyond the original paper, pushing performance more than 20% further while staying within 30% of the dense attention kernels used in production.
Highlights
- Claimed ~2x speed over BF16 FlashAttention on B200 for models like MiniMax H3
- Improvements beyond the original paper push performance more than 20% further
- Performance now within 30% of production dense attention kernels
- Requires significantly less [resource detail truncated in source]
- Released as open source