Key Info

MiniMax M3 is now supported on NVIDIA Vera Rubin NVL72 through vLLM, with collaboration from Inferact, NVIDIA AI, and Red Hat AI.

Highlights

  • Early vLLM results show more than 7.8x the throughput of GB200 on MiniMax M3 for AgentX workloads
  • vLLM has been updated to support NVIDIA Vera Rubin since its announcement
  • Multiple partners contributed to bringing the model up on the new hardware