Key Info

MiMo 2.6 Flash's native mxfp4 weights can run at 40 tokens/sec across heterogeneous hardware (an RTX 6000 GPU and an M5 laptop over 10 GbE), with llama.cpp supporting this out of the box.

Highlights

  • Native mxfp4 weights of a state-of-the-art model
  • Runs across heterogeneous hardware (RTX 6000 + M5 laptop) linked over 10 GbE
  • Achieves 40 tokens/sec throughput
  • Supported out of the box in llama.cpp