Key Info
MiMo 2.6 Flash's native mxfp4 weights can run at 40 tokens/sec across heterogeneous hardware (an RTX 6000 GPU and an M5 laptop over 10 GbE), with llama.cpp supporting this out of the box.
Highlights
- Native mxfp4 weights of a state-of-the-art model
- Runs across heterogeneous hardware (RTX 6000 + M5 laptop) linked over 10 GbE
- Achieves 40 tokens/sec throughput
- Supported out of the box in llama.cpp