Key Info
llama.cpp can distribute inference across heterogeneous devices via its ggml RPC backend, demonstrated running MiMo 2.6 Flash at 40 tokens/sec across an RTX 6000 GPU and an M5 laptop over 10 GbE.
Highlights
- Uses the ggml RPC backend for cross-device inference distribution
- Native mxfp4 weights of a state-of-the-art model run on heterogeneous hardware
- Supported out of the box in llama.cpp
- Currently an advanced setting, with plans to make it more accessible to regular users