Key Info

llama.cpp can distribute inference across heterogeneous devices via its ggml RPC backend, demonstrated running MiMo 2.6 Flash at 40 tokens/sec across an RTX 6000 GPU and an M5 laptop over 10 GbE.

Highlights

  • Uses the ggml RPC backend for cross-device inference distribution
  • Native mxfp4 weights of a state-of-the-art model run on heterogeneous hardware
  • Supported out of the box in llama.cpp
  • Currently an advanced setting, with plans to make it more accessible to regular users