RT @ivanfioravanti: Sherry quantization method by Tencent, combined with large Hy4 Preview is able to deliver incredible results!
1.5TB to 214GB using 1.25 bits weights while keeping incredible accuracy!
BF16 vs Sherry:
- MCP Atlas 83.7→83.2
- SWE-Bench multi 82.9→81.3
- MRCR 81.3→81.1
- IFBench 73.5→72.5
引用推文
1.5TB → 214GB. Seven times smaller, barely a dent.
That's Hy4 preview. The trick is Sherry — our quantization method that packs weights down to 1.25 bits each. (The image shows how.)
It also unlocks a new way to run it: stitch the GPUs you already have
across machines, and they work as one.
GGUF (quantized): https://t.co/EC2ZYpVnC4
Original model: https://t.co/olB85l7dxy