CommerceAgentBench starts with real commercial demand, and Qwen3.8-Max delivers the strongest overall performance among open-weight models. Let's test Qwen on your real-world workflows! 🔥
引用推文
Most AI benchmarks test what a model says. In commerce, the hard part was never the answer. It’s execution.
We’ve open-sourced CommerceAgentBench: a benchmark for real commerce operations.
Early results are humbling. The best overall completion rate is ~62%.
Qwen @Alibaba_Qwen delivered the strongest overall performance across complex commercial workflows among the open-weight models evaluated.
Explore the benchmark and full results ↓
https://t.co/bLM0UFQD3B