核心信息
阿里最新模型Qwen3.8-Max-0902在CodeArena WebDev排行榜上以1691分升至第一,较前代Qwen3.8-Max的1669分提升22分,刷新智能体编码工作流纪录。
要点
- 分别领先Claude Opus 5(Max)3分、Kimi K3(Max)17分,并较前代提升22分。
- 在多步推理、工具调用和完整应用生成方面表现突出。
- 混合定价约每百万Token 5美元,在帕累托前沿上占据最高得分位置。
- 所有WebDev类别均表现强劲,Agent Arena评分即将公布。
阿里最新模型Qwen3.8-Max-0902在CodeArena WebDev排行榜上以1691分升至第一,较前代Qwen3.8-Max的1669分提升22分,刷新智能体编码工作流纪录。
🏆 #1 on CodeArena: WebDev leaderboard. Qwen3.8-Max-0902 jumps from 1669 to 1691, setting a new record for agentic coding (WebDev) workflows, with standout strength in multistep reasoning, tool use, and full app generation. Thanks for the recognition! @arena
Big news: Qwen3.8-Max-0902 by @Alibaba_Qwen just debuted at #1 overall in the Code Arena: WebDev with 1691 pts! It scores 3 pts above Claude Opus 5 (Max), 17 pts above Kimi K3 (Max), and 22 pts above the previous Qwen3.8-Max. Priced at a blended $5/MToken, Qwen3.8-Max-0902 also claims the highest-scoring position on the Pareto frontier! Stay tuned for a closer look at its Pareto positioning, and for Agent Arena scores coming soon. Its strength carries across every Code Arena: WebDev category: - #1 in Data & Analytics and Consumer Product - #2 in Brand & Marketing, Gaming, and Simulations - #3 in Content Creation Tools and Reference-Based Design Congrats to the @Alibaba_Qwen team on this huge update!