核心信息
Inception AI 的 Mercury 2.5 预览版现已独家上线 OpenRouter,通过并行令牌生成实现每秒 1107 tokens,号称最快的推理大模型。
要点
- 通过并行令牌生成达到每秒 1107 tokens,专为延迟敏感型负载设计。
- 支持可调推理深度与并行工具调用。
- 输出符合 schema 的 JSON,便于结构化应用。
- 目前仅在 OpenRouter 上提供。
Inception AI 的 Mercury 2.5 预览版现已独家上线 OpenRouter,通过并行令牌生成实现每秒 1107 tokens,号称最快的推理大模型。
1/ The fastest reasoning LLM is now live exclusively on OpenRouter. Mercury 2.5 Preview from @_inception_ai reaches 1,107 tokens/sec through parallel token generation, with tunable reasoning, parallel tool calls, and schema-aligned JSON. Built for latency-sensitive workloads. https://t.co/xuGXhBDwiL