AI Daily · Sep 25, 2026: Meituan releases 1.6T-parameter LongCat-2.5-Preview; Gemini 3.8 Flash arrives free in Cline
Meituan's LongCat-2.5-Preview launches with 1.6T params and 1M context, granting existing users 5M free tokens; Cline adds free Gemini 3.8 Flash topping its price tier; B.AI and GMI Cloud run limited-time model discounts.
Pick a date
-
Meituan releases LongCat-2.5-Preview with 1.6T parameters and 1M context
Meituan's LongCat-2.5-Preview went live as a natively multimodal model with 1.6 trillion total parameters, about 48 billion active parameters, and a 1 million token context window, targeting long-horizon agentic tasks across terminals, browsers, GUIs, and spreadsheets; all existing users receive 5 million free tokens.
-
Gemini 3.8 Flash available free with strong benchmark scores
Cline made Gemini 3.8 Flash available for free in its editor, reporting a score of 41 on the Artificial Analysis Intelligence Index — the top of its price tier — along with 291 tokens per second throughput and a 1M-token context window.
-
B.AI launches discounted tiers for MiMo V2.6 and Flash models
B.AI opened new discount tiers: MiMo-V2.6-Flash at 90% off, MiMo-V2.6-Pro at 50% off, and DeepSeek-V4.1-Flash, GLM-5.3-Flash (320B), Qwen3.8-Flash (1M context) at 70% off, live from 15:00 SGT.
-
Vercel AI Gateway data shows Anthropic spend share falling from 69% to 40%
Vercel AI Gateway's last two months of spend data shows Anthropic's share dropping from 69% to 40% while OpenAI rose from 10% to 24% and now leads in token count; Kimi K3 and DeepSeek absorbed about half of Anthropic's losses, Opus 5.5 reached 10% of spend within two days, and OpenAI accounts for 62% of image generation.
-
OpenRouter launches Server Tools Marketplace with search and shell tools
OpenRouter introduced a Server Tools Marketplace offering grounding tools like web search APIs, Shell, Advisor, Subagent, Web fetch, Image generation, Apply patch, and Datetime, most running server-side during requests.
-
Pruna open-sources LoRA adapters accelerating Qwen-Image 6.3x
PrunaAI open-sourced Pruna-Qwen-Image-2.1, a set of few-step LoRA adapters that make Alibaba's Qwen-Image-2.1 up to 6.3× faster, generating or editing images in 5 or 8 steps instead of 40.
-
Tencent Hy Translation launches with 33 languages offline support
Tencent's Hy Translation, powered by the Hy-MT2 model, launched with support for 33 languages including 5 Chinese minority languages and dialects, offering fully offline on-device voice and photo translation already live in 12 countries and regions.
-
OpenCode makes $60 DeepSeek usage grant permanent
OpenCode announced that the $60 usage grant for DeepSeek v4.1 Flash, part of its Operation Cheepseek Phase 2, has been made permanent for users.
-
Hugging Face release opens SmolDataEnvs dataset with 5000 RL tasks
The SmolDataEnvs release provides over 5,000 verifiable reinforcement learning environment tasks for training small models in coding and data science, fully open-sourced including environments, evaluations, and training code.
-
DSPy 3.4.0 adds Jev and System one model support plus ReAnchor optimizer
DSPy 3.4.0 was released with native support for Jev and System one model types inside signatures, along with a new ReAnchor optimizer designed to calibrate outputs using confidence scores.