AI Daily · Oct 5, 2026: Same model hits 62% on Mini-SWE-Agent but only 33% on Claude Code
Same model weights score 62% under Mini-SWE-Agent versus 33% under Claude Code, showing harness choice matters hugely; Requesty reveals Claude Opus 5.5 cache hits cost 20x less than misses, Nvidia open-sources Lyra 2.0 image-to-3D-world tool, and Security-One open-weight 27B security decision model
Pick a date
-
Same model scores 62% on Mini-SWE-Agent versus 33% on Claude Code
Researchers converted coding harnesses like Claude Code, Codex, Hermes, Pi, and opencode into RL environments with no changes to harness or training code; identical model weights scored 62% under Mini-SWE-Agent but only 33% under Claude Code.
-
Claude Opus 5.5 cache hits priced 20x below misses
Requesty reports Claude Opus 5.5 caches input tokens at $0.20 per million on a hit versus $4.00 per million on a miss, a 20x gap, which matters because agent workloads resend the full conversation every turn.
-
Security-One open-weight 27B model released for security decisions
The team released Security-One, an open-weight 27B model that ingests a document, tool call, or code change and outputs probabilities over user-defined answers, letting downstream code decide whether to allow, block, or escalate for review.
-
Nvidia open-sources Lyra 2.0 image-to-3D-world tool
Nvidia released Lyra 2.0, a fully open-source tool that converts any image into a walkable, explorable 3D world, with the model hosted on Hugging Face and UI code published on GitHub; it supports placing a robot into the scene for simulation.
-
Hugging Face addresses multi-harness RL generalization gap
Hugging Face published guidance on multi-harness reinforcement learning, noting that a single open-weight model can drop accuracy or emit invalid tool calls when moved between interfaces because training within one harness overfits to that harness's tool naming conventions.
-
Hugging Face adds unified RL environment hub to its platform
Hugging Face announced RL environments now live on its Hub, letting users discover environments across frameworks like OpenEnv, Verifiers, Harbor, and NeMo Gym in one place instead of scattered registries.
-
Ling 3.1 Flash now free to use via Requesty gateway
Requesty announced that AntLing's Ling 3.1 Flash model, hosted by Novita Labs, is now available for free through its AI gateway.
-
Cline pauses free DeepSeek-V4.1-Flash promo after abuse
Cline is temporarily halting its free DeepSeek-V4.1-Flash promotion due to abnormally high abuse, and says it is investigating and working on mitigating the situation.
-
NVFP4-quantized Qwen3.8-Flash-Next checkpoint released with benchmarks
Red Hat AI released an NVFP4 checkpoint of Qwen3.8-Flash-Next on Hugging Face, with MoE experts quantized to FP4 and other weights kept in BF16, ready for vLLM, and benchmarked against other popular checkpoints; it accepts text, image, and video input.
-
Command Code GOAT plan offers $60 credits across four models
Command Code's $10/month GOAT plan bundles $60 in usage credits usable with DeepSeek V4.1 Flash, DeepSeek V4.1 Flash Fast, GLM 5.3 Flash, and Kimi K3 models, described as billions of tokens.