Key Info
Mistral Large 4 reportedly beats Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because it refuses far fewer security-related tasks while Opus and Astra have around 40% of tasks blocked by their own safety filters.
Highlights
- Mistral Large 4 leads Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks
- The edge is attributed to refusing far fewer security tasks
- Opus and Astra reportedly block about 40% of tasks with safety filters
- The model is now available in Cline for security-related work
- Available via CLI (npm i -g cline) and desktop