Key Info

Mistral Large 4 reportedly beats Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because it refuses far fewer security-related tasks while Opus and Astra have around 40% of tasks blocked by their own safety filters.

Highlights

  • Mistral Large 4 leads Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks
  • The edge is attributed to refusing far fewer security tasks
  • Opus and Astra reportedly block about 40% of tasks with safety filters
  • The model is now available in Cline for security-related work
  • Available via CLI (npm i -g cline) and desktop