Key Info

Surge AI ran a blind human evaluation of five frontier models on coding quality for the Mistral Large 4 release, with expert software engineers rating them; Le Chonk ranked #1 among open-weight models and #2 overall, behind only Opus 5.

Highlights

  • Surge AI's blind eval had expert software engineers judge five frontier models on coding quality
  • Mistral Large 4 ('Le Chonk') placed first among open-weight models
  • It placed second overall, behind only Opus 5
  • The evaluation framed the test as measuring 'would you merge it?' code review judgment, not just whether code passes unit tests