Key Info

A benchmark compared OpenAI's Decisions API with Jev on a production task serving 2.5 million users, where the model predicts relevance of a user query/resume to a job description on a 1-10 scale.

Highlights

  • OpenAI Decisions API measured 2x more expensive than Jev
  • OpenAI performed 5-10% worse on the relevance-scoring task
  • Task: score query/resume relevance to a job description on a 1-10 scale
  • Benchmark run on a system serving 2.5 million users
  • TypeSafe AI frames this as rebutting claims that Jev has been superseded