Key Info
TypeSafe AI's Jev model outperforms frontier LLMs at lower cost in shadow-testing on live production traffic, achieving up to 4x speed and higher accuracy on tasks like repeat-question matching and expense categorization.
Highlights
- Repeat-question matching accuracy improved from 70% to 97%.
- Expense categorization accuracy went from 50% to 86% (vs human reviewers).
- Escalation decisions caught the same issues with fewer false alarms.
- Speed measured while shadowing live production traffic: up to 4× faster; offline tests: 2 to 3×.
- Use cases include selecting from known sets: matching analytics questions to approved answers, picking 1 of 36 metrics, blocking PII requests, deciding when a support chat needs a human, tagging tickets across a 3-level taxonomy, and sorting expenses into about 55 categories.