Key Info
On the same coding task using the same model, LiteLLM's internal coding agent saw costs differ 13x ($126 vs $9.59) purely based on the harness used, with prompt caching hit rates (0% vs 96%) accounting for nearly all of the gap.
Highlights
- Same model, same task: $126 vs $9.59 depending on harness
- Prompt caching hit rate was 0% in one case vs 96% in the other
- Nearly all of the price difference attributed to caching behavior
- Practical takeaway: harness design strongly affects LLM cost efficiency