Key Info

On the same coding task using the same model, LiteLLM's internal coding agent saw costs differ 13x ($126 vs $9.59) purely based on the harness used, with prompt caching hit rates (0% vs 96%) accounting for nearly all of the gap.

Highlights

  • Same model, same task: $126 vs $9.59 depending on harness
  • Prompt caching hit rate was 0% in one case vs 96% in the other
  • Nearly all of the price difference attributed to caching behavior
  • Practical takeaway: harness design strongly affects LLM cost efficiency