Key Info

Perplexity post-trained a Computer model using hint-guided self-distillation, allowing it to learn from its own errors. In live A/B tests, a later checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint.

Highlights

  • With hints, the unchanged model avoided the original failure in 93.7% of cases, up from 75.1%.
  • In a separate live test, tool-call failures fell from 2.24% to 1.77% between trained versions, without hints at inference time.
  • The full article is available via the provided link.