Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation

Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% […]

This article has been indexed from MarkTechPost

Read the original article: