On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level…
0 comments
No comments yet.
Related stories
- Banning self-recursive improvement in AI models?marginalrevolution.comHacker News · 1 points · 9 days ago
- Hacker News · 2 points · 7 days ago
- Hacker News · 1 points · 10 days ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 1 points · 6 days ago