On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level…

2 points•simonpure•2 days ago•0 comments•

0 comments

No comments yet.

Related stories