Skip to content

On-Policy Distillation (OPD) Overview ​

Translation pending

This page has not been translated yet. Please read the Chinese version.

Why On-Policy ​

Method: Per-Token Reverse KL as a Dense Reward ​

Where OPD Sits Among SFT, Distillation, and RL ​

Timeline ​

Industrial Recipes ​

When OPD Works and When It Fails ​

Implementation Notes ​

References ​