| |
Self-Distillation Enables Continual Learning [PDF]
Researchers introduce Self-Distillation Fine-Tuning (SDFT), a method that enables models to learn new skills from expert demonstrations while preserving existing capabilities without requiring explicit reward functions. Unlike traditional supervised fine-tuning, SDFT uses a demonstration-conditioned model as its own teacher to generate on-policy training signals, achieving higher accuracy on new tasks while significantly reducing catastrophic forgetting. The approach allows a single model to accumulate multiple skills sequentially without performance degradation.
Read Full Article →
← More Tech news