Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova
AWS's Self-Distilled Reasoning technique prevents catastrophic forgetting during SFT without a teacher model
AWS researchers introduce Self-Distilled Reasoning (SDR), a technique that reuses a model's own chain-of-thought traces during supervised fine-tuning to prevent catastrophic forgetting. Vanilla SFT caused math performance to drop from 70% to 6% on Amazon Nova 2 Lite; SDR recovered nearly all of it while matching or improving target task performance. The method requires no human annotation or separate teacher model, making it practically appealing for domain-specific fine-tuning.