
E2060 - On-Policy Delta Distillation
Published: July 21, 2026
Duration: 19:41
🤗 Upvotes: 28 | cs.LG, cs.CL
Authors:
Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han
Title:
On-Policy Delta Distillation
Arxiv:
http://arxiv.org/abs/2607.15161v1
Abstract:
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output dis...