Daily Paper Cast

E2060 - On-Policy Delta Distillation

Published: July 21, 2026

Duration: 19:41

🤗 Upvotes: 28 | cs.LG, cs.CL

Authors:
Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

Title:
On-Policy Delta Distillation

Arxiv:
http://arxiv.org/abs/2607.15161v1

Abstract:
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output dis...