Adversarial Dual On-Policy Distillation from Expressive Flow-based Teacher

Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal expert actions. Yet these methods remain offline supervised learners: the policy is trained only on expert states and rec...

Read Original Article →

Source

http://arxiv.org/abs/2605.27095v1