The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchAugust 13, 2026

CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens accordi...

Read Original Article →

Source

http://arxiv.org/abs/2608.13387v1