The500Feed.Live
Everything going on in AI - updated daily from 500+ sources
📄 ResearchJuly 20, 2026
OR Else: A Differentiable Trust Region for Policy Optimization
PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduces an abrupt change in the scalar objective's derivative. We ask whether Output Reset (OR), a smooth one-sided saturation rule, offers a useful alternative for large language model po...
Read Original Article →