The500Feed.Live
Everything going on in AI - updated daily from 500+ sources
📄 ResearchJuly 27, 2026
What do Reward Models Memorize?
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, u...
Read Original Article →