The500Feed.Live
Everything going on in AI - updated daily from 500+ sources
📄 ResearchAugust 12, 2026
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option labels while committing to a...
Read Original Article →