The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchAugust 12, 2026

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option labels while committing to a...

Read Original Article →

Source

http://arxiv.org/abs/2608.11947v1