The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchAugust 12, 2026

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer i...

Read Original Article →

Source

http://arxiv.org/abs/2608.11694v1