Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish useful visual evidence from incidental image context in lexical judgments. We use human concreteness and imagery ratings ...

Read Original Article →

Source

http://arxiv.org/abs/2605.27315v1