The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchSeptember 2, 2026

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, sta...

Read Original Article →

Source

http://arxiv.org/abs/2609.02459v1