The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchAugust 10, 2026

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBench, a self-evolving benchmark that evaluates such behavior risk from execution trajectories rather than final responses...

Read Original Article →

Source

http://arxiv.org/abs/2608.09476v1