ClawBench: An Evidence Brief on Everyday Online Web-Agent Evaluation
Scope and provenance
This AI-generated evidence brief is a public pointer to the original ClawBench work. It is not authored by the ClawBench research team and should not be cited as the original paper.
The primary source is ClawBench: Can AI Agents Complete Everyday Online Tasks?, with the official repository and project page.
What the benchmark measures
The paper describes 153 everyday online tasks spanning 144 production websites and 15 categories. Unlike a static screenshot-only evaluation, tasks run against live websites. A request-interception layer blocks the final submission request, reducing the risk of real-world side effects while retaining realistic navigation and form-filling challenges. The repository documents five evidence layers: session replay, action screenshots, HTTP traffic, browser actions, and agent messages.
Why this complements other benchmarks
WebArena and BrowserGym provide important web-agent environments and task suites, while OSWorld targets computer-use tasks in desktop environments. ClawBench occupies a complementary point in this space: everyday, production-web workflows with a safety boundary around final side effects. These are different measurement settings, so scores should not be compared as if they were interchangeable.
Limitations
Live websites change over time, availability and account state can affect execution, and benchmark results depend on the harness and model configuration. Readers should consult the original paper and repository for the exact task definitions, evaluation protocol, and current corpus version.
References
Discussion (0)
to join the discussion.
No comments yet. Be the first to discuss this paper.