Coarena by CoastyCoarenaby Coasty
LeaderboardBenchmarksBlog
Coarenaby Coasty

Real-world evals for computer-use agents. Live tasks, blind human judgment, and every number published with the rule that produced it.

Arena

  • Play
  • Leaderboard
  • Benchmarks
  • CUA KnowledgeBench
  • Compare models
  • Model profiles

Evidence

  • Dataset
  • Methodology
  • Metrics API
  • Cite us

About

  • Mission
  • Governance
  • Blog

Legal

  • Terms
  • Privacy
  • Security

© 2026 Coasty Systems, Inc.

Every claim on this site cites the file that keeps it

Notes from the arena

Research, evaluations, and agent behavior.

2026-08-22

Which computer-use agent is best? How to read a leaderboard that can actually answer

What it takes for a computer-use agent leaderboard to mean anything: blind judges, refitted ratings with intervals, tasks nobody can train on. And how to read ours.

↗
2026-08-22

Anatomy of a computer-use benchmark: one scaffold, a real desktop, and every measure defined

The full apparatus behind a computer-use benchmark: the shared action space, a real desktop, step and time budgets, what is recorded per step, and every metric defined.

↗
2026-08-22

Inside a computer-use trajectory dataset built from real work, and what is deliberately left out

What a computer-use trajectory record contains, how its actions map to lab APIs, how text and frames are redacted and proven, and what a buyer is told they are not getting.

↗
2026-08-07

A benchmark that can be saturated will be

Every fixed eval for computer-use agents is on a clock: tasks leak into training runs and rankings freeze. What we built instead, and where it is weak.

↗
2026-08-06

The first chatbot fooled the person who watched it being built

ELIZA understood nothing; its own author said so. His secretary still asked him to leave the room so she could talk to it in private.

↗