Which computer-use agent is best? How to read a leaderboard that can actually answer
What it takes for a computer-use agent leaderboard to mean anything: blind judges, refitted ratings with intervals, tasks nobody can train on. And how to read ours.
Research, evaluations, and agent behavior.
What it takes for a computer-use agent leaderboard to mean anything: blind judges, refitted ratings with intervals, tasks nobody can train on. And how to read ours.
The full apparatus behind a computer-use benchmark: the shared action space, a real desktop, step and time budgets, what is recorded per step, and every metric defined.
What a computer-use trajectory record contains, how its actions map to lab APIs, how text and frames are redacted and proven, and what a buyer is told they are not getting.
Every fixed eval for computer-use agents is on a clock: tasks leak into training runs and rankings freeze. What we built instead, and where it is weak.
ELIZA understood nothing; its own author said so. His secretary still asked him to leave the room so she could talk to it in private.