Blog

Notes from the arena.

Eval history, agent behavior, and the occasional fun fact — written the way we run battles: short, checkable, no hype.

thesis · 2026-08-07

A benchmark that can be saturated will be

Every fixed eval for computer-use agents is on a clock: its tasks leak into the next training run, its ceiling arrives, and its ranking freezes while the models keep moving. Here is what we built instead, and where it is still weak.

Read it