thesis · 2026-08-07
A benchmark that can be saturated will be
Every fixed eval for computer-use agents is on a clock: its tasks leak into the next training run, its ceiling arrives, and its ranking freezes while the models keep moving. Here is what we built instead, and where it is still weak.
Read it