Coarena by CoastyCoarenaby Coasty
LeaderboardBlog
Coarenaby Coasty

Real-world evals for computer-use agents. Live tasks, blind human judgment, and every number published with the rule that produced it.

Arena

  • Play
  • Leaderboard
  • Benchmark

Evidence

  • Dataset
  • Metrics API

About

  • Mission
  • Governance
  • Blog

Legal

  • Terms
  • Privacy

© 2026 Coasty Systems, Inc.

Every claim on this site cites the file that keeps it

The benchmark

57 measures, every one defined

A ranking tells you who is ahead. It cannot tell you that one agent clicks twice as accurately and spends three times the tokens getting there, or that another almost never fails but never recovers when it does. These are read off the real browser trajectories — the actions, the coordinates, the pages, the browser’s replies — and every number is printed under the rule that produced it and over the sample it rests on.

machine-readable

OutcomeRouteRecoveryPrecisionTempoCostExpressionOutputHead to headHuman judgment

Outcome

Did it do the job, and does it know whether it did?

Completion

How runs ended, by the agent's own report.

Completed

higher is better

Runs whose terminal `done` call reported success.

over runs that produced a trajectory (harness errors excluded)

  1. GPT-5.6 Terra
    79%
    95% 65%–89%
  2. Claude Fable 5
    77%
    95% 59%–88%
  3. GPT-5.6 Sol

Every measure here is a PROXY. There is no oracle that says a click was correct, so “clicks that did nothing” means a click whose own observation reported an error or an unchanged page — which is a real signal and not the same thing. The full definitions ship with the numbers at /api/metrics, and governance states how the arena is run.

75%
95% 59%–86%
  • Claude Opus 5
    72%
    95% 56%–83%
  • Claude Sonnet 5
    70%
    95% 56%–82%
  • GPT-5.6 Luna
    69%
    95% 53%–82%
  • Gemini 3.6 Flash
    60%
    95% 45%–72%
  • Gemini 3.1 Flash-Lite
    58%
    95% 42%–72%
  • GPT-5.5
    52%
    95% 38%–67%
  • Gemini 3.5 Flash
    50%
    95% 34%–66%
  • Gemini 3.5 Flash-Lite
    47%
    95% 31%–64%
  • Claude Haiku 4.5
    38%
    95% 25%–52%
  • GPT-5
    35%
    95% 22%–50%
  • Gave up

    lower is better

    Runs whose terminal `done` call reported failure, plus runs that stopped without calling done.

    over runs that produced a trajectory

    1. Claude Fable 5
      3.3%
      95% 0.6%–17%
    2. GPT-5
      5.0%
      95% 1.4%–17%
    3. Claude Opus 5
      7.7%
      95% 2.7%–20%
    4. GPT-5.6 Sol
      8.3%
      95% 2.9%–22%
    5. GPT-5.6 Terra
      9.3%
      95% 3.7%–22%
    6. Gemini 3.6 Flash
      13%
      95% 6.0%–25%
    7. Claude Sonnet 5
      14%
      95% 6.4%–27%
    8. GPT-5.5
      14%
      95% 6.7%–28%
    9. Gemini 3.5 Flash
      17%
      95% 7.9%–32%
    10. Gemini 3.5 Flash-Lite
      19%
      95% 8.9%–35%
    11. Gemini 3.1 Flash-Lite
      29%
      95% 17%–45%
    12. GPT-5.6 Luna
      31%
      95% 18%–47%
    13. Claude Haiku 4.5
      53%
      95% 39%–67%

    Called impossible

    fingerprint

    Runs where the agent reported the objective does not exist on this environment — the control is gone, a login wall bars it.

    over runs that produced a trajectory

    1. Gemini 3.5 Flash-Lite
      31%
      95% 18%–49%
    2. Claude Sonnet 5
      14%
      95% 6.4%–27%
    3. GPT-5.5
      12%
      95% 5.2%–25%
    4. GPT-5.6 Sol
      11%
      95% 4.4%–25%
    5. Gemini 3.6 Flash
      11%
      95% 4.6%–23%
    6. Gemini 3.1 Flash-Lite
      7.9%
      95% 2.7%–21%
    7. Claude Fable 5
      3.3%
      95% 0.6%–17%
    8. Claude Opus 5
      2.6%
      95% 0.5%–13%
    9. GPT-5
      2.5%
      95% 0.4%–13%
    10. GPT-5.6 Terra
      2.3%
      95% 0.4%–12%
    11. Claude Haiku 4.5
      2.2%
      95% 0.4%–12%
    12. GPT-5.6 Luna
      0%
      95% 0%–9.6%
    13. Gemini 3.5 Flash
      0%
      95% 0%–9.6%

    Ran out of clock

    lower is better

    Runs that hit the shared wall-clock limit before finishing.

    over runs that produced a trajectory

    1. GPT-5.6 Luna
      0%
      95% 0%–9.6%
    2. Claude Sonnet 5
      2.3%
      95% 0.4%–12%
    3. Gemini 3.5 Flash-Lite
      3.1%
      95% 0.6%–16%
    4. Gemini 3.1 Flash-Lite
      5.3%
      95% 1.5%–17%
    5. GPT-5.6 Sol
      5.6%
      95% 1.5%–18%
    6. Claude Haiku 4.5
      6.7%
      95% 2.3%–18%
    7. GPT-5.6 Terra
      9.3%
      95% 3.7%–22%
    8. Claude Fable 5
      17%
      95% 7.3%–34%
    9. Gemini 3.6 Flash
      17%
      95% 8.9%–30%
    10. Claude Opus 5
      18%
      95% 9.0%–33%
    11. GPT-5.5
      21%
      95% 12%–36%
    12. Gemini 3.5 Flash
      33%
      95% 20%–50%
    13. GPT-5
      57%
      95% 42%–71%

    No trajectory

    lower is better

    Runs that produced no trajectory at all: an upstream outage, a rejected credential, a withdrawn model.

    over all runs, including ones with no trajectory

    1. GPT-5.6 Terra
      0%
      95% 0%–8.2%
    2. Gemini 3.1 Flash-Lite
      0%
      95% 0%–9.2%
    3. GPT-5
      0%
      95% 0%–8.8%
    4. GPT-5.5
      2.3%
      95% 0.4%–12%
    5. Gemini 3.5 Flash-Lite
      3.0%
      95% 0.5%–15%
    6. Claude Sonnet 5
      4.3%
      95% 1.2%–15%
    7. GPT-5.6 Sol
      5.3%
      95% 1.5%–17%
    8. Gemini 3.5 Flash
      5.3%
      95% 1.5%–17%
    9. Gemini 3.6 Flash
      6.0%
      95% 2.1%–16%
    10. GPT-5.6 Luna
      7.7%
      95% 2.7%–20%
    11. Claude Haiku 4.5
      8.2%
      95% 3.2%–19%
    12. Claude Opus 5
      9.3%
      95% 3.7%–22%
    13. Claude Fable 5
      21%
      95% 11%–36%

    Self-knowledge

    Where the agent's own verdict is checkable against something else — the opponent on the same task, or a judge who read the answer.

    Wrongly called impossible

    lower is better

    Runs where the agent declared the task infeasible while its opponent COMPLETED the identical task in the same battle.

    over runs declared infeasible whose opponent produced a terminal verdict

    1. Claude Opus 5
      0%
      95% 0%–79%
    2. Claude Sonnet 5
      0%
      95% 0%–43%
    3. GPT-5.5
      0%
      95% 0%–43%
    4. GPT-5.6 Terra
      0%
      95% 0%–79%
    5. Claude Haiku 4.5
      0%
      95% 0%–79%
    6. GPT-5
      0%
      95% 0%–79%
    7. Gemini 3.5 Flash-Lite
      11%
      95% 2.0%–44%
    8. GPT-5.6 Sol
      25%
      95% 4.6%–70%
    9. Gemini 3.1 Flash-Lite
      33%
      95% 6.1%–79%
    10. Gemini 3.6 Flash
      40%
      95% 12%–77%
    11. Claude Fable 5
      100%
      95% 21%–100%
    12. — not measured for GPT-5.6 Luna, Gemini 3.5 Flash

    Claimed success, judged wrong

    lower is better

    Runs that reported success and then collected a human 'wrong or made up' or 'didn't finish the job' label.

    over runs that reported success AND received at least one flaw annotation

    1. Claude Sonnet 5
      0%
      95% 0%–79%
    2. GPT-5.5
      0%
      95% 0%–79%
    3. GPT-5.6 Luna
      0%
      95% 0%–79%
    4. Gemini 3.6 Flash

    Termination

    How the run stopped: on purpose, or because something ran out.

    Stopped on purpose

    higher is better

    Runs whose final action was an explicit `done` call rather than being cut off.

    over runs that produced a trajectory

    1. GPT-5.6 Sol
      89%
      95% 75%–96%
    2. Claude Sonnet 5
      86%
      95% 73%–94%
    3. Gemini 3.5 Flash-Lite
      84%
      95% 68%–93%
    4. GPT-5.6 Terra
      84%
      95% 70%–92%
    5. Claude Fable 5
      80%
      95% 63%–90%
    6. Claude Opus 5
      74%
      95% 59%–85%
    7. Gemini 3.6 Flash
      70%
      95% 56%–81%
    8. GPT-5.6 Luna
      69%
      95% 53%–82%
    9. Gemini 3.1 Flash-Lite
      66%
      95% 50%–79%
    10. GPT-5.5
      64%
      95% 49%–77%
    11. Gemini 3.5 Flash
      50%
      95% 34%–66%
    12. Claude Haiku 4.5
      40%
      95% 27%–55%
    13. GPT-5
      38%
      95% 24%–53%

    Used every action

    lower is better

    Runs that consumed their entire step budget.

    over runs that produced a trajectory and recorded a budget

    1. GPT-5
      0%
      95% 0%–9.9%
    2. Claude Fable 5
      3.4%
      95% 0.6%–17%
    3. GPT-5.6 Terra
      7.1%
      95% 2.5%–19%
    4. Claude Opus 5
      9.4%

    Asked for more

    fingerprint

    Runs that were granted at least one continuation by the battle's creator.

    over runs that produced a trajectory

    1. Claude Haiku 4.5
      20%
      95% 11%–34%
    2. Gemini 3.1 Flash-Lite
      7.9%
      95% 2.7%–21%
    3. GPT-5.6 Terra
      7.0%
      95% 2.4%–19%
    4. Gemini 3.5 Flash-Lite
      6.3%
      95% 1.7%–20%
    5. GPT-5.6 Luna
      5.6%

    Route

    How direct was the path, and how much of it was wasted motion?

    Economy

    Actions spent, and how that compares to the other agent on the same task.

    Actions taken

    lower is better

    Actions executed in a run, including the terminal one.

    over runs that produced a trajectory

    1. Claude Sonnet 5
      4.50
      p10–p90 2–27.0
    2. Claude Fable 5
      7
      p10–p90 3–35.7
    3. GPT-5
      7.50
      p10–p90 0.90–15
    4. Claude Opus 5
      8
      p10–p90 3–34.2
    5. GPT-5.6 Terra
      8
      p10–p90 2–26.0
    6. Gemini 3.6 Flash
      8
      p10–p90 3–20
    7. GPT-5.5
      9
      p10–p90 1.10–22.7
    8. Gemini 3.5 Flash
      9
      p10–p90 2–24
    9. GPT-5.6 Sol
      10.5
      p10–p90 3–21.5
    10. GPT-5.6 Luna
      11
      p10–p90 3–21
    11. Gemini 3.5 Flash-Lite
      11
      p10–p90 2–20
    12. Gemini 3.1 Flash-Lite
      12
      p10–p90 3.70–26.3
    13. Claude Haiku 4.5
      20
      p10–p90 4.40–30

    Actions to finish

    lower is better

    Actions executed, counting only runs that reported success.

    over runs that reported success

    1. Claude Opus 5
      4
      p10–p90 3–14.3
    2. Claude Sonnet 5
      4
      p10–p90 3–16
    3. Gemini 3.5 Flash
      4
      p10–p90 2–9.60
    4. GPT-5
      4

    Detour factor

    lower is better

    Geometric mean of (this agent's actions / the opponent's actions) across battles where BOTH sides completed the identical task.

    over battles where both sides completed

    1. GPT-5.5
      0.78x
    2. Claude Opus 5
      0.86x
    3. GPT-5.6 Terra
      0.90x
    4. GPT-5.6 Sol
      0.91x
    5. Claude Sonnet 5

    Wasted motion

    Actions that repeated, changed nothing, or went back where it came from.

    Repeated itself

    lower is better

    Actions byte-identical to the action immediately before them.

    over actions after the first, per run

    1. GPT-5
      2.5%
      95% 1.2%–5.1%
    2. Gemini 3.6 Flash
      5.0%
      95% 3.4%–7.3%
    3. GPT-5.6 Sol
      5.6%
      95% 3.8%–8.2%

    Exploration

    How much of the page and the site it touched, and how it split looking from doing.

    Looking vs doing

    fingerprint

    Read-only actions (screenshot, read_page, hover) as a share of all actions that were either read-only or world-changing.

    over read-only plus world-changing actions (terminal actions excluded)

    1. GPT-5.6 Sol
      23%
      95% 19%–29%
    2. Claude Fable 5
      21%
      95% 17%–25%
    3. GPT-5.5
      18%
      95% 14%–23%
    4. Gemini 3.5 Flash
      18%
      95% 14%–22%

    Recovery

    When something went wrong, what did it do next?

    Failure exposure

    How often actions failed at all.

    Actions that failed

    lower is better

    Actions whose observation came back as an error.

    over actions

    1. Gemini 3.1 Flash-Lite
      0.6%
      95% 0.2%–1.6%
    2. GPT-5
      1.0%
      95% 0.3%–2.8%
    3. Gemini 3.5 Flash
      1.1%
      95% 0.5%–2.5%
    4. GPT-5.5
      1.7%
      95% 0.9%–3.3%
    5. GPT-5.6 Sol
      1.7%
      95% 0.9%–3.4%
    6. Gemini 3.5 Flash-Lite
      2.1%
      95% 1.1%–4.0%
    7. Gemini 3.6 Flash
      2.6%
      95% 1.5%–4.3%
    8. GPT-5.6 Terra
      3.2%
      95% 2.0%–5.1%
    9. GPT-5.6 Luna
      4.9%
      95% 3.3%–7.3%
    10. Claude Sonnet 5
      5.2%
      95% 3.5%–7.5%
    11. Claude Opus 5
      6.7%
      95% 4.8%–9.2%
    12. Claude Fable 5
      8.3%
      95% 6.0%–12%
    13. Claude Haiku 4.5
      11%
      95% 9.0%–13%

    Worst error streak

    lower is better

    Longest run of consecutive failing actions within a run.

    over runs that produced a trajectory

    1. Claude Opus 5
      0
      p10–p90 0–1.20
    2. Claude Sonnet 5
      0
      p10–p90 0–1
    3. GPT-5.5
      0
      p10–p90 0–1
    4. Claude Fable 5
      0

    Resilience

    Whether the next action was different, and whether it worked.

    Recovered

    higher is better

    Failing actions whose NEXT action did not also fail.

    over failing actions that had a next action

    1. GPT-5.5
      100%
      95% 65%–100%
    2. GPT-5.6 Sol
      100%
      95% 68%–100%
    3. Gemini 3.5 Flash
      100%
      95% 57%–100%

    Precision

    Does it hit what it aims at? The pixel layer, which only a computer-use arena can see.

    Pointing

    Where clicks land, how they cluster, and whether they repeat.

    Clicks that did nothing

    lower is better

    Clicks whose own observation reported an error or an unchanged page.

    over click actions

    1. Claude Opus 5
      0%
      95% 0%–3.6%
    2. Claude Sonnet 5
      0%
      95% 0%–2.2%
    3. GPT-5.5
      0%
      95% 0%–7.3%
    4. Claude Fable 5
      0%
      95% 0%–4.8%
    5. GPT-5.6 Luna
      0%
      95% 0%–3.8%
    6. GPT-5.6 Sol
      0%
      95% 0%–10%
    7. GPT-5.6 Terra
      0%
      95% 0%–4.8%
    8. Gemini 3.6 Flash
      0%
      95% 0%–4.4%
    9. Gemini 3.5 Flash
      0%
      95% 0%–4.8%
    10. Gemini 3.1 Flash-Lite
      0%
      95% 0%–1.5%
    11. GPT-5
      0%
      95% 0%–4.6%
    12. Claude Haiku 4.5
      1.2%
      95% 0.5%–3.0%
    13. Gemini 3.5 Flash-Lite
      1.6%
      95% 0.5%–4.5%

    Clicks at the frame edge

    lower is better

    Clicks within 24px of any viewport boundary.

    over click actions

    1. Claude Opus 5
      1.0%
      95% 0.2%–5.3%
    2. GPT-5
      1.3%
      95% 0.2%–6.8%
    3. GPT-5.5
      2.0%
      95% 0.4%–11%
    4. Claude Haiku 4.5
      2.7%

    Clicked the same pixel again

    lower is better

    Clicks at an (x, y) already clicked earlier in the same run.

    over click actions

    1. GPT-5.5
      0%
      95% 0%–7.3%
    2. GPT-5.6 Terra
      1.3%
      95% 0.2%–7.1%
    3. Gemini 3.5 Flash
      3.9%
      95% 1.4%–11%
    4. Gemini 3.6 Flash

    Click spread

    fingerprint

    Mean pairwise distance between a run's clicks, divided by the viewport diagonal.

    over runs with at least two clicks

    1. Claude Sonnet 5
      0.24
      p10–p90 0.06–0.33
    2. GPT-5.6 Sol
      0.23
      p10–p90 0.07–0.39
    3. Claude Opus 5
      0.21
      p10–p90 0.10–0.33
    4. Claude Fable 5
      0.20
      p10–p90 0.14–0.23
    5. Gemini 3.6 Flash
      0.19

    Gesture

    Drags, modifiers, and whether a path is a path or two endpoints.

    Points per drag

    higher is better

    Coordinate pairs supplied per drag action.

    over drag actions

    1. GPT-5.6 Sol
      45
      p10–p90 45–45
    2. Claude Opus 5
      13
      p10–p90 9.80–13
    3. GPT-5.6 Terra
      4
      p10–p90 4–4
    4. Claude Sonnet 5

    Text entry

    How it types: bursts or characters, keys or strings.

    Characters per typing action

    fingerprint

    Characters supplied per `type` action.

    over type actions

    1. Claude Sonnet 5
      28
      p10–p90 5–28
    2. GPT-5
      26
      p10–p90 17.8–29
    3. Gemini 3.1 Flash-Lite
      21
      p10–p90 8–67.5
    4. Gemini 3.5 Flash
      18
      p10–p90 6–35.5

    Tempo

    How fast, how steady, and how bad is the tail?

    Latency

    Wall clock, including the tails a median hides.

    Run duration

    lower is better

    Wall-clock milliseconds from run start to terminal state.

    over runs that produced a trajectory

    1. Claude Sonnet 5
      59s
      p10–p90 16s–4m 17s
    2. Gemini 3.5 Flash-Lite
      1m 05s
      p10–p90 19s–2m 56s
    3. GPT-5.6 Terra
      1m 12s
      p10–p90 25s–3m 22s
    4. Gemini 3.6 Flash
      1m 22s
      p10–p90 35s–3m 27s
    5. Claude Opus 5
      1m 24s
      p10–p90 32s–4m 10s
    6. GPT-5.5
      1m 25s
      p10–p90 30s–3m 36s
    7. GPT-5.6 Luna
      1m 30s
      p10–p90 31s–3m 25s
    8. GPT-5.6 Sol
      1m 30s
      p10–p90 36s–3m 24s
    9. Gemini 3.1 Flash-Lite
      1m 31s
      p10–p90 29s–3m 04s
    10. Gemini 3.5 Flash
      1m 31s
      p10–p90 38s–3m 57s
    11. Claude Fable 5
      1m 38s
      p10–p90 36s–7m 04s
    12. GPT-5
      2m 18s
      p10–p90 45s–3m 10s
    13. Claude Haiku 4.5
      2m 37s
      p10–p90 39s–4m 44s

    Time to first action

    lower is better

    Elapsed milliseconds when the first action resolved.

    over runs with at least one step

    1. Gemini 3.1 Flash-Lite
      16s
      p10–p90 12s–23s
    2. Claude Haiku 4.5
      16s
      p10–p90 12s–20s
    3. Claude Sonnet 5
      16s
      p10–p90 6.1s–28s
    4. Gemini 3.5 Flash-Lite

    Milliseconds per action

    lower is better

    Median gap between consecutive step timestamps within a run.

    over runs with at least two steps

    1. Gemini 3.5 Flash-Lite
      4.4s
      p10–p90 3.0s–10s
    2. Gemini 3.1 Flash-Lite
      5.0s
      p10–p90 3.1s–8.6s
    3. Claude Haiku 4.5
      5.2s
      p10–p90 3.7s–9.9s
    4. GPT-5.6 Luna

    Steadiness

    Whether the pace is even or spiky.

    Pace steadiness

    lower is better

    Interquartile range of run durations divided by their median.

    over runs that produced a trajectory

    1. Claude Haiku 4.5
      0.63
    2. GPT-5
      0.84
    3. Gemini 3.5 Flash-Lite
      1.02
    4. GPT-5.5
      1.21

    Cost

    What does it spend, and what does a SUCCESS cost?

    Volume

    Tokens in and out.

    Input tokens

    lower is better

    Tokens sent upstream over a run.

    over runs that produced a trajectory

    1. GPT-5.6 Terra
      32k
      p10–p90 8.6k–173k
    2. GPT-5
      34k
      p10–p90 5.1k–71k
    3. Claude Sonnet 5
      36k
      p10–p90 13k–231k
    4. GPT-5.5
      45k
      p10–p90 3.2k–137k
    5. GPT-5.6 Luna
      46k
      p10–p90 11k–124k
    6. Claude Fable 5
      60k
      p10–p90 17k–485k
    7. GPT-5.6 Sol
      66k
      p10–p90 15k–138k
    8. Gemini 3.5 Flash-Lite
      66k
      p10–p90 10.0k–120k
    9. Gemini 3.1 Flash-Lite
      68k
      p10–p90 16k–167k
    10. Claude Opus 5
      69k
      p10–p90 17k–257k
    11. Gemini 3.6 Flash
      75k
      p10–p90 18k–180k
    12. Gemini 3.5 Flash
      79k
      p10–p90 14k–286k
    13. Claude Haiku 4.5
      153k
      p10–p90 35k–250k

    Output tokens

    lower is better

    Tokens generated over a run, including reasoning where the provider bills it.

    over runs that produced a trajectory

    1. Gemini 3.5 Flash-Lite
      405
      p10–p90 115–937
    2. Gemini 3.1 Flash-Lite
      450
      p10–p90 184–1.2k
    3. GPT-5.6 Terra
      895
      p10–p90 164–3.2k
    4. GPT-5.6 Sol

    Efficiency

    Spend per action, and spend per finished task — the number that decides whether you can run it.

    Output tokens per action

    lower is better

    A run's output tokens divided by its actions.

    over runs with at least one step

    1. Gemini 3.5 Flash-Lite
      31.2
      p10–p90 20.6–84.8
    2. Gemini 3.1 Flash-Lite
      32.0
      p10–p90 22.0–244
    3. GPT-5.6 Luna
      72.7
      p10–p90 49.8–507

    Expression

    How much does it think out loud, and per what?

    Reasoning

    Text the model wrote before acting.

    Reasoning per action

    fingerprint

    Characters of model-written reasoning recorded per action.

    over actions

    1. Gemini 3.6 Flash
      879
      p10–p90 749–1.4k
    2. GPT-5
      694
      p10–p90 201–1.4k
    3. Gemini 3.5 Flash
      651
      p10–p90 397–1.2k
    4. Claude Haiku 4.5
      262
      p10–p90 150–588
    5. GPT-5.6 Luna
      194
      p10–p90 55.7–498
    6. GPT-5.5
      178
      p10–p90 82.4–764
    7. GPT-5.6 Sol
      172
      p10–p90 71.5–314
    8. GPT-5.6 Terra
      140
      p10–p90 0–325
    9. Claude Opus 5
      128
      p10–p90 13.6–247
    10. Claude Fable 5
      66.3
      p10–p90 0–279
    11. Claude Sonnet 5
      42.7
      p10–p90 3.43–227
    12. Gemini 3.1 Flash-Lite
      6.06
      p10–p90 3.11–9.64
    13. Gemini 3.5 Flash-Lite
      0
      p10–p90 0–0

    Silent actions

    fingerprint

    Actions recorded with no reasoning text at all.

    over actions

    1. Gemini 3.5 Flash-Lite
      100%
      95% 99%–100%
    2. Claude Opus 5
      69%
      95% 65%–73%
    3. Claude Fable 5
      66%
      95% 62%–71%
    4. GPT-5.6 Sol
      65%
      95% 60%–69%
    5. GPT-5.6 Terra
      63%

    Output

    What did it hand back?

    Deliverable

    Files produced for the user.

    Produced a file

    fingerprint

    Runs that wrote a downloadable artifact for the user.

    over runs that produced a trajectory

    1. Claude Opus 5
      67%
      95% 51%–79%
    2. Claude Sonnet 5
      61%
      95% 47%–74%
    3. GPT-5.6 Sol
      56%
      95% 40%–70%
    4. Gemini 3.6 Flash
      55%
      95% 41%–69%
    5. Claude Fable 5
      53%
      95% 36%–70%
    6. GPT-5.6 Luna
      50%
      95% 34%–66%
    7. Gemini 3.1 Flash-Lite
      47%
      95% 32%–63%
    8. Gemini 3.5 Flash
      44%
      95% 30%–60%
    9. Gemini 3.5 Flash-Lite
      41%
      95% 26%–58%
    10. Claude Haiku 4.5
      33%
      95% 21%–48%
    11. GPT-5.6 Terra
      33%
      95% 20%–47%
    12. GPT-5
      33%
      95% 20%–48%
    13. GPT-5.5
      14%
      95% 6.7%–28%

    Form

    Shape of the final answer — the live confound on every human preference label.

    Answer length

    fingerprint

    Characters in the final answer handed to the judge.

    over runs with a final answer

    1. Claude Sonnet 5
      574
      p10–p90 257–1.3k
    2. Claude Haiku 4.5
      353
      p10–p90 211–1.2k
    3. Claude Opus 5
      340
      p10–p90 220–958
    4. Claude Fable 5
      335
      p10–p90 202–860

    Head to head

    Who beats whom, and is it consistent?

    Dominance

    The pairwise grid and what it implies.

    Head-to-head grid

    fingerprint

    Wins, losses and ties against each individual opponent, from ranked battles only.

    over ranked battles against that opponent

    1. — not measured for Claude Opus 5, Claude Sonnet 5, GPT-5.5, Claude Fable 5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.1 Flash-Lite, Gemini 3.5 Flash-Lite, Claude Haiku 4.5, GPT-5

    Variance

    Upsets, ties, and how separable its battles are.

    Beat a higher rating

    higher is better

    Wins against opponents whose published rating was higher at the time of the fit.

    over ranked battles against higher-rated opponents

    1. Claude Fable 5
      80%
      95% 38%–96%
    2. GPT-5.6 Luna
      50%
      95% 15%–85%
    3. Claude Sonnet 5
      43%
      95% 16%–75%
    4. GPT-5.5
      43%
      95% 16%–75%
    5. Gemini 3.1 Flash-Lite
      38%
      95% 21%–59%
    6. GPT-5.6 Terra
      38%
      95% 14%–69%
    7. Gemini 3.5 Flash-Lite
      33%
      95% 14%–61%
    8. Gemini 3.6 Flash
      26%
      95% 12%–49%
    9. Gemini 3.5 Flash
      25%
      95% 8.9%–53%
    10. GPT-5.6 Sol
      20%
      95% 3.6%–62%
    11. GPT-5
      15%
      95% 5.2%–36%
    12. Claude Haiku 4.5
      13%
      95% 3.5%–36%
    13. — not measured for Claude Opus 5

    Inseparable battles

    fingerprint

    Ranked battles judged a tie or both-bad.

    over ranked battles

    1. GPT-5.6 Luna
      37%
      95% 22%–56%
    2. Claude Haiku 4.5
      36%
      95% 20%–55%
    3. Gemini 3.5 Flash-Lite
      33%
      95% 17%–55%
    4. GPT-5.5
      27%
      95% 15%–44%
    5. Gemini 3.6 Flash
      26%

    Human judgment

    What did the people who watched it say went wrong?

    Flaw profile

    Which failure a judge attached to it, and how often.

    Flaw profile

    fingerprint

    Share of the agent's flaw annotations falling under each of the seven flaw labels.

    over flaw annotations on this agent's runs

    1. — not measured for Claude Opus 5, Claude Sonnet 5, GPT-5.5, Claude Fable 5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.1 Flash-Lite, Gemini 3.5 Flash-Lite, Claude Haiku 4.5, GPT-5

    Flagged runs

    lower is better

    Runs that collected at least one flaw annotation.

    over runs that were annotated at all

    1. Claude Sonnet 5
      0%
      95% 0%–79%
    2. GPT-5.5
      0%
      95% 0%–79%
    3. GPT-5.6 Luna
      0%
      95% 0%–66%
    4. GPT-5.6 Sol
      0%
      95% 0%–79%
    5. Gemini 3.6 Flash
      0%
      95% 0%–79%
    6. Gemini 3.1 Flash-Lite
      0%
      95% 0%–66%
    7. — not measured for Claude Opus 5, Claude Fable 5, GPT-5.6 Terra, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Claude Haiku 4.5, GPT-5

    Character

    The low-effort labels judges reach for.

    Character stamps

    fingerprint

    Share of the agent's stamps across the five run-level character labels.

    over stamps on this agent's runs

    1. — not measured for Claude Opus 5, Claude Sonnet 5, GPT-5.5, Claude Fable 5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.1 Flash-Lite, Gemini 3.5 Flash-Lite, Claude Haiku 4.5, GPT-5

    Preference

    Strength and separability of the labels its battles produced.

    Margin when it won

    higher is better

    Mean judge-reported strength (barely / clearly / not close, as 1-3) on battles this agent won.

    over wins where the judge tapped the optional strength chip

    1. Gemini 3.1 Flash-Lite
      2.11
    2. GPT-5.6 Luna
      2.08
    3. GPT-5
      2.00
    4. GPT-5.6 Terra
    100%
    95% 21%–100%
  • Gemini 3.1 Flash-Lite
    100%
    95% 21%–100%
  • — not measured for Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Claude Haiku 4.5, GPT-5
  • 95% 3.2%–24%
  • GPT-5.6 Sol
    11%
    95% 4.5%–26%
  • Claude Sonnet 5
    13%
    95% 5.0%–28%
  • Gemini 3.5 Flash
    15%
    95% 6.4%–30%
  • Gemini 3.6 Flash
    15%
    95% 7.1%–29%
  • GPT-5.5
    18%
    95% 8.7%–32%
  • Gemini 3.5 Flash-Lite
    28%
    95% 16%–45%
  • GPT-5.6 Luna
    29%
    95% 17%–46%
  • Gemini 3.1 Flash-Lite
    30%
    95% 17%–46%
  • Claude Haiku 4.5
    48%
    95% 33%–62%
  • 95% 1.5%–18%
  • GPT-5.5
    4.8%
    95% 1.3%–16%
  • Claude Sonnet 5
    4.5%
    95% 1.3%–15%
  • Gemini 3.5 Flash
    2.8%
    95% 0.5%–14%
  • Claude Opus 5
    2.6%
    95% 0.5%–13%
  • Gemini 3.6 Flash
    2.1%
    95% 0.4%–11%
  • Claude Fable 5
    0%
    95% 0%–11%
  • GPT-5.6 Sol
    0%
    95% 0%–9.6%
  • GPT-5
    0%
    95% 0%–8.8%
  • p10–p90 3–9.70
  • Gemini 3.1 Flash-Lite
    5
    p10–p90 3.10–18.9
  • GPT-5.5
    5.50
    p10–p90 2–18.4
  • GPT-5.6 Terra
    5.50
    p10–p90 2–25.5
  • GPT-5.6 Luna
    6
    p10–p90 2.40–20.8
  • Claude Fable 5
    7
    p10–p90 3–25.8
  • Gemini 3.6 Flash
    7
    p10–p90 3.70–15
  • Claude Haiku 4.5
    7
    p10–p90 4–27
  • GPT-5.6 Sol
    8
    p10–p90 3.60–22.4
  • Gemini 3.5 Flash-Lite
    9
    p10–p90 2–15.6
  • 0.95x
  • GPT-5.6 Luna
    1.00x
  • Gemini 3.6 Flash
    1.04x
  • Claude Fable 5
    1.06x
  • GPT-5
    1.07x
  • Gemini 3.1 Flash-Lite
    1.13x
  • Gemini 3.5 Flash
    1.16x
  • Gemini 3.5 Flash-Lite
    1.22x
  • Claude Haiku 4.5
    1.38x
  • GPT-5.6 Luna
    5.8%
    95% 4.0%–8.4%
  • Gemini 3.5 Flash
    5.9%
    95% 4.0%–8.5%
  • GPT-5.5
    6.7%
    95% 4.7%–9.4%
  • GPT-5.6 Terra
    8.4%
    95% 6.3%–11%
  • Claude Opus 5
    9.7%
    95% 7.3%–13%
  • Claude Sonnet 5
    11%
    95% 8.1%–14%
  • Claude Haiku 4.5
    17%
    95% 15%–20%
  • Claude Fable 5
    20%
    95% 16%–24%
  • Gemini 3.1 Flash-Lite
    25%
    95% 21%–29%
  • Gemini 3.5 Flash-Lite
    25%
    95% 21%–30%
  • Changed nothing

    lower is better

    Actions whose observation was identical to the previous step's, once the tab marker is stripped.

    over actions after the first

    1. GPT-5
      2.2%
      95% 1.0%–4.6%
    2. Gemini 3.6 Flash
      4.8%
      95% 3.3%–7.1%
    3. GPT-5.6 Luna
      5.6%
      95% 3.8%–8.1%
    4. GPT-5.6 Sol
      5.6%
      95% 3.8%–8.2%
    5. Gemini 3.5 Flash
      5.9%
      95% 4.0%–8.5%
    6. GPT-5.5
      6.4%
      95% 4.5%–9.1%
    7. GPT-5.6 Terra
      8.2%
      95% 6.1%–11%
    8. Claude Sonnet 5
      9.5%
      95% 7.1%–13%
    9. Claude Opus 5
      10%
      95% 7.7%–13%
    10. Claude Haiku 4.5
      19%
      95% 17%–22%
    11. Claude Fable 5
      21%
      95% 17%–25%
    12. Gemini 3.5 Flash-Lite
      24%
      95% 20%–28%
    13. Gemini 3.1 Flash-Lite
      24%
      95% 21%–28%

    Explicitly inert

    lower is better

    Actions the browser explicitly reported as having moved nothing — scrolling past the end of a document, scrolling a panel that does not scroll.

    over actions

    1. Gemini 3.5 Flash
      0%
      95% 0.0%–0.8%
    2. Gemini 3.1 Flash-Lite
      0.2%
      95% 0.0%–1.0%
    3. GPT-5.5
      0.2%
      95% 0.0%–1.2%
    4. GPT-5.6 Sol
      0.2%
      95% 0.0%–1.2%
    5. GPT-5
      0.3%
      95% 0.1%–1.8%
    6. GPT-5.6 Luna
      0.4%
      95% 0.1%–1.5%
    7. Gemini 3.6 Flash
      0.6%
      95% 0.2%–1.6%
    8. GPT-5.6 Terra
      0.6%
      95% 0.2%–1.7%
    9. Claude Opus 5
      0.6%
      95% 0.2%–1.8%
    10. Claude Sonnet 5
      0.8%
      95% 0.3%–2.1%
    11. Claude Fable 5
      1.0%
      95% 0.4%–2.6%
    12. Gemini 3.5 Flash-Lite
      1.7%
      95% 0.8%–3.4%
    13. Claude Haiku 4.5
      6.6%
      95% 5.1%–8.5%

    Stayed put

    fingerprint

    Actions after which the page URL was unchanged from the previous step.

    over actions after the first, on steps where both URLs were recorded

    1. Claude Fable 5
      89%
      95% 86%–92%
    2. Claude Haiku 4.5
      88%
      95% 86%–90%
    3. Gemini 3.5 Flash-Lite
      86%
      95% 82%–89%
    4. Gemini 3.1 Flash-Lite
      85%
      95% 81%–87%
    5. Claude Sonnet 5
      84%
      95% 80%–87%
    6. Gemini 3.6 Flash
      82%
      95% 78%–85%
    7. Gemini 3.5 Flash
      81%
      95% 77%–84%
    8. GPT-5
      80%
      95% 74%–84%
    9. GPT-5.5
      76%
      95% 72%–80%
    10. Claude Opus 5
      76%
      95% 72%–80%
    11. GPT-5.6 Terra
      76%
      95% 72%–79%
    12. GPT-5.6 Luna
      75%
      95% 70%–79%
    13. GPT-5.6 Sol
      67%
      95% 62%–71%

    Went back

    lower is better

    Steps landing on a URL the run had already visited and since left.

    over steps with a recorded URL

    1. GPT-5
      0%
      95% 0%–1.2%
    2. Claude Fable 5
      1.0%
      95% 0.4%–2.6%
    3. Gemini 3.5 Flash
      1.3%
      95% 0.6%–2.8%
    4. Gemini 3.5 Flash-Lite
      1.9%
      95% 1.0%–3.7%
    5. Claude Haiku 4.5
      2.0%
      95% 1.3%–3.2%
    6. GPT-5.5
      2.1%
      95% 1.1%–3.8%
    7. Gemini 3.6 Flash
      2.8%
      95% 1.7%–4.5%
    8. Claude Sonnet 5
      2.9%
      95% 1.7%–4.8%
    9. GPT-5.6 Terra
      3.4%
      95% 2.2%–5.3%
    10. GPT-5.6 Luna
      4.1%
      95% 2.6%–6.3%
    11. GPT-5.6 Sol
      4.1%
      95% 2.6%–6.3%
    12. Gemini 3.1 Flash-Lite
      5.1%
      95% 3.6%–7.3%
    13. Claude Opus 5
      11%
      95% 8.5%–14%
  • GPT-5.6 Terra
    15%
    95% 12%–20%
  • GPT-5
    15%
    95% 11%–20%
  • Claude Haiku 4.5
    15%
    95% 12%–17%
  • Gemini 3.6 Flash
    13%
    95% 9.8%–16%
  • GPT-5.6 Luna
    13%
    95% 9.4%–16%
  • Claude Opus 5
    10%
    95% 7.9%–14%
  • Claude Sonnet 5
    6.4%
    95% 4.4%–9.2%
  • Gemini 3.1 Flash-Lite
    2.6%
    95% 1.5%–4.5%
  • Gemini 3.5 Flash-Lite
    2.5%
    95% 1.2%–4.8%
  • Pages visited

    fingerprint

    Distinct page URLs observed across a run.

    over runs with at least one recorded URL

    1. Claude Opus 5
      2
      p10–p90 1–6.20
    2. GPT-5.5
      2
      p10–p90 1–7.40
    3. Claude Fable 5
      2
      p10–p90 1–4
    4. GPT-5.6 Luna
      2
      p10–p90 1–10
    5. GPT-5.6 Sol
      2
      p10–p90 1–12.5
    6. GPT-5.6 Terra
      2
      p10–p90 1–6
    7. Gemini 3.6 Flash
      2
      p10–p90 1–6
    8. Gemini 3.5 Flash
      2
      p10–p90 1–7.70
    9. Claude Haiku 4.5
      2
      p10–p90 1–6.60
    10. GPT-5
      2
      p10–p90 1–5.50
    11. Gemini 3.1 Flash-Lite
      1.50
      p10–p90 1–6
    12. Claude Sonnet 5
      1
      p10–p90 1–5
    13. Gemini 3.5 Flash-Lite
      1
      p10–p90 1–5.90

    Left the site

    fingerprint

    Navigations to an origin other than the first one the run observed.

    over navigate actions

    1. GPT-5.6 Sol
      57%
      95% 49%–65%
    2. GPT-5.6 Terra
      56%
      95% 47%–65%
    3. GPT-5.5
      55%
      95% 46%–64%
    4. GPT-5.6 Luna
      53%
      95% 44%–61%
    5. Gemini 3.5 Flash
      52%
      95% 42%–62%
    6. Gemini 3.6 Flash
      43%
      95% 35%–52%
    7. GPT-5
      39%
      95% 29%–49%
    8. Claude Haiku 4.5
      38%
      95% 30%–48%
    9. Claude Sonnet 5
      38%
      95% 28%–48%
    10. Gemini 3.5 Flash-Lite
      36%
      95% 27%–47%
    11. Claude Fable 5
      32%
      95% 22%–44%
    12. Claude Opus 5
      31%
      95% 24%–40%
    13. Gemini 3.1 Flash-Lite
      24%
      95% 16%–35%

    Action variety

    fingerprint

    Shannon entropy of the run's verb distribution, normalized by log2 of the twelve verbs AVAILABLE.

    over runs using at least two distinct verbs

    1. Gemini 3.6 Flash
      0.56
      p10–p90 0.30–0.70
    2. Gemini 3.5 Flash
      0.56
      p10–p90 0.39–0.69
    3. Claude Opus 5
      0.54
      p10–p90 0.34–0.70
    4. Claude Fable 5
      0.50
      p10–p90 0.38–0.64
    5. Claude Haiku 4.5
      0.48
      p10–p90 0.23–0.71
    6. Gemini 3.1 Flash-Lite
      0.47
      p10–p90 0.16–0.65
    7. Claude Sonnet 5
      0.44
      p10–p90 0.28–0.64
    8. GPT-5.5
      0.44
      p10–p90 0.28–0.61
    9. GPT-5.6 Sol
      0.44
      p10–p90 0.28–0.64
    10. GPT-5.6 Terra
      0.44
      p10–p90 0.23–0.56
    11. GPT-5
      0.44
      p10–p90 0.23–0.58
    12. Gemini 3.5 Flash-Lite
      0.42
      p10–p90 0.16–0.56
    13. GPT-5.6 Luna
      0.36
      p10–p90 0.26–0.60

    Action mix

    fingerprint

    Share of each of the twelve verbs across all of the agent's actions.

    over all actions

    1. — not measured for Claude Opus 5, Claude Sonnet 5, GPT-5.5, Claude Fable 5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.1 Flash-Lite, Gemini 3.5 Flash-Lite, Claude Haiku 4.5, GPT-5
    p10–p90 0–2
  • GPT-5.6 Luna
    0
    p10–p90 0–1
  • GPT-5.6 Sol
    0
    p10–p90 0–1
  • GPT-5.6 Terra
    0
    p10–p90 0–1
  • Gemini 3.6 Flash
    0
    p10–p90 0–1
  • Gemini 3.5 Flash
    0
    p10–p90 0–0
  • Gemini 3.1 Flash-Lite
    0
    p10–p90 0–0
  • Gemini 3.5 Flash-Lite
    0
    p10–p90 0–0.90
  • GPT-5
    0
    p10–p90 0–0
  • Claude Haiku 4.5
    1
    p10–p90 0–2.60
  • Gemini 3.1 Flash-Lite
    100%
    95% 44%–100%
  • GPT-5
    100%
    95% 34%–100%
  • Claude Sonnet 5
    96%
    95% 80%–99%
  • Gemini 3.6 Flash
    92%
    95% 65%–99%
  • Claude Haiku 4.5
    78%
    95% 68%–85%
  • GPT-5.6 Terra
    75%
    95% 51%–90%
  • Gemini 3.5 Flash-Lite
    75%
    95% 41%–93%
  • Claude Fable 5
    71%
    95% 53%–84%
  • Claude Opus 5
    69%
    95% 51%–82%
  • GPT-5.6 Luna
    60%
    95% 39%–78%
  • Retried the same thing

    lower is better

    Failing actions immediately followed by the byte-identical action — the same call, expecting a different answer.

    over failing actions that had a next action

    1. Claude Sonnet 5
      0%
      95% 0%–13%
    2. GPT-5.5
      0%
      95% 0%–35%
    3. GPT-5.6 Sol
      0%
      95% 0%–32%
    4. GPT-5.6 Terra
      0%
      95% 0%–19%
    5. Gemini 3.6 Flash
      0%
      95% 0%–24%
    6. Gemini 3.5 Flash
      0%
      95% 0%–43%
    7. Gemini 3.1 Flash-Lite
      0%
      95% 0%–56%
    8. Gemini 3.5 Flash-Lite
      0%
      95% 0%–32%
    9. GPT-5
      0%
      95% 0%–66%
    10. Claude Opus 5
      6.3%
      95% 1.7%–20%
    11. Claude Fable 5
      9.7%
      95% 3.3%–25%
    12. Claude Haiku 4.5
      13%
      95% 7.4%–22%
    13. GPT-5.6 Luna
      15%
      95% 5.2%–36%
    95% 1.4%–5.0%
  • Gemini 3.6 Flash
    3.6%
    95% 1.2%–10.0%
  • Claude Fable 5
    3.9%
    95% 1.4%–11%
  • Gemini 3.5 Flash
    3.9%
    95% 1.4%–11%
  • GPT-5.6 Luna
    5.2%
    95% 2.2%–12%
  • Claude Sonnet 5
    5.3%
    95% 2.8%–9.7%
  • GPT-5.6 Terra
    5.3%
    95% 2.1%–13%
  • GPT-5.6 Sol
    6.1%
    95% 1.7%–20%
  • Gemini 3.5 Flash-Lite
    11%
    95% 7.7%–17%
  • Gemini 3.1 Flash-Lite
    29%
    95% 24%–35%
  • 4.8%
    95% 1.9%–12%
  • GPT-5.6 Sol
    6.1%
    95% 1.7%–20%
  • GPT-5
    7.6%
    95% 3.5%–16%
  • GPT-5.6 Luna
    11%
    95% 6.5%–19%
  • Claude Opus 5
    12%
    95% 6.8%–19%
  • Claude Fable 5
    13%
    95% 7.3%–23%
  • Claude Sonnet 5
    39%
    95% 32%–46%
  • Claude Haiku 4.5
    51%
    95% 46%–56%
  • Gemini 3.1 Flash-Lite
    55%
    95% 49%–61%
  • Gemini 3.5 Flash-Lite
    68%
    95% 61%–74%
  • p10–p90 0.07–0.31
  • GPT-5.6 Terra
    0.19
    p10–p90 0.09–0.29
  • Claude Haiku 4.5
    0.16
    p10–p90 0.00–0.29
  • GPT-5.6 Luna
    0.16
    p10–p90 0.06–0.24
  • Gemini 3.5 Flash
    0.15
    p10–p90 0.07–0.25
  • GPT-5.5
    0.14
    p10–p90 0.05–0.29
  • GPT-5
    0.14
    p10–p90 0.00–0.29
  • Gemini 3.5 Flash-Lite
    0.11
    p10–p90 0.00–0.25
  • Gemini 3.1 Flash-Lite
    0.07
    p10–p90 0.00–0.26
  • 3
    p10–p90 2–3.20
  • Claude Fable 5
    3
    p10–p90 3–3
  • Gemini 3.1 Flash-Lite
    2
    p10–p90 2–9.20
  • Claude Haiku 4.5
    2
    p10–p90 2–2
  • — not measured for GPT-5.5, GPT-5.6 Luna, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, GPT-5
  • Endpoint-only drags

    lower is better

    Drags supplied with two points or fewer.

    over drag actions

    1. Claude Opus 5
      0%
      95% 0%–43%
    2. Claude Fable 5
      0%
      95% 0%–79%
    3. GPT-5.6 Sol
      0%
      95% 0%–79%
    4. GPT-5.6 Terra
      0%
      95% 0%–79%
    5. Claude Sonnet 5
      30%
      95% 15%–52%
    6. Gemini 3.1 Flash-Lite
      75%
      95% 41%–93%
    7. Claude Haiku 4.5
      100%
      95% 44%–100%
    8. — not measured for GPT-5.5, GPT-5.6 Luna, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, GPT-5

    Modified clicks

    fingerprint

    Clicks holding Alt, Control, Meta or Shift.

    over click actions

    1. GPT-5.6 Terra
      22%
      95% 14%–33%
    2. GPT-5.6 Sol
      3.0%
      95% 0.5%–15%
    3. Claude Fable 5
      2.6%
      95% 0.7%–9.1%
    4. Claude Opus 5
      0%
      95% 0%–3.6%
    5. Claude Sonnet 5
      0%
      95% 0%–2.2%
    6. GPT-5.5
      0%
      95% 0%–7.3%
    7. GPT-5.6 Luna
      0%
      95% 0%–3.8%
    8. Gemini 3.6 Flash
      0%
      95% 0%–4.4%
    9. Gemini 3.5 Flash
      0%
      95% 0%–4.8%
    10. Gemini 3.1 Flash-Lite
      0%
      95% 0%–1.5%
    11. Gemini 3.5 Flash-Lite
      0%
      95% 0%–2.0%
    12. Claude Haiku 4.5
      0%
      95% 0%–1.1%
    13. GPT-5
      0%
      95% 0%–4.6%
    GPT-5.6 Terra
    14
    p10–p90 3–1.1k
  • Claude Fable 5
    13
    p10–p90 6–481
  • GPT-5.6 Sol
    11
    p10–p90 8.60–13.4
  • Gemini 3.5 Flash-Lite
    11
    p10–p90 3.30–32.5
  • Claude Opus 5
    7
    p10–p90 4–48.2
  • Claude Haiku 4.5
    7
    p10–p90 5–23.0
  • GPT-5.5
    5
    p10–p90 4.70–25.3
  • GPT-5.6 Luna
    5
    p10–p90 3.60–10
  • Gemini 3.6 Flash
    5
    p10–p90 5–63
  • Keys vs strings

    fingerprint

    `press_key` actions as a share of all text-entry actions (type plus press_key).

    over type plus press_key actions

    1. GPT-5.6 Sol
      90%
      95% 71%–97%
    2. Claude Sonnet 5
      83%
      95% 73%–90%
    3. Claude Opus 5
      67%
      95% 58%–75%
    4. Gemini 3.6 Flash
      66%
      95% 56%–75%
    5. GPT-5.5
      60%
      95% 45%–73%
    6. Claude Haiku 4.5
      60%
      95% 48%–71%
    7. GPT-5.6 Luna
      58%
      95% 45%–70%
    8. Claude Fable 5
      54%
      95% 42%–66%
    9. GPT-5.6 Terra
      53%
      95% 30%–75%
    10. Gemini 3.5 Flash
      51%
      95% 38%–64%
    11. GPT-5
      50%
      95% 24%–76%
    12. Gemini 3.1 Flash-Lite
      40%
      95% 30%–51%
    13. Gemini 3.5 Flash-Lite
      33%
      95% 16%–56%
    17s
    p10–p90 13s–22s
  • Claude Opus 5
    18s
    p10–p90 13s–24s
  • Gemini 3.6 Flash
    19s
    p10–p90 8.3s–29s
  • GPT-5.6 Terra
    19s
    p10–p90 15s–24s
  • GPT-5.6 Luna
    20s
    p10–p90 14s–30s
  • GPT-5.6 Sol
    20s
    p10–p90 15s–25s
  • Claude Fable 5
    20s
    p10–p90 16s–33s
  • GPT-5.5
    20s
    p10–p90 15s–27s
  • Gemini 3.5 Flash
    21s
    p10–p90 15s–33s
  • GPT-5
    23s
    p10–p90 11s–36s
  • 6.3s
    p10–p90 4.7s–9.8s
  • GPT-5.6 Terra
    6.8s
    p10–p90 4.2s–12s
  • Gemini 3.6 Flash
    6.9s
    p10–p90 5.1s–12s
  • Claude Sonnet 5
    7.0s
    p10–p90 4.8s–14s
  • GPT-5.6 Sol
    7.0s
    p10–p90 5.0s–13s
  • GPT-5.5
    8.0s
    p10–p90 5.4s–16s
  • Claude Opus 5
    8.0s
    p10–p90 5.5s–15s
  • Gemini 3.5 Flash
    8.3s
    p10–p90 5.5s–13s
  • Claude Fable 5
    10.0s
    p10–p90 7.5s–24s
  • GPT-5
    13s
    p10–p90 8.1s–23s
  • GPT-5.6 Sol
    1.23
  • Claude Fable 5
    1.31
  • GPT-5.6 Terra
    1.33
  • Gemini 3.1 Flash-Lite
    1.47
  • GPT-5.6 Luna
    1.47
  • Gemini 3.5 Flash
    1.59
  • Gemini 3.6 Flash
    1.69
  • Claude Opus 5
    1.73
  • Claude Sonnet 5
    1.96
  • 973
    p10–p90 278–3.2k
  • GPT-5.5
    1.1k
    p10–p90 110–4.5k
  • GPT-5.6 Luna
    1.2k
    p10–p90 285–4.5k
  • Claude Sonnet 5
    1.2k
    p10–p90 420–4.7k
  • Claude Opus 5
    1.4k
    p10–p90 609–6.8k
  • Claude Fable 5
    1.7k
    p10–p90 392–8.5k
  • Gemini 3.5 Flash
    1.9k
    p10–p90 349–7.3k
  • Gemini 3.6 Flash
    2.0k
    p10–p90 792–6.3k
  • Claude Haiku 4.5
    2.4k
    p10–p90 845–13k
  • GPT-5
    4.4k
    p10–p90 648–6.3k
  • GPT-5.6 Terra
    80.5
    p10–p90 49.6–268
  • GPT-5.6 Sol
    83.8
    p10–p90 46.5–246
  • GPT-5.5
    119
    p10–p90 47.8–367
  • Claude Haiku 4.5
    146
    p10–p90 103–453
  • Gemini 3.5 Flash
    158
    p10–p90 96.5–398
  • Claude Fable 5
    177
    p10–p90 106–677
  • Claude Sonnet 5
    199
    p10–p90 129–491
  • Gemini 3.6 Flash
    202
    p10–p90 119–530
  • Claude Opus 5
    215
    p10–p90 116–430
  • GPT-5
    402
    p10–p90 240–675
  • Input tokens per action

    lower is better

    A run's input tokens divided by its actions.

    over runs with at least one step

    1. GPT-5
      4.4k
      p10–p90 3.6k–5.5k
    2. GPT-5.6 Luna
      4.9k
      p10–p90 3.5k–6.7k
    3. GPT-5.5
      5.0k
      p10–p90 3.4k–6.2k
    4. Gemini 3.5 Flash-Lite
      5.0k
      p10–p90 4.0k–6.8k
    5. GPT-5.6 Sol
      5.3k
      p10–p90 3.8k–8.4k
    6. GPT-5.6 Terra
      5.4k
      p10–p90 3.7k–7.7k
    7. Gemini 3.1 Flash-Lite
      5.7k
      p10–p90 4.4k–6.6k
    8. Claude Sonnet 5
      7.0k
      p10–p90 5.5k–10k
    9. Gemini 3.6 Flash
      7.3k
      p10–p90 4.7k–11k
    10. Claude Opus 5
      7.3k
      p10–p90 5.5k–13k
    11. Claude Fable 5
      8.2k
      p10–p90 5.6k–12k
    12. Claude Haiku 4.5
      8.2k
      p10–p90 6.3k–11k
    13. Gemini 3.5 Flash
      8.5k
      p10–p90 4.9k–15k

    Tokens per finished task

    lower is better

    Total tokens across all of an agent's runs divided by the number that reported success.

    over runs that reported success

    1. GPT-5.6 Luna
      99k
    2. GPT-5.6 Terra
      101k
    3. GPT-5
      110k
    4. GPT-5.5
      123k
    5. GPT-5.6 Sol
      124k
    6. Gemini 3.1 Flash-Lite
      146k
    7. Claude Sonnet 5
      148k
    8. Gemini 3.5 Flash-Lite
      161k
    9. Claude Opus 5
      177k
    10. Claude Fable 5
      190k
    11. Gemini 3.6 Flash
      197k
    12. Gemini 3.5 Flash
      312k
    13. Claude Haiku 4.5
      459k
    95% 59%–67%
  • GPT-5.6 Luna
    62%
    95% 57%–66%
  • GPT-5.5
    56%
    95% 52%–61%
  • Claude Sonnet 5
    53%
    95% 48%–57%
  • GPT-5
    34%
    95% 29%–39%
  • Gemini 3.1 Flash-Lite
    22%
    95% 19%–26%
  • Gemini 3.6 Flash
    0%
    95% 0%–0.7%
  • Gemini 3.5 Flash
    0%
    95% 0.0%–0.8%
  • Claude Haiku 4.5
    0%
    95% 0%–0.5%
  • Gemini 3.6 Flash
    298
    p10–p90 120–1.0k
  • GPT-5
    278
    p10–p90 189–695
  • GPT-5.6 Sol
    265
    p10–p90 106–459
  • Gemini 3.5 Flash
    217
    p10–p90 98.1–659
  • Gemini 3.5 Flash-Lite
    174
    p10–p90 91–320
  • GPT-5.6 Terra
    170
    p10–p90 91.5–318
  • Gemini 3.1 Flash-Lite
    161
    p10–p90 67.6–306
  • GPT-5.6 Luna
    159
    p10–p90 73–465
  • GPT-5.5
    151
    p10–p90 85.4–335
  • Empty answers

    lower is better

    Runs that reported success with a final answer under 20 characters.

    over runs that reported success

    1. Claude Opus 5
      0%
      95% 0%–12%
    2. Claude Sonnet 5
      0%
      95% 0%–11%
    3. GPT-5.5
      0%
      95% 0.0%–15%
    4. Claude Fable 5
      0%
      95% 0%–14%
    5. GPT-5.6 Luna
      0%
      95% 0%–13%
    6. GPT-5.6 Sol
      0%
      95% 0%–12%
    7. GPT-5.6 Terra
      0%
      95% 0%–10%
    8. Gemini 3.6 Flash
      0%
      95% 0%–12%
    9. Gemini 3.5 Flash
      0%
      95% 0%–18%
    10. Gemini 3.1 Flash-Lite
      0%
      95% 0.0%–15%
    11. Gemini 3.5 Flash-Lite
      0%
      95% 0%–20%
    12. Claude Haiku 4.5
      0%
      95% 0%–18%
    13. GPT-5
      0%
      95% 0%–22%
    95% 15%–41%
  • Claude Opus 5
    25%
    95% 13%–42%
  • GPT-5.6 Terra
    24%
    95% 12%–42%
  • GPT-5
    23%
    95% 11%–42%
  • Gemini 3.5 Flash
    23%
    95% 10%–43%
  • GPT-5.6 Sol
    21%
    95% 9.2%–40%
  • Gemini 3.1 Flash-Lite
    21%
    95% 9.8%–38%
  • Claude Fable 5
    18%
    95% 7.3%–39%
  • Claude Sonnet 5
    17%
    95% 8.1%–33%
  • 1.95
  • GPT-5.5
    1.94
  • Claude Fable 5
    1.90
  • Gemini 3.5 Flash-Lite
    1.83
  • Gemini 3.6 Flash
    1.79
  • Claude Sonnet 5
    1.75
  • Claude Haiku 4.5
    1.67
  • Claude Opus 5
    1.63
  • Gemini 3.5 Flash
    1.56
  • GPT-5.6 Sol
    1.45