Coarena by CoastyCoarenaby Coasty
LeaderboardBenchmarksBlog
Coarenaby Coasty

Real-world evals for computer-use agents. Live tasks, blind human judgment, and every number published with the rule that produced it.

Arena

  • Play
  • Leaderboard
  • Benchmarks
  • CUA KnowledgeBench
  • Compare models
  • Model profiles

Evidence

  • Dataset
  • Methodology
  • Metrics API
  • Cite us

About

  • Mission
  • Governance
  • Blog

Legal

  • Terms
  • Privacy
  • Security

© 2026 Coasty Systems, Inc.

Every claim on this site cites the file that keeps it

Benchmarks

Real-world evaluations of computer-use agents.

Computer Use Arena

Real-world tasks, two agents, and blind human comparisons. Explore how models use a computer and how their performance is measured. Public rankings and published metrics.

CUA KnowledgeBench 12-Hours & 1-Day

Reasoning, current information, and connected work across more than 100 applications. Private evaluation: contact us to arrange one.

CUA KnowledgeBench 3-Days & 7-Days

Coming soon: longer evaluations of sustained knowledge work. Availability and evaluation details will be announced.

Published arena metrics

57 measures in 10 families, each with its definition, denominator and direction. They describe recorded behaviour in Computer Use Arena, not task correctness; KnowledgeBench results are separate. Also published as JSON.

Outcome10 metrics

Did it do the job, and does it know whether it did?

Completion

How runs ended, by the agent's own report.

Completed
higher is better

Runs whose terminal `done` call reported success.

Over runs that produced a trajectory (harness errors excluded)

  1. 1GPT-5.6 Luna
    96%
    95% 88%–99%
  2. 2GPT-5.6 Terra
    89%
    95% 74%–95%
  3. 3Gemini 3.5 Flash
    80%
    95% 68%–88%
  4. 4Claude Sonnet 5
    80%
    95% 67%–88%
  5. 5Inkling
    78%
    95% 67%–86%
  6. 6Claude Opus 5
    76%
    95% 63%–86%
  7. 7GPT-5.6 Sol
    75%
    95% 57%–87%
  8. 8Qwen3.8 Max
    72%
    95% 68%–76%
  9. 9Claude Fable 5
    70%
    95% 53%–83%
  10. 10Gemini 3.8 Flash
    67%
    95% 63%–71%
  11. 11Claude Fable 5.1
    61%
    95% 51%–71%
  12. 12Claude Haiku 4.5
    59%
    95% 45%–72%
  13. 13Muse Spark 1.2
    56%
    95% 42%–68%
  14. 14Kimi K3
    54%
    95% 40%–67%
  15. 15Grok 4.6
    35%
    95% 23%–48%
  16. 16GPT-6 Astra
    33%
    95% 24%–43%
  17. 17GLM 5.3 Flash
    16%
    95% 9.1%–26%
Gave up
lower is better

Runs whose terminal `done` call reported failure, plus runs that stopped without calling done.

Over runs that produced a trajectory

  1. 1Claude Opus 5
    0%
    95% 0%–7.1%
  2. 2Claude Fable 5.1
    1.2%
    95% 0.2%–6.5%
  3. 3GPT-5.6 Luna
    1.8%
    95% 0.3%–9.3%
  4. 4Claude Sonnet 5
    1.9%
    95% 0.3%–9.8%
Called impossible
fingerprint

Runs where the agent reported the objective does not exist on this environment: the control is gone, a login wall bars it.

Over runs that produced a trajectory

  1. Gemini 3.8 Flash
    4.5%
    95% 3.1%–6.5%
  2. Qwen3.8 Max
    2.9%
    95% 1.8%–4.7%
  3. Claude Fable 5
    6.1%
    95% 1.7%–20%
  4. GPT-5.6 Sol
    0%
    95% 0%–12%
  5. Claude Fable 5.1
    3.6%
    95% 1.2%–10%
  6. GPT-5.6 Luna
Ran out of clock
lower is better

Runs that hit the shared wall-clock limit before finishing.

Over runs that produced a trajectory

  1. 1Inkling
    0%
    95% 0%–5.3%
  2. 2GPT-5.6 Luna
    1.8%
    95% 0.3%–9.3%
  3. 3GPT-5.6 Terra
    5.7%
    95% 1.6%–19%
  4. 4Claude Sonnet 5
    9.3%
    95% 4.0%–20%
No trajectory
lower is better

Runs that produced no trajectory at all: an upstream outage, a rejected credential, a withdrawn model.

Over all runs, including ones with no trajectory

  1. 1GPT-5.6 Sol
    0%
    95% 0%–12%
  2. 1GPT-5.6 Luna
    0%
    95% 0%–6.3%
  3. 1Claude Haiku 4.5
    0%
    95% 0%–7.3%
  4. 1Muse Spark 1.2
    0%
    95% 0%–6.9%

Self-knowledge

Where the agent's own verdict is checkable against something else: the opponent on the same task, or a judge who read the answer.

Wrongly called impossible
lower is better

Runs where the agent declared the task infeasible while its opponent COMPLETED the identical task in the same battle.

Over runs declared infeasible whose opponent produced a terminal verdict

  1. 1Qwen3.8 Max
    0%
    95% 0%–19%
  2. 1Claude Fable 5
    0%
    95% 0%–66%
  3. 1Gemini 3.5 Flash
    0%
    95% 0%–66%

Termination

How the run stopped: on purpose, or because something ran out.

Stopped on purpose
higher is better

Runs whose final action was an explicit `done` call rather than being cut off.

Over runs that produced a trajectory

  1. 1GPT-5.6 Luna
    98%
    95% 91%–100%
  2. 2GPT-5.6 Terra
    94%
    95% 81%–98%
  3. 3Claude Sonnet 5
    91%
    95% 80%–96%
  4. 4

Route13 metrics

How direct was the path, and how much of it was wasted motion?

Economy

Actions spent, and how that compares to the other agent on the same task.

Actions taken
lower is better

Actions executed in a run, including the terminal one.

Over runs that produced a trajectory

  1. 1Claude Sonnet 5
    13.5
    p10–p90 2.30–69.1
  2. 2Inkling
    17
    p10–p90 8–50.3
  3. 3GPT-5.6 Luna

Recovery4 metrics

When something went wrong, what did it do next?

Failure exposure

How often actions failed at all.

Actions that failed
lower is better

Actions whose observation came back as an error.

Over actions

  1. 1GPT-5.6 Sol
    0.8%
    95% 0.4%–1.6%
  2. 2Claude Fable 5.1
    1.1%
    95% 0.7%–1.6%
  3. 3GPT-6 Astra

Precision9 metrics

Does it hit what it aims at? The pixel layer, which only a computer-use arena can see.

Pointing

Where clicks land, how they cluster, and whether they repeat.

Clicks that did nothing
lower is better

Clicks whose own observation reported an error or an unchanged page.

Over click actions

Clicks that did nothing figures on request: founders@coasty.ai

Clicks at the frame edge
lower is better

Clicks within 24px of any viewport boundary.

Over click actions

Clicks at the frame edge figures on request: founders@coasty.ai

Clicked the same pixel again

Tempo4 metrics

How fast, how steady, and how bad is the tail?

Latency

Wall clock, including the tails a median hides.

Run duration
lower is better

Wall-clock milliseconds from run start to terminal state.

Over runs that produced a trajectory

  1. 1Inkling
    1m 58s
    p10–p90 42s–4m 05s
  2. 2GPT-5.6 Luna
    2m 33s
    p10–p90 57s–6m 04s
  3. 3Claude Sonnet 5

Cost5 metrics

What does it spend, and what does a SUCCESS cost?

Volume

Tokens in and out.

Input tokens
lower is better

Tokens sent upstream over a run.

Over runs that produced a trajectory

  1. 1GPT-5.6 Terra
    178k
    p10–p90 17k–974k
  2. 2Claude Sonnet 5
    184k
    p10–p90 19k–1.9M
  3. 3GPT-5.6 Luna

Expression2 metrics

How much does it think out loud, and per what?

Reasoning

Text the model wrote before acting.

Reasoning per action
fingerprint

Characters of model-written reasoning recorded per action.

Over actions

  1. Gemini 3.8 Flash
    2.0k
    p10–p90 1.5k–3.3k
  2. Qwen3.8 Max
    3.7k
    p10–p90 1.4k–8.7k
  3. Claude Fable 5
    226
    p10–p90 99.3–356
  4. GPT-5.6 Sol

Output3 metrics

What did it hand back?

Deliverable

Files produced for the user.

Produced a file
fingerprint

Runs that wrote a downloadable artifact for the user.

Over runs that produced a trajectory

  1. Gemini 3.8 Flash
    72%
    95% 68%–76%
  2. Qwen3.8 Max
    77%
    95% 74%–81%
  3. Claude Fable 5
    73%
    95% 56%–85%
  4. GPT-5.6 Sol

Head to head3 metrics

Who beats whom, and is it consistent?

Dominance

The pairwise grid and what it implies.

Head-to-head grid
fingerprint

Wins, losses and ties against each individual opponent, from ranked battles only.

Over ranked battles against that opponent

Gemini 3.8 FlashBy opponent
Claude Fable 5
71.4%
Claude Fable 5.1
43.1%
Gemini 3.5 Flash
50.0%
gemini-3-7-flash
50.0%
GLM 5.3 Flash
70.0%
GPT-5.6 Luna
72.7%
GPT-5.6 Sol
48.8%
GPT-5.6 Terra
81.8%
GPT-6 Astra
79.2%
Grok 4.6
70.6%
Claude Haiku 4.5
66.7%
Inkling
36.4%
Kimi K3
45.5%
Muse Spark 1.2
55.9%
Claude Opus 5
85.7%
Qwen3.8 Max
47.7%

Human judgment4 metrics

What did the people who watched it say went wrong?

Flaw profile

Which failure a judge attached to it, and how often.

Flaw profile
fingerprint

Share of the agent's flaw annotations falling under each of the seven flaw labels.

Over flaw annotations on this agent's runs

Gemini 3.8 FlashComposition
output contamination
0.0%
wrong answer
7.1%
incomplete
83.9%
unnecessary navigation
1.8%
repeated action
0.0%
misclick
7.1%
slow recovery
0.0%
Qwen3.8 MaxComposition
output contamination
0.0%
wrong answer
1.8%
incomplete
81.8%
  • 5GPT-5.6 Terra
    2.9%
    95% 0.5%–15%
  • 6Claude Fable 5
    3.0%
    95% 0.5%–15%
  • 7Gemini 3.8 Flash
    3.8%
    95% 2.5%–5.7%
  • 8Kimi K3
    6.0%
    95% 2.1%–16%
  • 9Gemini 3.5 Flash
    6.8%
    95% 2.7%–16%
  • 10GLM 5.3 Flash
    7.2%
    95% 3.1%–16%
  • 11Grok 4.6
    7.3%
    95% 2.9%–17%
  • 12Qwen3.8 Max
    9.2%
    95% 7.0%–12%
  • 13GPT-5.6 Sol
    11%
    95% 3.7%–27%
  • 14Muse Spark 1.2
    21%
    95% 12%–34%
  • 15Inkling
    22%
    95% 14%–33%
  • 16Claude Haiku 4.5
    24%
    95% 15%–38%
  • 17GPT-6 Astra
    31%
    95% 22%–41%
  • 0%
    95% 0%–6.3%
  • Claude Haiku 4.5
    2.0%
    95% 0.4%–11%
  • Kimi K3
    0%
    95% 0%–7.1%
  • Muse Spark 1.2
    0%
    95% 0%–6.9%
  • Gemini 3.5 Flash
    3.4%
    95% 0.9%–12%
  • Claude Sonnet 5
    9.3%
    95% 4.0%–20%
  • Claude Opus 5
    0%
    95% 0%–7.1%
  • GPT-5.6 Terra
    2.9%
    95% 0.5%–15%
  • Inkling
    0%
    95% 0%–5.3%
  • Grok 4.6
    9.1%
    95% 3.9%–20%
  • GPT-6 Astra
    4.3%
    95% 1.7%–10%
  • GLM 5.3 Flash
    0%
    95% 0%–5.3%
  • 5Gemini 3.5 Flash
    10%
    95% 4.7%–20%
  • 6GPT-5.6 Sol
    14%
    95% 5.7%–31%
  • 6Claude Haiku 4.5
    14%
    95% 7.1%–27%
  • 8Qwen3.8 Max
    16%
    95% 13%–19%
  • 9Claude Fable 5
    21%
    95% 11%–38%
  • 10Muse Spark 1.2
    23%
    95% 14%–36%
  • 11Claude Opus 5
    24%
    95% 14%–37%
  • 12Gemini 3.8 Flash
    24%
    95% 21%–28%
  • 13GPT-6 Astra
    32%
    95% 23%–42%
  • 14Claude Fable 5.1
    34%
    95% 24%–44%
  • 15Kimi K3
    40%
    95% 28%–54%
  • 16Grok 4.6
    49%
    95% 36%–62%
  • 17GLM 5.3 Flash
    77%
    95% 66%–85%
  • 1Claude Sonnet 5
    0%
    95% 0%–6.6%
  • 1GPT-5.6 Terra
    0%
    95% 0%–9.9%
  • 1Inkling
    0%
    95% 0%–5.3%
  • 8Gemini 3.8 Flash
    0.7%
    95% 0.3%–1.8%
  • 9GPT-6 Astra
    1.1%
    95% 0.2%–5.7%
  • 10Qwen3.8 Max
    1.1%
    95% 0.5%–2.3%
  • 11Grok 4.6
    1.8%
    95% 0.3%–9.4%
  • 12Kimi K3
    2.0%
    95% 0.3%–10%
  • 12Claude Opus 5
    2.0%
    95% 0.3%–10%
  • 14Gemini 3.5 Flash
    3.3%
    95% 0.9%–11%
  • 15Claude Fable 5.1
    5.7%
    95% 2.5%–13%
  • 16GLM 5.3 Flash
    14%
    95% 7.8%–23%
  • 17Claude Fable 5
    18%
    95% 8.7%–32%
  • 1Grok 4.6
    0%
    95% 0%–43%
  • 5Gemini 3.8 Flash
    19%
    95% 8.5%–38%
  • 6Claude Fable 5.1
    33%
    95% 6.1%–79%
  • 7Claude Sonnet 5
    40%
    95% 12%–77%
  • 8GPT-6 Astra
    50%
    95% 15%–85%
  • 9Claude Haiku 4.5
    100%
    95% 21%–100%
  • 9GPT-5.6 Terra
    100%
    95% 21%–100%
  • Claimed success, judged wrong
    lower is better

    Runs that reported success and then collected a human 'wrong or made up' or 'didn't finish the job' label.

    Over runs that reported success AND received at least one flaw annotation

    1. 1Qwen3.8 Max
      77%
      95% 61%–88%
    2. 2GPT-5.6 Terra
      79%
      95% 57%–91%
    3. 3Kimi K3
      83%
      95% 55%–95%
    4. 4Claude Fable 5.1
      84%
      95% 67%–93%
    5. 5Claude Fable 5
      88%
      95% 53%–98%
    6. 5Muse Spark 1.2
      88%
      95% 64%–97%
    7. 5GLM 5.3 Flash
      88%
      95% 53%–98%
    8. 8Gemini 3.8 Flash
      91%
      95% 76%–97%
    9. 9Inkling
      92%
      95% 75%–98%
    10. 10GPT-6 Astra
      92%
      95% 67%–99%
    11. 11Claude Opus 5
      93%
      95% 69%–99%
    12. 12Claude Sonnet 5
      93%
      95% 70%–99%
    13. 13Gemini 3.5 Flash
      96%
      95% 79%–99%
    14. 14GPT-5.6 Luna
      96%
      95% 80%–99%
    15. 15GPT-5.6 Sol
      100%
      95% 76%–100%
    16. 15Claude Haiku 4.5
      100%
      95% 78%–100%
    17. 15Grok 4.6
      100%
      95% 72%–100%
    GPT-5.6 Sol
    86%
    95% 69%–94%
  • 5Gemini 3.5 Flash
    83%
    95% 72%–91%
  • 6Inkling
    79%
    95% 68%–87%
  • 7Claude Opus 5
    76%
    95% 63%–86%
  • 8Qwen3.8 Max
    76%
    95% 72%–79%
  • 9Claude Fable 5
    76%
    95% 59%–87%
  • 10Gemini 3.8 Flash
    72%
    95% 68%–75%
  • 11GPT-6 Astra
    68%
    95% 58%–77%
  • 12Claude Fable 5.1
    66%
    95% 56%–76%
  • 13Claude Haiku 4.5
    65%
    95% 51%–77%
  • 14Kimi K3
    56%
    95% 42%–69%
  • 15Muse Spark 1.2
    56%
    95% 42%–68%
  • 16Grok 4.6
    51%
    95% 38%–64%
  • 17GLM 5.3 Flash
    17%
    95% 10%–28%
  • Used every action
    lower is better

    Runs that consumed their entire step budget.

    Over runs that produced a trajectory and recorded a budget

    1. 1Claude Fable 5
      0%
      95% 0%–10%
    2. 1GPT-5.6 Sol
      0%
      95% 0%–12%
    3. 1Claude Fable 5.1
      0%
      95% 0%–4.4%
    4. 1GPT-5.6 Luna
      0%
      95% 0%–6.3%
    5. 1Claude Sonnet 5
      0%
      95% 0%–6.6%
    6. 1Claude Opus 5
      0%
      95% 0%–7.1%
    7. 1GPT-5.6 Terra
      0%
      95% 0%–9.9%
    8. 1Grok 4.6
      0%
      95% 0%–6.5%
    9. 1GPT-6 Astra
      0%
      95% 0.0%–3.9%
    10. 10Gemini 3.8 Flash
      3.6%
      95% 2.4%–5.5%
    11. 11Muse Spark 1.2
      3.8%
      95% 1.1%–13%
    12. 12Kimi K3
      4.0%
      95% 1.1%–13%
    13. 13Inkling
      4.4%
      95% 1.5%–12%
    14. 14GLM 5.3 Flash
      5.8%
      95% 2.3%–14%
    15. 15Gemini 3.5 Flash
      8.5%
      95% 3.7%–18%
    16. 16Qwen3.8 Max
      9.0%
      95% 6.9%–12%
    17. 17Claude Haiku 4.5
      14%
      95% 7.1%–27%
    Asked for more
    fingerprint

    Runs that were granted at least one continuation by the battle's creator.

    Over runs that produced a trajectory

    1. Gemini 3.8 Flash
      0.2%
      95% 0.0%–1.0%
    2. Qwen3.8 Max
      0.2%
      95% 0.0%–1.0%
    3. Claude Fable 5
      0%
      95% 0%–10%
    4. GPT-5.6 Sol
      0%
      95% 0%–12%
    5. Claude Fable 5.1
      0%
      95% 0%–4.4%
    6. GPT-5.6 Luna
      0%
      95% 0%–6.3%
    7. Claude Haiku 4.5
      4.1%
      95% 1.1%–14%
    8. Kimi K3
      0%
      95% 0%–7.1%
    9. Muse Spark 1.2
      0%
      95% 0%–6.9%
    10. Gemini 3.5 Flash
      1.7%
      95% 0.3%–9.0%
    11. Claude Sonnet 5
      0%
      95% 0%–6.6%
    12. Claude Opus 5
      0%
      95% 0%–7.1%
    13. GPT-5.6 Terra
      0%
      95% 0%–9.9%
    14. Inkling
      0%
      95% 0%–5.3%
    15. Grok 4.6
      0%
      95% 0%–6.5%
    16. GPT-6 Astra
      0%
      95% 0.0%–3.9%
    17. GLM 5.3 Flash
      0%
      95% 0%–5.3%
    19
    p10–p90 7.20–49.4
  • 4Claude Fable 5.1
    23
    p10–p90 2.20–65.2
  • 5Kimi K3
    25
    p10–p90 2–84.1
  • 6GPT-5.6 Terra
    26
    p10–p90 2.80–79.2
  • 7Qwen3.8 Max
    27
    p10–p90 3–98.5
  • 8Gemini 3.8 Flash
    28
    p10–p90 4–82
  • 9GPT-5.6 Sol
    29.5
    p10–p90 2.70–64.2
  • 10Claude Fable 5
    30
    p10–p90 4–65.4
  • 11Claude Opus 5
    34
    p10–p90 4–66.7
  • 12Claude Haiku 4.5
    40
    p10–p90 7.60–100
  • 13GLM 5.3 Flash
    42
    p10–p90 0–87.2
  • 14GPT-6 Astra
    43.5
    p10–p90 4–67.7
  • 15Grok 4.6
    49
    p10–p90 2.40–86.2
  • 16Muse Spark 1.2
    54.5
    p10–p90 20.3–87.9
  • 17Gemini 3.5 Flash
    69
    p10–p90 20.6–99
  • Actions to finish
    lower is better

    Actions executed, counting only runs that reported success.

    Over runs that reported success

    1. 1Claude Sonnet 5
      15
      p10–p90 4–69.4
    2. 2Inkling
      18
      p10–p90 8–43.0
    3. 3GPT-5.6 Luna
      19
      p10–p90 6.80–49.6
    4. 4Qwen3.8 Max
      23
      p10–p90 5.30–82
    5. 5Gemini 3.8 Flash
      26
      p10–p90 4–72.4
    6. 5GPT-5.6 Terra
      26
      p10–p90 4–82
    7. 7GLM 5.3 Flash
      28
      p10–p90 13–86
    8. 8GPT-5.6 Sol
      31
      p10–p90 5–74
    9. 9Kimi K3
      33
      p10–p90 2–84
    10. 10Claude Opus 5
      36.5
      p10–p90 6.70–56.6
    11. 11Claude Fable 5.1
      39
      p10–p90 5–66
    12. 12Claude Fable 5
      40
      p10–p90 5.40–76.4
    13. 12GPT-6 Astra
      40
      p10–p90 7–71
    14. 14Muse Spark 1.2
      41
      p10–p90 20.8–63
    15. 15Grok 4.6
      54
      p10–p90 26–88
    16. 16Claude Haiku 4.5
      60
      p10–p90 11–95.2
    17. 17Gemini 3.5 Flash
      62
      p10–p90 21–89.4
    Detour factor
    lower is better

    Geometric mean of (this agent's actions / the opponent's actions) across battles where BOTH sides completed the identical task.

    Over battles where both sides completed

    1. 1Claude Sonnet 5
      0.36x
    2. 2GPT-5.6 Luna
      0.54x
    3. 3GPT-5.6 Sol
      0.56x
    4. 4GPT-5.6 Terra
      0.69x
    5. 5Claude Fable 5
      0.75x
    6. 6Claude Opus 5
      0.87x
    7. 7Claude Fable 5.1
      0.89x
    8. 8Kimi K3
      0.89x
    9. 9Gemini 3.8 Flash
      0.97x
    10. 10Inkling
      1.04x
    11. 11Qwen3.8 Max
      1.10x
    12. 12Muse Spark 1.2
      1.19x
    13. 13GPT-6 Astra
      1.23x
    14. 14GLM 5.3 Flash
      1.38x
    15. 15Gemini 3.5 Flash
      1.84x
    16. 16Claude Haiku 4.5
      1.87x
    17. 17Grok 4.6
      2.13x

    Wasted motion

    Actions that repeated, changed nothing, or went back where it came from.

    Repeated itself
    lower is better

    Actions byte-identical to the action immediately before them.

    Over actions after the first, per run

    1. 1GPT-5.6 Sol
      0.1%
      95% 0.0%–0.6%
    2. 2Grok 4.6
      0.4%
      95% 0.2%–0.7%
    3. 3GPT-6 Astra
      0.4%
      95% 0.2%–0.6%
    4. 4GPT-5.6 Terra
      0.6%
      95% 0.3%–1.3%
    5. 5GPT-5.6 Luna
      0.6%
      95% 0.3%–1.2%
    6. 6Kimi K3
      0.7%
      95% 0.4%–1.1%
    7. 7Muse Spark 1.2
      1.0%
      95% 0.7%–1.4%
    8. 8Gemini 3.5 Flash
      1.5%
      95% 1.1%–1.9%
    9. 9Gemini 3.8 Flash
      1.5%
      95% 1.3%–1.7%
    10. 10Inkling
      1.6%
      95% 1.1%–2.4%
    11. 11Claude Sonnet 5
      2.2%
      95% 1.5%–3.2%
    12. 12Qwen3.8 Max
      2.7%
      95% 2.4%–2.9%
    13. 13Claude Fable 5
      2.7%
      95% 1.8%–3.9%
    14. 14Claude Fable 5.1
      2.8%
      95% 2.2%–3.6%
    15. 15Claude Opus 5
      2.9%
      95% 2.2%–3.8%
    16. 16GLM 5.3 Flash
      5.1%
      95% 4.4%–6.0%
    17. 17Claude Haiku 4.5
      10%
      95% 9.1%–12%
    Changed nothing
    lower is better

    Actions whose observation was identical to the previous step's, once the tab marker is stripped.

    Over actions after the first

    1. 1GPT-6 Astra
      0.3%
      95% 0.1%–0.5%
    2. 2GPT-5.6 Sol
      2.5%
      95% 1.6%–3.7%
    3. 3Claude Fable 5.1
      2.5%
      95% 1.9%–3.2%
    4. 4Gemini 3.8 Flash
      2.5%
      95% 2.3%–2.8%
    Explicitly inert
    lower is better

    Actions the browser explicitly reported as having moved nothing: scrolling past the end of a document, scrolling a panel that does not scroll.

    Over actions

    1. 1Gemini 3.8 Flash
      0%
      95% 0%–0.0%
    2. 1Qwen3.8 Max
      0%
      95% 0%–0.0%
    3. 1Claude Fable 5
      0%
      95% 0%–0.4%
    4. 1GPT-5.6 Sol
      0%
      95% 0.0%–0.4%
    5. 1Claude Fable 5.1
      0%
      95% 0%–0.2%
    Stayed put
    fingerprint

    Actions after which the page URL was unchanged from the previous step.

    Over actions after the first, on steps where both URLs were recorded

    1. Gemini 3.8 Flash
      91%
      95% 90%–91%
    2. Qwen3.8 Max
      91%
      95% 90%–91%
    3. Claude Fable 5
      95%
      95% 94%–96%
    4. GPT-5.6 Sol
      94%
      95% 92%–96%
    5. Claude Fable 5.1
      97%
      95% 96%–98%
    6. GPT-5.6 Luna
    Went back
    lower is better

    Steps landing on a URL the run had already visited and since left.

    Over steps with a recorded URL

    1. 1Claude Fable 5.1
      0.1%
      95% 0.0%–0.3%
    2. 2Claude Opus 5
      0.2%
      95% 0.1%–0.5%
    3. 3Claude Fable 5
      0.2%
      95% 0.1%–0.7%
    4. 4GPT-5.6 Sol
      0.2%
      95% 0.1%–0.8%

    Exploration

    How much of the page and the site it touched, and how it split looking from doing.

    Looking vs doing
    fingerprint

    Read-only actions (screenshot, read_page, hover) as a share of all actions that were either read-only or world-changing.

    Over read-only plus world-changing actions (terminal actions excluded)

    1. Gemini 3.8 Flash
      27%
      95% 25%–29%
    2. Qwen3.8 Max
      30%
      95% 28%–32%
    3. Claude Fable 5
      36%
      95% 27%–48%
    4. GPT-5.6 Sol
      27%
      95% 17%–38%
    5. Claude Fable 5.1
      25%
      95% 18%–34%
    6. GPT-5.6 Luna
      21%
      95% 15%–27%
    7. Claude Haiku 4.5
      63%
      95% 58%–67%
    8. Kimi K3
      20%
      95% 12%–31%
    9. Muse Spark 1.2
      24%
      95% 17%–33%
    10. Gemini 3.5 Flash
      31%
      95% 26%–36%
    11. Claude Sonnet 5
      24%
      95% 19%–30%
    12. Claude Opus 5
      42%
      95% 35%–50%
    13. GPT-5.6 Terra
      44%
      95% 32%–56%
    14. Inkling
      30%
      95% 25%–36%
    15. Grok 4.6
      18%
      95% 14%–23%
    16. GPT-6 Astra
      29%
      95% 25%–33%
    17. GLM 5.3 Flash
      25%
      95% 19%–31%
    Pages visited
    fingerprint

    Distinct page URLs observed across a run.

    Over runs with at least one recorded URL

    1. Gemini 3.8 Flash
      2
      p10–p90 1–9
    2. Qwen3.8 Max
      2
      p10–p90 1–10
    3. Claude Fable 5
      1
      p10–p90 1–4
    4. GPT-5.6 Sol
      1
      p10–p90 1–4.30
    5. Claude Fable 5.1
      1
      p10–p90 1–2
    6. GPT-5.6 Luna
      1
    Left the site
    fingerprint

    Navigations to an origin other than the first one the run observed.

    Over navigate actions

    1. Gemini 3.8 Flash
      77%
      95% 75%–79%
    2. Qwen3.8 Max
      70%
      95% 68%–72%
    3. Claude Fable 5
      41%
      95% 28%–55%
    4. GPT-5.6 Sol
      85%
      95% 72%–93%
    5. Claude Fable 5.1
      74%
      95% 64%–82%
    6. GPT-5.6 Luna
      58%
    Action variety
    fingerprint

    Shannon entropy of the run's verb distribution, normalized by log2 of the twelve verbs AVAILABLE.

    Over runs using at least two distinct verbs

    1. Gemini 3.8 Flash
      0.37
      p10–p90 0.26–0.51
    2. Qwen3.8 Max
      0.40
      p10–p90 0.25–0.54
    3. Claude Fable 5
      0.28
      p10–p90 0.28–0.52
    4. GPT-5.6 Sol
      0.30
      p10–p90 0.26–0.51
    5. Claude Fable 5.1
      0.42
      p10–p90 0.28–0.45
    6. GPT-5.6 Luna
    Action mix
    fingerprint

    Share of each of the twelve verbs across all of the agent's actions.

    Over all actions

    Gemini 3.8 FlashComposition
    navigate
    51.9%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.7%
    read page
    18.4%
    screenshot
    1.0%
    deliver
    16.2%
    done
    11.8%
    Qwen3.8 MaxComposition
    navigate
    49.9%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    1.2%
    read page
    21.3%
    screenshot
    0.6%
    deliver
    17.0%
    done
    10.1%
    Claude Fable 5Composition
    1.2%
    95% 0.9%–1.6%
  • 4GPT-5.6 Terra
    1.5%
    95% 0.9%–2.4%
  • 5Grok 4.6
    2.1%
    95% 1.6%–2.7%
  • 6GPT-5.6 Luna
    2.3%
    95% 1.6%–3.2%
  • 7Claude Fable 5
    2.6%
    95% 1.8%–3.7%
  • 8Gemini 3.5 Flash
    2.8%
    95% 2.3%–3.4%
  • 9Gemini 3.8 Flash
    2.9%
    95% 2.7%–3.1%
  • 10Qwen3.8 Max
    3.0%
    95% 2.8%–3.3%
  • 11Kimi K3
    3.0%
    95% 2.4%–3.9%
  • 12Claude Opus 5
    3.7%
    95% 2.9%–4.7%
  • 13Claude Sonnet 5
    4.2%
    95% 3.3%–5.4%
  • 14Muse Spark 1.2
    4.5%
    95% 3.8%–5.4%
  • 15Inkling
    4.7%
    95% 3.8%–5.8%
  • 16GLM 5.3 Flash
    8.5%
    95% 7.5%–9.5%
  • 17Claude Haiku 4.5
    11%
    95% 9.7%–12%
  • Worst error streak
    lower is better

    Longest run of consecutive failing actions within a run.

    Over runs that produced a trajectory

    1. 1Gemini 3.8 Flash
      0
      p10–p90 0–1
    2. 1Qwen3.8 Max
      0
      p10–p90 0–1
    3. 1Claude Fable 5
      0
      p10–p90 0–1
    4. 1GPT-5.6 Sol
      0
      p10–p90 0–1
    5. 1Claude Fable 5.1
      0
      p10–p90 0–1
    6. 1GPT-5.6 Luna
      0
      p10–p90 0–1
    7. 1Claude Sonnet 5
      0
      p10–p90 0–1
    8. 1GPT-5.6 Terra
      0
      p10–p90 0–1
    9. 1Grok 4.6
      0
      p10–p90 0–2
    10. 1GPT-6 Astra
      0
      p10–p90 0–1
    11. 11Claude Haiku 4.5
      1
      p10–p90 0.80–3
    12. 11Kimi K3
      1
      p10–p90 0–2
    13. 11Muse Spark 1.2
      1
      p10–p90 0–2.90
    14. 11Gemini 3.5 Flash
      1
      p10–p90 0–2
    15. 11Claude Opus 5
      1
      p10–p90 0–1
    16. 11Inkling
      1
      p10–p90 0–2
    17. 11GLM 5.3 Flash
      1
      p10–p90 0–3

    Resilience

    Whether the next action was different, and whether it worked.

    Recovered
    higher is better

    Failing actions whose NEXT action did not also fail.

    Over failing actions that had a next action

    1. 1GPT-5.6 Sol
      100%
      95% 61%–100%
    2. 2Claude Opus 5
      98%
      95% 91%–100%
    3. 3Gemini 3.8 Flash
      97%
      95% 95%–98%
    4. 4GPT-6 Astra
      95%
      95% 85%–99%
    5. 5Gemini 3.5 Flash
      92%
      95% 85%–96%
    6. 6Claude Fable 5
      92%
      95% 75%–98%
    7. 7GPT-5.6 Luna
      91%
      95% 76%–97%
    8. 8Claude Fable 5.1
      88%
      95% 70%–96%
    9. 9Qwen3.8 Max
      83%
      95% 80%–86%
    10. 10Inkling
      82%
      95% 72%–89%
    11. 11Claude Haiku 4.5
      80%
      95% 74%–84%
    12. 12GPT-5.6 Terra
      75%
      95% 51%–90%
    13. 13Claude Sonnet 5
      75%
      95% 62%–84%
    14. 14Kimi K3
      73%
      95% 61%–83%
    15. 15Grok 4.6
      70%
      95% 57%–81%
    16. 16Muse Spark 1.2
      62%
      95% 53%–71%
    17. 17GLM 5.3 Flash
      43%
      95% 37%–49%
    Retried the same thing
    lower is better

    Failing actions immediately followed by the byte-identical action: the same call, expecting a different answer.

    Over failing actions that had a next action

    1. 1Gemini 3.8 Flash
      0%
      95% 0%–0.7%
    2. 1GPT-5.6 Sol
      0%
      95% 0%–39%
    3. 1GPT-5.6 Luna
      0%
      95% 0%–11%
    4. 1Gemini 3.5 Flash
      0%
      95% 0%–3.7%
    lower is better

    Clicks at an (x, y) already clicked earlier in the same run.

    Over click actions

    Clicked the same pixel again figures on request: founders@coasty.ai

    Click spread
    fingerprint

    Mean pairwise distance between a run's clicks, divided by the viewport diagonal.

    Over runs with at least two clicks

    Click spread figures on request: founders@coasty.ai

    Gesture

    Drags, modifiers, and whether a path is a path or two endpoints.

    Points per drag
    higher is better

    Coordinate pairs supplied per drag action.

    Over drag actions

    Points per drag figures on request: founders@coasty.ai

    Endpoint-only drags
    lower is better

    Drags supplied with two points or fewer.

    Over drag actions

    Endpoint-only drags figures on request: founders@coasty.ai

    Modified clicks
    fingerprint

    Clicks holding Alt, Control, Meta or Shift.

    Over click actions

    Modified clicks figures on request: founders@coasty.ai

    Text entry

    How it types: bursts or characters, keys or strings.

    Characters per typing action
    fingerprint

    Characters supplied per `type` action.

    Over type actions

    Characters per typing action figures on request: founders@coasty.ai

    Keys vs strings
    fingerprint

    `press_key` actions as a share of all text-entry actions (type plus press_key).

    Over type plus press_key actions

    Keys vs strings figures on request: founders@coasty.ai

    2m 54s
    p10–p90 39s–10m 52s
  • 4GPT-5.6 Terra
    3m 07s
    p10–p90 36s–7m 09s
  • 5Qwen3.8 Max
    4m 56s
    p10–p90 1m 32s–18m 07s
  • 6GPT-5.6 Sol
    5m 09s
    p10–p90 1m 44s–8m 39s
  • 7Claude Fable 5
    6m 55s
    p10–p90 1m 09s–15m 51s
  • 8Gemini 3.8 Flash
    7m 34s
    p10–p90 1m 43s–18m 01s
  • 9Claude Fable 5.1
    7m 53s
    p10–p90 2m 11s–17m 00s
  • 10Claude Haiku 4.5
    8m 00s
    p10–p90 1m 46s–13m 49s
  • 11Claude Opus 5
    8m 50s
    p10–p90 2m 22s–18m 30s
  • 12Gemini 3.5 Flash
    10m 35s
    p10–p90 4m 29s–18m 15s
  • 13Kimi K3
    12m 05s
    p10–p90 2m 17s–20m 22s
  • 14Muse Spark 1.2
    12m 30s
    p10–p90 4m 21s–20m 30s
  • 15GPT-6 Astra
    12m 37s
    p10–p90 1m 54s–19m 05s
  • 16Grok 4.6
    12m 44s
    p10–p90 2m 30s–20m 45s
  • 17GLM 5.3 Flash
    16m 50s
    p10–p90 4m 30s–20m 38s
  • Time to first action
    lower is better

    Elapsed milliseconds when the first action resolved.

    Over runs with at least one step

    1. 1Inkling
      10s
      p10–p90 8.5s–1m 06s
    2. 2Claude Haiku 4.5
      11s
      p10–p90 9.4s–1m 03s
    3. 3GPT-5.6 Luna
      13s
      p10–p90 11s–1m 10s
    4. 4GPT-5.6 Sol
      14s
      p10–p90 11s–1m 09s
    5. 5Muse Spark 1.2
      14s
      p10–p90 11s–1m 08s
    6. 6GPT-5.6 Terra
      14s
      p10–p90 10s–1m 09s
    7. 7Claude Opus 5
      14s
      p10–p90 11s–1m 10s
    8. 8GPT-6 Astra
      15s
      p10–p90 12s–1m 09s
    9. 9Gemini 3.5 Flash
      16s
      p10–p90 12s–1m 10s
    10. 10Claude Fable 5.1
      16s
      p10–p90 13s–1m 07s
    11. 11Claude Sonnet 5
      16s
      p10–p90 11s–1m 08s
    12. 12Gemini 3.8 Flash
      16s
      p10–p90 12s–33s
    13. 13Claude Fable 5
      18s
      p10–p90 13s–1m 14s
    14. 14Kimi K3
      18s
      p10–p90 12s–1m 15s
    15. 15GLM 5.3 Flash
      19s
      p10–p90 12s–1m 30s
    16. 16Qwen3.8 Max
      19s
      p10–p90 10s–60s
    17. 17Grok 4.6
      23s
      p10–p90 13s–1m 14s
    Milliseconds per action
    lower is better

    Median gap between consecutive step timestamps within a run.

    Over runs with at least two steps

    1. 1Inkling
      2.4s
      p10–p90 1.8s–3.7s
    2. 2Claude Haiku 4.5
      3.7s
      p10–p90 3.1s–5.9s
    3. 3GPT-5.6 Luna
      3.8s
      p10–p90 1.2s–6.8s
    4. 4GPT-5.6 Terra
      4.0s
      p10–p90 1.5s–11s
    5. 5GPT-5.6 Sol
      5.1s
      p10–p90 4.0s–15s
    6. 6Qwen3.8 Max
      5.9s
      p10–p90 3.2s–14s
    7. 7Claude Sonnet 5
      6.3s
      p10–p90 4.1s–17s
    8. 8Gemini 3.5 Flash
      6.6s
      p10–p90 5.2s–10.0s
    9. 9Grok 4.6
      7.1s
      p10–p90 1.2s–20s
    10. 10Kimi K3
      7.6s
      p10–p90 5.3s–28s
    11. 11GLM 5.3 Flash
      7.8s
      p10–p90 3.5s–13s
    12. 12Gemini 3.8 Flash
      7.9s
      p10–p90 6.3s–20s
    13. 13Claude Fable 5
      8.5s
      p10–p90 6.2s–14s
    14. 14Claude Opus 5
      8.7s
      p10–p90 6.6s–14s
    15. 15Muse Spark 1.2
      10s
      p10–p90 7.1s–16s
    16. 16GPT-6 Astra
      11s
      p10–p90 5.7s–14s
    17. 17Claude Fable 5.1
      12s
      p10–p90 5.5s–19s

    Steadiness

    Whether the pace is even or spiky.

    Pace steadiness
    lower is better

    Interquartile range of run durations divided by their median.

    Over runs that produced a trajectory

    1. 1GPT-5.6 Sol
      0.47
    2. 2Gemini 3.5 Flash
      0.72
    3. 3Claude Haiku 4.5
      0.77
    4. 4GLM 5.3 Flash
      0.83
    5. 5Kimi K3
      0.84
    6. 6GPT-6 Astra
      0.87
    7. 7GPT-5.6 Luna
      0.91
    8. 8Claude Fable 5.1
      0.95
    9. 9Inkling
      0.96
    10. 10Grok 4.6
      0.98
    11. 11Muse Spark 1.2
      1.00
    12. 12Claude Opus 5
      1.03
    13. 13Gemini 3.8 Flash
      1.13
    14. 14GPT-5.6 Terra
      1.13
    15. 15Claude Fable 5
      1.46
    16. 16Claude Sonnet 5
      1.69
    17. 17Qwen3.8 Max
      1.96
    185k
    p10–p90 46k–696k
  • 4Inkling
    244k
    p10–p90 74k–863k
  • 5Qwen3.8 Max
    292k
    p10–p90 16k–1.8M
  • 6Kimi K3
    338k
    p10–p90 14k–2.4M
  • 7GPT-5.6 Sol
    368k
    p10–p90 16k–1.6M
  • 8Gemini 3.8 Flash
    430k
    p10–p90 23k–2.6M
  • 9Claude Fable 5.1
    454k
    p10–p90 22k–2.3M
  • 10Grok 4.6
    577k
    p10–p90 16k–1.8M
  • 11GLM 5.3 Flash
    632k
    p10–p90 0–2.3M
  • 12Claude Fable 5
    747k
    p10–p90 36k–1.7M
  • 13Claude Haiku 4.5
    941k
    p10–p90 86k–3.6M
  • 14Claude Opus 5
    980k
    p10–p90 41k–2.8M
  • 15Muse Spark 1.2
    1.4M
    p10–p90 317k–3.8M
  • 16GPT-6 Astra
    1.6M
    p10–p90 30k–3.1M
  • 17Gemini 3.5 Flash
    2.3M
    p10–p90 319k–4.3M
  • Output tokens
    lower is better

    Tokens generated over a run, including reasoning where the provider bills it.

    Over runs that produced a trajectory

    1. 1GPT-5.6 Luna
      4.0k
      p10–p90 2.6k–6.8k
    2. 2Inkling
      4.4k
      p10–p90 1.3k–12k
    3. 3GPT-5.6 Terra
      5.1k
      p10–p90 925–9.6k
    4. 4GPT-5.6 Sol
      5.1k
      p10–p90 694–8.1k
    5. 5Claude Sonnet 5
      5.6k
      p10–p90 1.3k–26k
    6. 6Claude Fable 5.1
      8.6k
      p10–p90 635–35k
    7. 7Kimi K3
      8.6k
      p10–p90 1.3k–21k
    8. 8Claude Fable 5
      10k
      p10–p90 758–28k
    9. 9GLM 5.3 Flash
      13k
      p10–p90 0–26k
    10. 10Gemini 3.8 Flash
      15k
      p10–p90 3.7k–42k
    11. 11Claude Opus 5
      17k
      p10–p90 2.6k–38k
    12. 12GPT-6 Astra
      18k
      p10–p90 544–31k
    13. 13Claude Haiku 4.5
      21k
      p10–p90 1.8k–37k
    14. 14Grok 4.6
      22k
      p10–p90 901–48k
    15. 15Gemini 3.5 Flash
      23k
      p10–p90 10k–44k
    16. 16Qwen3.8 Max
      33k
      p10–p90 5.1k–130k
    17. 17Muse Spark 1.2
      62k
      p10–p90 24k–98k

    Efficiency

    Spend per action, and spend per finished task: the number that decides whether you can run it.

    Output tokens per action
    lower is better

    A run's output tokens divided by its actions.

    Over runs with at least one step

    1. 1GPT-5.6 Sol
      139
      p10–p90 97.4–436
    2. 2GPT-5.6 Luna
      190
      p10–p90 92.2–641
    3. 3GPT-5.6 Terra
      204
      p10–p90 104–510
    4. 4Inkling
      240
      p10–p90 78.5–620
    5. 5Kimi K3
      242
      p10–p90 129–637
    6. 6Claude Haiku 4.5
      291
      p10–p90 111–1.1k
    7. 7GLM 5.3 Flash
      304
      p10–p90 118–456
    8. 8Gemini 3.5 Flash
      318
      p10–p90 234–1.2k
    9. 9Claude Fable 5
      403
      p10–p90 127–671
    10. 10Claude Sonnet 5
      413
      p10–p90 188–1.2k
    11. 11GPT-6 Astra
      416
      p10–p90 85.5–670
    12. 12Grok 4.6
      468
      p10–p90 190–907
    13. 13Claude Opus 5
      501
      p10–p90 168–954
    14. 14Claude Fable 5.1
      579
      p10–p90 161–981
    15. 15Gemini 3.8 Flash
      599
      p10–p90 219–1.9k
    16. 16Muse Spark 1.2
      1.0k
      p10–p90 745–1.7k
    17. 17Qwen3.8 Max
      1.2k
      p10–p90 546–3.3k
    Input tokens per action
    lower is better

    A run's input tokens divided by its actions.

    Over runs with at least one step

    1. 1GPT-5.6 Luna
      8.0k
      p10–p90 5.5k–14k
    2. 2GPT-5.6 Terra
      8.9k
      p10–p90 5.5k–15k
    3. 3Qwen3.8 Max
      11k
      p10–p90 6.2k–23k
    4. 4GPT-5.6 Sol
      12k
      p10–p90 4.3k–23k
    Tokens per finished task
    lower is better

    Total tokens across all of an agent's runs divided by the number that reported success.

    Over runs that reported success

    1. 1GPT-5.6 Luna
      263k
    2. 2GPT-5.6 Terra
      439k
    3. 3Inkling
      487k
    4. 4GPT-5.6 Sol
      767k
    5. 5Claude Sonnet 5
    98.5
    p10–p90 42.4–292
  • Claude Fable 5.1
    191
    p10–p90 81.0–381
  • GPT-5.6 Luna
    177
    p10–p90 93.9–297
  • Claude Haiku 4.5
    192
    p10–p90 96.3–528
  • Kimi K3
    273
    p10–p90 135–924
  • Muse Spark 1.2
    115
    p10–p90 93.6–133
  • Gemini 3.5 Flash
    1.5k
    p10–p90 1.3k–2.3k
  • Claude Sonnet 5
    175
    p10–p90 96.1–361
  • Claude Opus 5
    175
    p10–p90 104–341
  • GPT-5.6 Terra
    130
    p10–p90 84.9–248
  • Inkling
    197
    p10–p90 42.9–428
  • Grok 4.6
    301
    p10–p90 161–593
  • GPT-6 Astra
    328
    p10–p90 51.1–458
  • GLM 5.3 Flash
    909
    p10–p90 283–1.4k
  • Silent actions
    fingerprint

    Actions recorded with no reasoning text at all.

    Over actions

    1. Gemini 3.8 Flash
      0.2%
      95% 0.1%–0.3%
    2. Qwen3.8 Max
      15%
      95% 14%–15%
    3. Claude Fable 5
      42%
      95% 39%–45%
    4. GPT-5.6 Sol
      77%
      95% 74%–79%
    5. Claude Fable 5.1
      42%
      95% 40%–44%
    6. GPT-5.6 Luna
      68%
      95% 66%–70%
    7. Claude Haiku 4.5
      1.5%
      95% 1.1%–2.0%
    8. Kimi K3
      60%
      95% 58%–62%
    9. Muse Spark 1.2
      1.5%
      95% 1.1%–2.0%
    10. Gemini 3.5 Flash
      0.3%
      95% 0.1%–0.5%
    11. Claude Sonnet 5
      22%
      95% 20%–24%
    12. Claude Opus 5
      36%
      95% 34%–38%
    13. GPT-5.6 Terra
      73%
      95% 71%–76%
    14. Inkling
      47%
      95% 44%–49%
    15. Grok 4.6
      41%
      95% 40%–43%
    16. GPT-6 Astra
      39%
      95% 38%–41%
    17. GLM 5.3 Flash
      45%
      95% 43%–46%
    82%
    95% 64%–92%
  • Claude Fable 5.1
    63%
    95% 52%–72%
  • GPT-5.6 Luna
    96%
    95% 88%–99%
  • Claude Haiku 4.5
    63%
    95% 49%–75%
  • Kimi K3
    58%
    95% 44%–71%
  • Muse Spark 1.2
    79%
    95% 66%–88%
  • Gemini 3.5 Flash
    86%
    95% 75%–93%
  • Claude Sonnet 5
    85%
    95% 73%–92%
  • Claude Opus 5
    82%
    95% 69%–90%
  • GPT-5.6 Terra
    94%
    95% 81%–98%
  • Inkling
    79%
    95% 68%–87%
  • Grok 4.6
    38%
    95% 27%–51%
  • GPT-6 Astra
    72%
    95% 63%–80%
  • GLM 5.3 Flash
    20%
    95% 12%–31%
  • Form

    Shape of the final answer: the live confound on every human preference label.

    Answer length
    fingerprint

    Characters in the final answer handed to the judge.

    Over runs with a final answer

    1. Gemini 3.8 Flash
      1.5k
      p10–p90 422–3.3k
    2. Qwen3.8 Max
      1.4k
      p10–p90 785–2.6k
    3. Claude Fable 5
      1.3k
      p10–p90 690–1.9k
    4. GPT-5.6 Sol
      707
      p10–p90 337–959
    5. Claude Fable 5.1
      1.7k
      p10–p90 615–2.3k
    6. GPT-5.6 Luna
      592
      p10–p90 301–849
    7. Claude Haiku 4.5
      2.0k
      p10–p90 846–4.8k
    8. Kimi K3
      1.3k
      p10–p90 613–1.8k
    9. Muse Spark 1.2
      1.6k
      p10–p90 1.0k–3.3k
    10. Gemini 3.5 Flash
      1.2k
      p10–p90 589–2.0k
    11. Claude Sonnet 5
      1.5k
      p10–p90 664–2.3k
    12. Claude Opus 5
      2.0k
      p10–p90 929–2.8k
    13. GPT-5.6 Terra
      497
      p10–p90 240–789
    14. Inkling
      1.3k
      p10–p90 470–2.9k
    15. Grok 4.6
      514
      p10–p90 265–896
    16. GPT-6 Astra
      938
      p10–p90 367–1.1k
    17. GLM 5.3 Flash
      1.8k
      p10–p90 873–2.5k
    Empty answers
    lower is better

    Runs that reported success with a final answer under 20 characters.

    Over runs that reported success

    1. 1Gemini 3.8 Flash
      0%
      95% 0%–1.0%
    2. 1Qwen3.8 Max
      0%
      95% 0.0%–1.0%
    3. 1Claude Fable 5
      0%
      95% 0%–14%
    4. 1GPT-5.6 Sol
      0%
      95% 0%–15%
    5. 1Claude Fable 5.1
      0%
      95% 0%–7.0%
    Claude Sonnet 5
    60.0%
    Qwen3.8 MaxBy opponent
    Claude Fable 5
    50.0%
    Claude Fable 5.1
    55.6%
    Gemini 3.5 Flash
    48.1%
    gemini-3-7-flash
    46.6%
    Gemini 3.8 Flash
    52.3%
    GLM 5.3 Flash
    81.3%
    GPT-5.6 Luna
    39.1%
    GPT-5.6 Sol
    51.7%
    GPT-5.6 Terra
    42.5%
    GPT-6 Astra
    73.1%
    Grok 4.6
    51.7%
    Claude Haiku 4.5
    45.0%
    Inkling
    55.6%
    Kimi K3
    48.6%
    Muse Spark 1.2
    50.0%
    Claude Opus 5
    52.5%
    Claude Sonnet 5
    48.1%
    Claude Fable 5By opponent
    Claude Fable 5.1
    50.0%
    Gemini 3.5 Flash
    63.9%
    gemini-3-7-flash
    46.5%
    Gemini 3.8 Flash
    28.6%
    GLM 5.3 Flash
    50.0%
    GPT-5.6 Luna
    51.2%
    GPT-5.6 Sol
    56.8%
    GPT-5.6 Terra
    65.3%
    GPT-6 Astra
    70.0%
    Grok 4.6
    71.7%
    Claude Haiku 4.5
    61.1%
    Inkling
    56.3%
    Kimi K3
    46.2%
    Muse Spark 1.2
    72.0%
    muse-spark-1-3-contributor
    100.0%
    Claude Opus 5
    52.5%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    65.0%
    GPT-5.6 SolBy opponent
    Claude Fable 5
    43.2%
    Claude Fable 5.1
    60.7%
    Gemini 3.5 Flash
    65.6%
    gemini-3-7-flash
    43.1%
    Gemini 3.8 Flash
    51.2%
    GLM 5.3 Flash
    55.6%
    GPT-5.6 Luna
    43.9%
    GPT-5.6 Terra
    55.7%
    GPT-6 Astra
    67.9%
    Grok 4.6
    69.0%
    Claude Haiku 4.5
    58.0%
    Inkling
    63.2%
    Kimi K3
    59.9%
    Muse Spark 1.2
    62.5%
    Claude Opus 5
    52.4%
    Qwen3.8 Max
    48.3%
    Claude Sonnet 5
    56.4%
    Claude Fable 5.1By opponent
    Claude Fable 5
    50.0%
    Gemini 3.5 Flash
    45.8%
    gemini-3-7-flash
    37.5%
    Gemini 3.8 Flash
    56.9%
    GLM 5.3 Flash
    78.6%
    GPT-5.6 Luna
    53.8%
    GPT-5.6 Sol
    39.3%
    GPT-5.6 Terra
    62.5%
    GPT-6 Astra
    34.6%
    Grok 4.6
    58.3%
    Claude Haiku 4.5
    46.2%
    Inkling
    58.7%
    Kimi K3
    37.5%
    Muse Spark 1.2
    59.1%
    Claude Opus 5
    50.0%
    Qwen3.8 Max
    44.4%
    Claude Sonnet 5
    52.5%
    GPT-5.6 LunaBy opponent
    Claude Fable 5
    48.8%
    Claude Fable 5.1
    46.2%
    Gemini 3.5 Flash
    58.3%
    gemini-3-7-flash
    38.0%
    Gemini 3.8 Flash
    27.3%
    GLM 5.3 Flash
    68.8%
    GPT-5.6 Sol
    56.1%
    GPT-5.6 Terra
    50.7%
    GPT-6 Astra
    52.9%
    Grok 4.6
    46.1%
    Claude Haiku 4.5
    55.5%
    Inkling
    64.4%
    Kimi K3
    48.5%
    Muse Spark 1.2
    57.8%
    Claude Opus 5
    42.0%
    Qwen3.8 Max
    60.9%
    Claude Sonnet 5
    50.0%
    Claude Haiku 4.5By opponent
    Claude Fable 5
    38.9%
    Claude Fable 5.1
    53.8%
    Gemini 3.5 Flash
    53.7%
    gemini-3-7-flash
    50.0%
    Gemini 3.8 Flash
    33.3%
    GLM 5.3 Flash
    54.5%
    GPT-5.6 Luna
    44.5%
    GPT-5.6 Sol
    42.0%
    GPT-5.6 Terra
    39.8%
    GPT-6 Astra
    50.0%
    Grok 4.6
    56.3%
    Inkling
    52.9%
    Kimi K3
    52.9%
    Muse Spark 1.2
    39.1%
    Claude Opus 5
    47.7%
    Qwen3.8 Max
    55.0%
    Claude Sonnet 5
    39.6%
    Kimi K3By opponent
    Claude Fable 5
    53.8%
    Claude Fable 5.1
    62.5%
    Gemini 3.5 Flash
    53.2%
    gemini-3-7-flash
    50.0%
    Gemini 3.8 Flash
    54.5%
    GLM 5.3 Flash
    50.0%
    GPT-5.6 Luna
    51.5%
    GPT-5.6 Sol
    40.1%
    GPT-5.6 Terra
    53.2%
    GPT-6 Astra
    63.6%
    Grok 4.6
    54.8%
    Claude Haiku 4.5
    47.1%
    Inkling
    54.9%
    Muse Spark 1.2
    67.9%
    Claude Opus 5
    50.0%
    Qwen3.8 Max
    51.4%
    Claude Sonnet 5
    37.9%
    Muse Spark 1.2By opponent
    Claude Fable 5
    28.0%
    Claude Fable 5.1
    40.9%
    Gemini 3.5 Flash
    52.2%
    gemini-3-7-flash
    15.6%
    Gemini 3.8 Flash
    44.1%
    GLM 5.3 Flash
    54.2%
    GPT-5.6 Luna
    42.2%
    GPT-5.6 Sol
    37.5%
    GPT-5.6 Terra
    30.0%
    GPT-6 Astra
    71.1%
    Grok 4.6
    45.5%
    Claude Haiku 4.5
    60.9%
    Inkling
    31.5%
    Kimi K3
    32.1%
    Claude Opus 5
    52.2%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    32.6%
    Gemini 3.5 FlashBy opponent
    Claude Fable 5
    36.1%
    Claude Fable 5.1
    54.2%
    gemini-3-7-flash
    34.7%
    Gemini 3.8 Flash
    50.0%
    GLM 5.3 Flash
    52.9%
    GPT-5.6 Luna
    41.7%
    GPT-5.6 Sol
    34.4%
    GPT-5.6 Terra
    52.6%
    GPT-6 Astra
    50.0%
    Grok 4.6
    73.5%
    Claude Haiku 4.5
    46.3%
    Inkling
    42.6%
    Kimi K3
    46.8%
    Muse Spark 1.2
    47.8%
    Claude Opus 5
    43.6%
    Qwen3.8 Max
    51.9%
    Claude Sonnet 5
    55.6%
    Claude Sonnet 5By opponent
    Claude Fable 5
    35.0%
    Claude Fable 5.1
    47.5%
    Gemini 3.5 Flash
    44.4%
    gemini-3-7-flash
    37.7%
    Gemini 3.8 Flash
    40.0%
    GLM 5.3 Flash
    75.0%
    GPT-5.6 Luna
    50.0%
    GPT-5.6 Sol
    43.6%
    GPT-5.6 Terra
    47.2%
    GPT-6 Astra
    53.1%
    Grok 4.6
    53.6%
    Claude Haiku 4.5
    60.4%
    Inkling
    68.2%
    Kimi K3
    62.1%
    Muse Spark 1.2
    67.4%
    Claude Opus 5
    35.5%
    Qwen3.8 Max
    51.9%
    Claude Opus 5By opponent
    Claude Fable 5
    47.5%
    Claude Fable 5.1
    50.0%
    Gemini 3.5 Flash
    56.4%
    gemini-3-7-flash
    47.9%
    Gemini 3.8 Flash
    14.3%
    GLM 5.3 Flash
    78.6%
    GPT-5.6 Luna
    58.0%
    GPT-5.6 Sol
    47.6%
    GPT-5.6 Terra
    53.3%
    GPT-6 Astra
    40.9%
    Grok 4.6
    63.0%
    Claude Haiku 4.5
    52.3%
    Inkling
    50.0%
    Kimi K3
    50.0%
    Muse Spark 1.2
    47.8%
    Qwen3.8 Max
    47.5%
    Claude Sonnet 5
    64.5%
    GPT-5.6 TerraBy opponent
    Claude Fable 5
    34.7%
    Claude Fable 5.1
    37.5%
    Gemini 3.5 Flash
    47.4%
    gemini-3-7-flash
    50.0%
    Gemini 3.8 Flash
    18.2%
    GLM 5.3 Flash
    62.5%
    GPT-5.6 Luna
    49.3%
    GPT-5.6 Sol
    44.3%
    GPT-6 Astra
    69.2%
    Grok 4.6
    38.5%
    Claude Haiku 4.5
    60.2%
    Inkling
    51.6%
    Kimi K3
    46.8%
    Muse Spark 1.2
    70.0%
    muse-spark-1-3-contributor
    100.0%
    Claude Opus 5
    46.7%
    Qwen3.8 Max
    57.5%
    Claude Sonnet 5
    52.8%
    InklingBy opponent
    Claude Fable 5
    43.8%
    Claude Fable 5.1
    41.3%
    Gemini 3.5 Flash
    57.4%
    gemini-3-7-flash
    51.5%
    Gemini 3.8 Flash
    63.6%
    GLM 5.3 Flash
    50.0%
    GPT-5.6 Luna
    35.6%
    GPT-5.6 Sol
    36.8%
    GPT-5.6 Terra
    48.4%
    GPT-6 Astra
    43.3%
    Grok 4.6
    44.4%
    Claude Haiku 4.5
    47.1%
    Kimi K3
    45.1%
    Muse Spark 1.2
    68.5%
    Claude Opus 5
    50.0%
    Qwen3.8 Max
    44.4%
    Claude Sonnet 5
    31.8%
    Grok 4.6By opponent
    Claude Fable 5
    28.3%
    Claude Fable 5.1
    41.7%
    Gemini 3.5 Flash
    26.5%
    gemini-3-7-flash
    39.4%
    Gemini 3.8 Flash
    29.4%
    GLM 5.3 Flash
    60.0%
    GPT-5.6 Luna
    53.9%
    GPT-5.6 Sol
    31.0%
    GPT-5.6 Terra
    61.5%
    GPT-6 Astra
    47.2%
    Claude Haiku 4.5
    43.8%
    Inkling
    55.6%
    Kimi K3
    45.2%
    Muse Spark 1.2
    54.5%
    Claude Opus 5
    37.0%
    Qwen3.8 Max
    48.3%
    Claude Sonnet 5
    46.4%
    GPT-6 AstraBy opponent
    Claude Fable 5
    30.0%
    Claude Fable 5.1
    65.4%
    Gemini 3.5 Flash
    50.0%
    gemini-3-7-flash
    0.0%
    Gemini 3.8 Flash
    20.8%
    GLM 5.3 Flash
    66.7%
    GPT-5.6 Luna
    47.1%
    GPT-5.6 Sol
    32.1%
    GPT-5.6 Terra
    30.8%
    Grok 4.6
    52.8%
    Claude Haiku 4.5
    50.0%
    Inkling
    56.7%
    Kimi K3
    36.4%
    Muse Spark 1.2
    28.9%
    Claude Opus 5
    59.1%
    Qwen3.8 Max
    26.9%
    Claude Sonnet 5
    46.9%
    GLM 5.3 FlashBy opponent
    Claude Fable 5
    50.0%
    Claude Fable 5.1
    21.4%
    Gemini 3.5 Flash
    47.1%
    gemini-3-7-flash
    50.0%
    Gemini 3.8 Flash
    30.0%
    GPT-5.6 Luna
    31.3%
    GPT-5.6 Sol
    44.4%
    GPT-5.6 Terra
    37.5%
    GPT-6 Astra
    33.3%
    Grok 4.6
    40.0%
    Claude Haiku 4.5
    45.5%
    Inkling
    50.0%
    Kimi K3
    50.0%
    Muse Spark 1.2
    45.8%
    muse-spark-1-3-contributor
    100.0%
    Claude Opus 5
    21.4%
    Qwen3.8 Max
    18.8%
    Claude Sonnet 5
    25.0%

    Variance

    Upsets, ties, and how separable its battles are.

    Beat a higher rating
    higher is better

    Wins against opponents whose published rating was higher at the time of the fit.

    Over ranked battles against higher-rated opponents

    1. 1Gemini 3.8 Flash
      55%
      95% 38%–71%
    2. 2Qwen3.8 Max
      50%
      95% 45%–55%
    3. 3Kimi K3
      47%
      95% 40%–54%
    4. 4GPT-5.6 Luna
      46%
      95% 40%–53%
    5. 5Claude Opus 5
      46%
      95% 42%–51%
    6. 6Claude Haiku 4.5
      43%
      95% 38%–49%
    7. 7Gemini 3.5 Flash
      43%
      95% 38%–49%
    8. 7GPT-5.6 Terra
      43%
      95% 39%–48%
    9. 9Claude Fable 5.1
      43%
      95% 34%–53%
    10. 10GPT-5.6 Sol
      43%
      95% 38%–48%
    11. 11Inkling
      43%
      95% 37%–49%
    12. 12Claude Sonnet 5
      42%
      95% 37%–47%
    13. 13Grok 4.6
      40%
      95% 34%–45%
    14. 14GPT-6 Astra
      39%
      95% 31%–47%
    15. 15Muse Spark 1.2
      39%
      95% 33%–45%
    16. 16GLM 5.3 Flash
      35%
      95% 27%–43%
    Inseparable battles
    fingerprint

    Ranked battles judged a tie or both-bad.

    Over ranked battles

    1. Gemini 3.8 Flash
      23%
      95% 20%–27%
    2. Qwen3.8 Max
      26%
      95% 23%–29%
    3. Claude Fable 5
      23%
      95% 20%–25%
    4. GPT-5.6 Sol
      26%
      95% 24%–28%
    5. Claude Fable 5.1
      37%
      95% 32%–43%
    6. GPT-5.6 Luna
      23%
    unnecessary navigation
    10.9%
    repeated action
    0.0%
    misclick
    5.5%
    slow recovery
    0.0%
    Claude Fable 5Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    91.7%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    8.3%
    slow recovery
    0.0%
    GPT-5.6 SolComposition
    output contamination
    0.0%
    wrong answer
    13.3%
    incomplete
    86.7%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Claude Fable 5.1Composition
    output contamination
    0.0%
    wrong answer
    2.1%
    incomplete
    89.6%
    unnecessary navigation
    4.2%
    repeated action
    0.0%
    misclick
    4.2%
    slow recovery
    0.0%
    GPT-5.6 LunaComposition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    100.0%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Claude Haiku 4.5Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    100.0%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Kimi K3Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    95.2%
    unnecessary navigation
    4.8%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Muse Spark 1.2Composition
    output contamination
    0.0%
    wrong answer
    7.4%
    incomplete
    92.6%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Gemini 3.5 FlashComposition
    output contamination
    0.0%
    wrong answer
    8.0%
    incomplete
    88.0%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    4.0%
    slow recovery
    0.0%
    Claude Sonnet 5Composition
    output contamination
    0.0%
    wrong answer
    5.3%
    incomplete
    89.5%
    unnecessary navigation
    5.3%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Claude Opus 5Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    100.0%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    GPT-5.6 TerraComposition
    output contamination
    0.0%
    wrong answer
    13.6%
    incomplete
    68.2%
    unnecessary navigation
    9.1%
    repeated action
    0.0%
    misclick
    9.1%
    slow recovery
    0.0%
    InklingComposition
    output contamination
    0.0%
    wrong answer
    15.2%
    incomplete
    81.8%
    unnecessary navigation
    3.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Grok 4.6Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    100.0%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    GPT-6 AstraComposition
    output contamination
    0.0%
    wrong answer
    4.3%
    incomplete
    93.6%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    2.1%
    slow recovery
    0.0%
    GLM 5.3 FlashComposition
    output contamination
    0.0%
    wrong answer
    6.5%
    incomplete
    90.3%
    unnecessary navigation
    3.2%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Flagged runs
    lower is better

    Runs that collected at least one flaw annotation.

    Over runs that were annotated at all

    1. 1Muse Spark 1.2
      93%
      95% 78%–98%
    2. 2Qwen3.8 Max
      95%
      95% 86%–98%
    3. 3Claude Opus 5
      95%
      95% 76%–99%
    4. 4Kimi K3
      95%
      95% 78%–99%
    5. 5Gemini 3.5 Flash
      96%
      95% 81%–99%
    6. 6GPT-5.6 Luna
      96%
      95% 82%–99%
    7. 7Gemini 3.8 Flash
      97%
      95% 88%–99%
    8. 8GLM 5.3 Flash
      97%
      95% 84%–99%
    9. 9Inkling
      97%
      95% 85%–99%
    10. 10Claude Fable 5.1
      98%
      95% 89%–100%
    11. 11Claude Fable 5
      100%
      95% 76%–100%
    12. 11GPT-5.6 Sol
      100%
      95% 80%–100%
    13. 11Claude Haiku 4.5
      100%
      95% 85%–100%
    14. 11Claude Sonnet 5
      100%
      95% 83%–100%
    15. 11GPT-5.6 Terra
      100%
      95% 85%–100%
    16. 11Grok 4.6
      100%
      95% 86%–100%
    17. 11GPT-6 Astra
      100%
      95% 92%–100%

    Character

    The low-effort labels judges reach for.

    Character stamps
    fingerprint

    Share of the agent's stamps across the five run-level character labels.

    Over stamps on this agent's runs

    GPT-5.6 TerraComposition
    wandered
    0.0%
    recovered
    0.0%
    speedran
    0.0%
    made it up
    100.0%
    nailed it
    0.0%

    Preference

    Strength and separability of the labels its battles produced.

    Margin when it won
    higher is better

    Mean judge-reported strength (barely / clearly / not close, as 1-3) on battles this agent won.

    Over wins where the judge tapped the optional strength chip

    1. 1GPT-5.6 Luna
      2.08
    2. 2GPT-5.6 Terra
      1.92
    3. 3GPT-5.6 Sol
      1.85
    4. 4Claude Opus 5
      1.77
    5. 5Claude Fable 5
      1.72
    6. 6Gemini 3.5 Flash
      1.71
    7. 7Claude Sonnet 5
      1.58
    8. 8Claude Haiku 4.5
      1.55
  • 5Inkling
    2.6%
    95% 2.0%–3.5%
  • 6GPT-5.6 Luna
    2.8%
    95% 2.0%–3.8%
  • 7Grok 4.6
    2.8%
    95% 2.2%–3.5%
  • 8GPT-5.6 Terra
    2.9%
    95% 2.0%–4.1%
  • 9Muse Spark 1.2
    3.1%
    95% 2.5%–3.8%
  • 10Gemini 3.5 Flash
    3.3%
    95% 2.8%–3.9%
  • 11Qwen3.8 Max
    3.5%
    95% 3.3%–3.8%
  • 12Kimi K3
    4.0%
    95% 3.2%–5.0%
  • 13Claude Sonnet 5
    5.3%
    95% 4.2%–6.6%
  • 14Claude Opus 5
    5.4%
    95% 4.4%–6.6%
  • 15GLM 5.3 Flash
    8.1%
    95% 7.2%–9.1%
  • 16Claude Haiku 4.5
    11%
    95% 9.3%–12%
  • 17Claude Fable 5
    12%
    95% 9.9%–14%
  • 1GPT-5.6 Luna
    0%
    95% 0%–0.3%
  • 1Claude Haiku 4.5
    0%
    95% 0%–0.2%
  • 1Kimi K3
    0%
    95% 0%–0.2%
  • 1Muse Spark 1.2
    0%
    95% 0%–0.1%
  • 1Gemini 3.5 Flash
    0%
    95% 0%–0.1%
  • 1Claude Sonnet 5
    0%
    95% 0%–0.3%
  • 1Claude Opus 5
    0%
    95% 0%–0.2%
  • 1GPT-5.6 Terra
    0%
    95% 0%–0.4%
  • 1Inkling
    0%
    95% 0%–0.2%
  • 1Grok 4.6
    0%
    95% 0%–0.1%
  • 1GPT-6 Astra
    0%
    95% 0%–0.1%
  • 1GLM 5.3 Flash
    0%
    95% 0%–0.1%
  • 90%
    95% 89%–92%
  • Claude Haiku 4.5
    89%
    95% 88%–91%
  • Kimi K3
    97%
    95% 96%–98%
  • Muse Spark 1.2
    97%
    95% 97%–98%
  • Gemini 3.5 Flash
    95%
    95% 94%–95%
  • Claude Sonnet 5
    88%
    95% 86%–90%
  • Claude Opus 5
    95%
    95% 94%–96%
  • GPT-5.6 Terra
    96%
    95% 95%–97%
  • Inkling
    89%
    95% 87%–90%
  • Grok 4.6
    93%
    95% 92%–94%
  • GPT-6 Astra
    92%
    95% 91%–92%
  • GLM 5.3 Flash
    96%
    95% 95%–96%
  • 5Kimi K3
    0.3%
    95% 0.1%–0.6%
  • 6Muse Spark 1.2
    0.4%
    95% 0.2%–0.7%
  • 7GPT-5.6 Terra
    0.4%
    95% 0.1%–1.0%
  • 8GLM 5.3 Flash
    0.5%
    95% 0.3%–0.9%
  • 9Grok 4.6
    0.6%
    95% 0.4%–1.0%
  • 10Gemini 3.8 Flash
    0.7%
    95% 0.6%–0.8%
  • 11Gemini 3.5 Flash
    0.7%
    95% 0.5%–1.0%
  • 12GPT-6 Astra
    0.9%
    95% 0.7%–1.3%
  • 13Qwen3.8 Max
    1.4%
    95% 1.3%–1.6%
  • 14Claude Sonnet 5
    1.5%
    95% 1.0%–2.3%
  • 15Inkling
    1.6%
    95% 1.1%–2.3%
  • 16Claude Haiku 4.5
    2.0%
    95% 1.5%–2.7%
  • 17GPT-5.6 Luna
    2.1%
    95% 1.4%–3.0%
  • p10–p90 1–5
  • Claude Haiku 4.5
    2
    p10–p90 1–12
  • Kimi K3
    1
    p10–p90 1–4.20
  • Muse Spark 1.2
    1
    p10–p90 1–3
  • Gemini 3.5 Flash
    2
    p10–p90 1–8.20
  • Claude Sonnet 5
    1
    p10–p90 1–8
  • Claude Opus 5
    1
    p10–p90 1–5.20
  • GPT-5.6 Terra
    1
    p10–p90 1–3.70
  • Inkling
    1
    p10–p90 1–8
  • Grok 4.6
    2
    p10–p90 1–7.10
  • GPT-6 Astra
    2.50
    p10–p90 1–9.80
  • GLM 5.3 Flash
    1
    p10–p90 1–5
  • 95% 49%–65%
  • Claude Haiku 4.5
    80%
    95% 74%–85%
  • Kimi K3
    79%
    95% 66%–88%
  • Muse Spark 1.2
    80%
    95% 71%–87%
  • Gemini 3.5 Flash
    79%
    95% 72%–84%
  • Claude Sonnet 5
    70%
    95% 62%–77%
  • Claude Opus 5
    73%
    95% 63%–81%
  • GPT-5.6 Terra
    36%
    95% 22%–52%
  • Inkling
    78%
    95% 72%–83%
  • Grok 4.6
    81%
    95% 75%–86%
  • GPT-6 Astra
    81%
    95% 76%–85%
  • GLM 5.3 Flash
    73%
    95% 65%–80%
  • 0.28
    p10–p90 0.27–0.51
  • Claude Haiku 4.5
    0.42
    p10–p90 0.26–0.58
  • Kimi K3
    0.28
    p10–p90 0.23–0.41
  • Muse Spark 1.2
    0.39
    p10–p90 0.24–0.44
  • Gemini 3.5 Flash
    0.34
    p10–p90 0.26–0.53
  • Claude Sonnet 5
    0.36
    p10–p90 0.26–0.52
  • Claude Opus 5
    0.42
    p10–p90 0.28–0.54
  • GPT-5.6 Terra
    0.44
    p10–p90 0.28–0.54
  • Inkling
    0.42
    p10–p90 0.28–0.56
  • Grok 4.6
    0.28
    p10–p90 0.21–0.49
  • GPT-6 Astra
    0.41
    p10–p90 0.23–0.51
  • GLM 5.3 Flash
    0.28
    p10–p90 0.23–0.55
  • navigate
    37.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    14.5%
    screenshot
    6.9%
    deliver
    22.1%
    done
    19.1%
    GPT-5.6 SolComposition
    navigate
    39.8%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    14.4%
    screenshot
    0.0%
    deliver
    25.4%
    done
    20.3%
    Claude Fable 5.1Composition
    navigate
    37.0%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    11.0%
    screenshot
    1.4%
    deliver
    25.6%
    done
    25.1%
    GPT-5.6 LunaComposition
    navigate
    47.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    1.0%
    read page
    12.2%
    screenshot
    0.3%
    deliver
    20.7%
    done
    18.4%
    Claude Haiku 4.5Composition
    navigate
    29.0%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    1.0%
    read page
    18.4%
    screenshot
    31.9%
    deliver
    14.3%
    done
    5.4%
    Kimi K3Composition
    navigate
    37.8%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.8%
    read page
    7.9%
    screenshot
    1.6%
    deliver
    29.9%
    done
    22.0%
    Muse Spark 1.2Composition
    navigate
    38.6%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    5.8%
    screenshot
    6.3%
    deliver
    36.3%
    done
    13.0%
    Gemini 3.5 FlashComposition
    navigate
    48.9%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    21.6%
    screenshot
    0.2%
    deliver
    17.1%
    done
    12.2%
    Claude Sonnet 5Composition
    navigate
    47.9%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    2.6%
    read page
    12.0%
    screenshot
    3.9%
    deliver
    17.8%
    done
    15.9%
    Claude Opus 5Composition
    navigate
    36.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    10.3%
    screenshot
    16.2%
    deliver
    22.1%
    done
    15.0%
    GPT-5.6 TerraComposition
    navigate
    25.5%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    19.9%
    screenshot
    0.0%
    deliver
    31.2%
    done
    23.4%
    InklingComposition
    navigate
    49.0%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.5%
    read page
    14.6%
    screenshot
    6.7%
    deliver
    15.4%
    done
    13.8%
    Grok 4.6Composition
    navigate
    66.1%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.7%
    read page
    14.6%
    screenshot
    0.0%
    deliver
    9.2%
    done
    9.5%
    GPT-6 AstraComposition
    navigate
    50.8%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    20.4%
    screenshot
    0.3%
    deliver
    17.8%
    done
    10.6%
    GLM 5.3 FlashComposition
    navigate
    63.5%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    1.4%
    read page
    16.8%
    screenshot
    4.3%
    deliver
    8.2%
    done
    5.8%
  • 1GPT-5.6 Terra
    0%
    95% 0%–19%
  • 1Grok 4.6
    0%
    95% 0%–6.6%
  • 1GPT-6 Astra
    0%
    95% 0.0%–8.0%
  • 8Claude Opus 5
    1.6%
    95% 0.3%–8.6%
  • 9Kimi K3
    1.7%
    95% 0.3%–8.9%
  • 10Qwen3.8 Max
    2.7%
    95% 1.7%–4.2%
  • 11Claude Sonnet 5
    3.6%
    95% 1.0%–12%
  • 12Claude Fable 5.1
    4.0%
    95% 0.7%–20%
  • 13Muse Spark 1.2
    6.0%
    95% 2.9%–12%
  • 14Inkling
    7.7%
    95% 3.6%–16%
  • 15Claude Fable 5
    8.0%
    95% 2.2%–25%
  • 16Claude Haiku 4.5
    15%
    95% 11%–20%
  • 17GLM 5.3 Flash
    29%
    95% 23%–34%
  • 5Grok 4.6
    12k
    p10–p90 6.4k–23k
  • 6Inkling
    14k
    p10–p90 8.6k–20k
  • 7Kimi K3
    14k
    p10–p90 6.5k–29k
  • 8GLM 5.3 Flash
    15k
    p10–p90 6.2k–27k
  • 9Claude Fable 5
    15k
    p10–p90 8.8k–38k
  • 10Claude Sonnet 5
    15k
    p10–p90 8.4k–32k
  • 11Gemini 3.8 Flash
    16k
    p10–p90 6.7k–37k
  • 12Claude Fable 5.1
    18k
    p10–p90 9.1k–40k
  • 13Claude Haiku 4.5
    22k
    p10–p90 9.9k–38k
  • 14Claude Opus 5
    25k
    p10–p90 10k–44k
  • 15Muse Spark 1.2
    28k
    p10–p90 15k–49k
  • 16Gemini 3.5 Flash
    32k
    p10–p90 13k–49k
  • 17GPT-6 Astra
    36k
    p10–p90 6.2k–51k
  • 869k
  • 6Qwen3.8 Max
    968k
  • 7Claude Fable 5
    1.2M
  • 8Gemini 3.8 Flash
    1.4M
  • 9Claude Fable 5.1
    1.4M
  • 10Claude Opus 5
    1.5M
  • 11Kimi K3
    1.6M
  • 12Claude Haiku 4.5
    2.3M
  • 13Grok 4.6
    2.4M
  • 14Gemini 3.5 Flash
    2.9M
  • 15Muse Spark 1.2
    3.3M
  • 16GPT-6 Astra
    4.5M
  • 17GLM 5.3 Flash
    5.1M
  • 1GPT-5.6 Luna
    0%
    95% 0%–6.5%
  • 1Claude Haiku 4.5
    0%
    95% 0%–12%
  • 1Kimi K3
    0%
    95% 0%–12%
  • 1Muse Spark 1.2
    0%
    95% 0%–12%
  • 1Gemini 3.5 Flash
    0%
    95% 0.0%–7.6%
  • 1Claude Sonnet 5
    0%
    95% 0%–8.2%
  • 1Claude Opus 5
    0%
    95% 0%–9.2%
  • 1GPT-5.6 Terra
    0%
    95% 0%–11%
  • 1Inkling
    0%
    95% 0%–6.8%
  • 1Grok 4.6
    0%
    95% 0%–17%
  • 1GPT-6 Astra
    0%
    95% 0%–11%
  • 1GLM 5.3 Flash
    0%
    95% 0.0%–26%
  • 95% 20%–26%
  • Claude Haiku 4.5
    22%
    95% 19%–26%
  • Kimi K3
    23%
    95% 20%–27%
  • Muse Spark 1.2
    25%
    95% 21%–30%
  • Gemini 3.5 Flash
    25%
    95% 22%–29%
  • Claude Sonnet 5
    23%
    95% 20%–26%
  • Claude Opus 5
    24%
    95% 22%–26%
  • GPT-5.6 Terra
    26%
    95% 24%–29%
  • Inkling
    26%
    95% 22%–30%
  • Grok 4.6
    28%
    95% 25%–33%
  • GPT-6 Astra
    38%
    95% 32%–45%
  • GLM 5.3 Flash
    31%
    95% 25%–37%