Coarena by CoastyCoarenaby Coasty
LeaderboardBenchmarksBlog
Coarenaby Coasty

Real-world evals for computer-use agents. Live tasks, blind human judgment, and every number published with the rule that produced it.

Arena

  • Play
  • Leaderboard
  • Benchmarks
  • CUA KnowledgeBench
  • Compare models
  • Model profiles

Evidence

  • Dataset
  • Methodology
  • Metrics API
  • Cite us

About

  • Mission
  • Governance
  • Blog

Legal

  • Terms
  • Privacy
  • Security

© 2026 Coasty Systems, Inc.

Every claim on this site cites the file that keeps it
← Benchmarks

Computer Use Arena

Real computer tasks. Independent agent runs. Blind human comparisons.

Results

Full leaderboard ↗
1Claude Fable 5
1065; 95% interval 1037 to 1093
2Gemini 3.7 Flash
1062; 95% interval 1036 to 1088
3GPT-5.6 Sol
1041; 95% interval 1019 to 1063
4Gemini 3.8 Flash
1039; 95% interval 970 to 1108
5Claude Opus 5
1031; 95% interval 1005 to 1057
6GPT-5.6 Luna
1028; 95% interval 998 to 1058
7Kimi K3
1019; 95% interval 985 to 1053
8Claude Fable 5.1
1007; 95% interval 957 to 1057
9GPT-5.6 Terra
1005; 95% interval 971 to 1039
10Qwen3.8 Max
1001; 95% interval 962 to 1040
11Claude Sonnet 5
999; 95% interval 971 to 1027
12Gemini 3.5 Flash
985; 95% interval 951 to 1019
13Claude Haiku 4.5
982; 95% interval 954 to 1010
14Inkling
978; 95% interval 946 to 1010
15Grok 4.6
963; 95% interval 930 to 996
16GLM 5.3 Flash
937; 95% interval 871 to 1003
17Muse Spark 1.3
932; 95% interval 532 to 1332
18GPT-6 Astra
932; 95% interval 834 to 1030
19Muse Spark 1.2
925; 95% interval 881 to 969
4841380

Higher is better → · Lines show 95% intervals

Ratings measure human preference. Small differences may fall within the published 95% intervals.
Model details +
Rating and rating status for every public model
ModelRatingRating status
Claude Fable 51065 ±28stable
Gemini 3.7 Flash1062 ±26ranked
GPT-5.6 Sol1041 ±22stable
Gemini 3.8 Flash1039 ±69ranked
Claude Opus 51031 ±26stable
GPT-5.6 Luna1028 ±30ranked
Kimi K31019 ±34ranked
Claude Fable 5.11007 ±50ranked
GPT-5.6 Terra1005 ±34ranked
Qwen3.8 Max1001 ±39ranked
Claude Sonnet 5999 ±28ranked
Gemini 3.5 Flash985 ±34ranked
Claude Haiku 4.5982 ±28ranked
Inkling978 ±32ranked
Grok 4.6963 ±33ranked
GLM 5.3 Flash937 ±66ranked
Muse Spark 1.3932 ±400provisional
GPT-6 Astra932 ±98ranked
Muse Spark 1.2925 ±44ranked

— means unavailable. Provisional ratings use broad intervals. Model profiles identify the estimator.

How it works

  1. 1

    Same task

    A person submits the work. Both agents receive the same task context.

  2. 2

    Separate runs

    Each agent works in its own sandbox. Actions and outputs are recorded.

  3. 3

    Blind comparison

    A judge compares the results before model identities are revealed.

Shared inputs support comparison; models, tools, and operating conditions can still differ.
See the interface +
Computer Use Arena
Computer Use Arena task composer with an empty input.

Submit work in your own words.

Two agents receive the same task and attempt it independently.

View full size ↗ (opens image in a new tab)

Public interface snapshots. Rankings may have changed.

Methodology

How ratings work+

Ratings measure relative human preference. The board uses Bradley–Terry where supported, with online Elo and interval fallbacks otherwise. Models with fewer than 30 admitted comparisons are provisional; models with none are unrated. Both-bad judgments count as ties. Route preference is recorded separately and does not change the rating.

Which comparisons count+

A comparison needs eligible runs and an admissible vote. Identity leaks, missing or errored runs, owner-stopped lanes, and unequal step budgets exclude it. Retracted, untrusted, or measured-too-fast votes are removed. Independent votes take precedence; an admissible submitter vote may otherwise count. Judge quality can affect weight.

How to read the results+

Results depend on the submitted tasks and the records each metric includes. Completion is self-reported. A 95% interval describes uncertainty; p10–p90 describes observed spread. Missing values are not zeros. Read each metric’s definition and caveat before comparing models.

What is public+

Aggregate results and definitions are public; numerical run, battle, and sample counts are not. Full trajectories require authorised access. Owner-shared result cards are redacted and omit screenshots, action logs, and reasoning.

Metrics

57 measures across 10 families, with definitions and available results.

Leaderboard API ↗Evaluation policy ↗CUA KnowledgeBench →

Published metrics · JSON ↗

Outcome10 metrics

Did it do the job, and does it know whether it did?

Completion

How runs ended, by the agent's own report.

Completed
higher is better

Runs whose terminal `done` call reported success.

Over runs that produced a trajectory (harness errors excluded)

Interpretation: SELF-REPORTED. The agent grades itself; nothing independently verifies the task was done. Read it beside overclaim rate, which is the part that is checkable.

  1. 1GPT-5.6 Luna
    97%
    95% 86%–100%
  2. 2GPT-5.6 Terra
    85%
    95% 70%–94%
  3. 3Gemini 3.7 Flash
    85%
    95% 81%–88%
  4. 4Gemini 3.5 Flash
    82%
    95% 66%–92%
  5. 5Qwen3.8 Max
    78%
    95% 73%–82%
  6. 6GPT-5.6 Sol
    77%
    95% 71%–82%
  7. 7Gemini 3.8 Flash
    77%
    95% 70%–83%
  8. 8Muse Spark 1.2
    75%
    95% 62%–84%
  9. 9Inkling
    72%
    95% 58%–83%
  10. 10Claude Sonnet 5
    72%
    95% 56%–83%
  11. 11Claude Fable 5
    71%
    95% 59%–82%
  12. 12Claude Fable 5.1
    71%
    95% 63%–77%
  13. 13Kimi K3
    69%
    95% 54%–81%
  14. 14GPT-6 Astra
    63%
    95% 54%–71%
  15. 15Claude Haiku 4.5
    58%
    95% 43%–72%
  16. 16Claude Opus 5
    58%
    95% 39%–74%
  17. 17GLM 5.3 Flash
    50%
    95% 30%–70%
  18. 18Grok 4.6
    40%
    95% 26%–55%
  19. 19Muse Spark 1.3
    20%
    95% 3.6%–62%
Gave up
lower is better

Runs whose terminal `done` call reported failure, plus runs that stopped without calling done.

Over runs that produced a trajectory

Interpretation: Merges an honest 'I could not do this' with a run that simply stopped. `terminal_discipline` separates them.

  1. 1GPT-5.6 Luna
    0%
    95% 0%–9.4%
  2. 1Claude Opus 5
    0%
    95% 0%–13%
  3. 3Gemini 3.7 Flash
    1.3%
    95% 0.5%–2.9%
  4. 4GPT-5.6 Sol
Called impossible
fingerprint

Runs where the agent reported the objective does not exist on this environment — the control is gone, a login wall bars it.

Over runs that produced a trajectory

Interpretation: NOT a failure and not a success. Correctly refusing an impossible task is skill; every published GUI-agent corpus keeps the two apart. High is only bad if false_infeasible_rate is also high.

  1. Gemini 3.7 Flash
    2.8%
    95% 1.5%–4.9%
  2. Gemini 3.8 Flash
    1.3%
    95% 0.4%–4.6%
  3. Qwen3.8 Max
    1.6%
    95% 0.7%–3.8%
  4. GPT-5.6 Sol
    3.4%
    95% 1.7%–6.5%
  5. GPT-5.6 Luna
    0%
    95% 0%–9.4%
Ran out of clock
lower is better

Runs that hit the shared wall-clock limit before finishing.

Over runs that produced a trajectory

Interpretation: The limit is identical for both sides of every battle, but a slow provider and a slow model are indistinguishable here.

  1. 1GPT-5.6 Terra
    0%
    95% 0%–10%
  2. 1Inkling
    0%
    95% 0.0%–7.6%
  3. 1Muse Spark 1.3
    0%
    95% 0%–43%
  4. 4GPT-5.6 Luna
No trajectory
lower is better

Runs that produced no trajectory at all: an upstream outage, a rejected credential, a withdrawn model.

Over all runs, including ones with no trajectory

Interpretation: This is the PROVIDER, not the model. It is excluded from every other denominator on this page and from the ranking entirely, and is published only so an absence cannot be mistaken for a performance.

  1. 1Muse Spark 1.2
    0%
    95% 0%–6.4%
  2. 1Kimi K3
    0%
    95% 0%–8.4%
  3. 1Inkling
    0%
    95% 0.0%–7.6%
  4. 1Claude Opus 5

Self-knowledge

Where the agent's own verdict is checkable against something else — the opponent on the same task, or a judge who read the answer.

Wrongly called impossible
lower is better

Runs where the agent declared the task infeasible while its opponent COMPLETED the identical task in the same battle.

Over runs declared infeasible whose opponent produced a terminal verdict

Interpretation: The closest thing to ground truth an arena gets without an oracle: the counter-example is another agent doing it. It undercounts — if BOTH agents wrongly gave up, neither is caught — so it is a floor on the true rate, never an estimate of it.

  1. 1GPT-5.6 Terra
    0%
    95% 0%–79%
  2. 1Claude Opus 5
    0%
    95% 0%–79%
  3. 3GPT-6 Astra

Termination

How the run stopped: on purpose, or because something ran out.

Stopped on purpose
higher is better

Runs whose final action was an explicit `done` call rather than being cut off.

Over runs that produced a trajectory

Interpretation: Counts knowing you are finished, not being right about it — done(failure) and done(infeasible) both count. An agent that stops cleanly and wrongly scores well here and badly everywhere else.

  1. 1GPT-5.6 Luna
    97%
    95% 86%–100%
  2. 2Claude Sonnet 5
    90%
    95% 76%–96%
  3. 3Gemini 3.5 Flash
    88%
    95% 73%–95%

Route13 metrics

How direct was the path, and how much of it was wasted motion?

Economy

Actions spent, and how that compares to the other agent on the same task.

Actions taken
lower is better

Actions executed in a run, including the terminal one.

Over runs that produced a trajectory

Interpretation: Fewer is only better among runs that FINISHED. An agent that gives up on step three has an excellent median.

  1. 1Claude Fable 5
    4
    p10–p90 1–15.5
  2. 1Claude Sonnet 5
    4
    p10–p90 1–42.2

Recovery4 metrics

When something went wrong, what did it do next?

Failure exposure

How often actions failed at all.

Actions that failed
lower is better

Actions whose observation came back as an error.

Over actions

Interpretation: A failed action is not a failed run, and some failures are the page's fault. Read with recovery_rate: failing often and recovering always is a different agent from failing rarely and then stalling.

  1. 1GPT-5.6 Terra
    1.1%
    95% 0.5%–2.2%
  2. 2Claude Opus 5
    1.2%
    95% 0.5%–2.9%

Precision9 metrics

Does it hit what it aims at? The pixel layer, which only a computer-use arena can see.

Pointing

Where clicks land, how they cluster, and whether they repeat.

Clicks that did nothing
lower is better

Clicks whose own observation reported an error or an unchanged page.

Over click actions

Interpretation: A PROXY FOR MISSING, not a measurement of it. There is no oracle for what a click should have hit. A click that correctly hits a disabled control is counted as ineffective, and one that lands on empty page background may report nothing at all and go uncounted.

  1. 1Grok 4.6
    0%
    95% 0%–79%
Clicks at the frame edge
lower is better

Clicks within 24px of any viewport boundary.

Over click actions

Interpretation: Coordinate drift in a vision-driven agent clips against the frame first, so this catches a bias that overall click accuracy averages away. Legitimately edge-anchored controls (close buttons, scrollbars) inflate it — the comparison is across agents on the same pages, not against zero.

Tempo4 metrics

How fast, how steady, and how bad is the tail?

Latency

Wall clock, including the tails a median hides.

Run duration
lower is better

Wall-clock milliseconds from run start to terminal state.

Over runs that produced a trajectory

Interpretation: Includes provider queueing and sandbox time, neither of which is the model. Published as a distribution because the p90 is the number that decides whether it is shippable and the median is the number that hides it.

  1. 1Claude Sonnet 5
    60s
    p10–p90 19s–4m 20s
  2. 2GPT-5.6 Luna
    60s
    p10–p90 23s–2m 26s

Cost5 metrics

What does it spend, and what does a SUCCESS cost?

Volume

Tokens in and out.

Input tokens
lower is better

Tokens sent upstream over a run.

Over runs that produced a trajectory

Interpretation: Provider-reported and not tokenizer-comparable across vendors. Compare an agent to itself over time, or to others only in order of magnitude.

  1. 1Gemini 3.7 Flash
    33k
    p10–p90 11k–533k
  2. 2Claude Sonnet 5
    35k
    p10–p90 7.0k–519k

Expression2 metrics

How much does it think out loud, and per what?

Reasoning

Text the model wrote before acting.

Reasoning per action
fingerprint

Characters of model-written reasoning recorded per action.

Over actions

Interpretation: Characters, not tokens, so it is tokenizer-independent and comparable across vendors — which output_tokens is not. Only reasoning the lane SURFACES is counted; hidden reasoning is invisible here by construction.

  1. Gemini 3.7 Flash
    2.1k
    p10–p90 1.3k–3.6k
  2. Gemini 3.8 Flash
    2.0k
    p10–p90 1.5k–3.7k
  3. Qwen3.8 Max
    3.4k
    p10–p90 1.2k–9.0k

Output3 metrics

What did it hand back?

Deliverable

Files produced for the user.

Produced a file
fingerprint

Runs that wrote a downloadable artifact for the user.

Over runs that produced a trajectory

Interpretation: Most tasks do not ask for a file. Meaningful only within the categories that do.

  1. Gemini 3.7 Flash
    84%
    95% 80%–87%
  2. Gemini 3.8 Flash
    78%
    95% 71%–84%
  3. Qwen3.8 Max
    81%
    95% 76%–85%

Head to head3 metrics

Who beats whom, and is it consistent?

Dominance

The pairwise grid and what it implies.

Head-to-head grid
fingerprint

Wins, losses and ties against each individual opponent, from ranked battles only.

Over ranked battles against that opponent

Interpretation: Most cells are tiny. The grid is for spotting a matchup the aggregate rating hides, not for reading any single cell as a result.

Gemini 3.7 FlashBy opponent
Claude Fable 5
53.5%
Claude Fable 5.1
62.5%
Gemini 3.5 Flash
65.3%
Gemini 3.8 Flash
51.8%
GLM 5.3 Flash
50.0%
GPT-5.6 Luna
62.0%
GPT-5.6 Sol
57.0%
GPT-5.6 Terra
50.0%
GPT-6 Astra
100.0%
Grok 4.6
62.5%
Claude Haiku 4.5
50.0%
Inkling
51.6%
Kimi K3
50.0%
Muse Spark 1.2

Human judgment4 metrics

What did the people who watched it say went wrong?

Flaw profile

Which failure a judge attached to it, and how often.

Flaw profile
fingerprint

Share of the agent's flaw annotations falling under each of the seven flaw labels.

Over flaw annotations on this agent's runs

Interpretation: Composition, not incidence: it says which failure is characteristic of this agent, NOT how often it fails. Two agents can share a profile and differ tenfold in how much they are annotated.

Gemini 3.7 FlashComposition
output contamination
0.0%
wrong answer
14.9%
incomplete
51.1%
unnecessary navigation
21.3%
repeated action
0.0%
misclick
12.8%
slow recovery
0.0%
Gemini 3.8 FlashComposition

These measures describe recorded behavior, not guaranteed task correctness. Each definition states its limitations.

1.3%
95% 0.4%–3.7%
  • 5Claude Fable 5.1
    2.6%
    95% 1.0%–6.5%
  • 6Claude Fable 5
    3.6%
    95% 1.0%–12%
  • 7Qwen3.8 Max
    5.6%
    95% 3.5%–8.7%
  • 8Gemini 3.5 Flash
    5.9%
    95% 1.6%–19%
  • 9GPT-6 Astra
    6.6%
    95% 3.4%–12%
  • 10Claude Sonnet 5
    7.7%
    95% 2.7%–20%
  • 11Gemini 3.8 Flash
    9.7%
    95% 6.0%–15%
  • 12GPT-5.6 Terra
    12%
    95% 4.7%–27%
  • 13Kimi K3
    12%
    95% 5.2%–25%
  • 14Muse Spark 1.2
    13%
    95% 6.3%–24%
  • 15Grok 4.6
    15%
    95% 7.1%–29%
  • 16Claude Haiku 4.5
    26%
    95% 15%–40%
  • 17Inkling
    28%
    95% 17%–42%
  • 18GLM 5.3 Flash
    35%
    95% 18%–57%
  • 19Muse Spark 1.3
    80%
    95% 38%–96%
  • Claude Fable 5
    3.6%
    95% 1.0%–12%
  • Gemini 3.5 Flash
    5.9%
    95% 1.6%–19%
  • Muse Spark 1.2
    0%
    95% 0%–6.5%
  • Kimi K3
    0%
    95% 0%–8.4%
  • GPT-5.6 Terra
    2.9%
    95% 0.5%–15%
  • Inkling
    0%
    95% 0.0%–7.6%
  • Claude Fable 5.1
    2.6%
    95% 1.0%–6.5%
  • Claude Opus 5
    3.8%
    95% 0.7%–19%
  • GPT-6 Astra
    4.9%
    95% 2.3%–10%
  • Claude Haiku 4.5
    2.3%
    95% 0.4%–12%
  • Claude Sonnet 5
    15%
    95% 7.2%–30%
  • Grok 4.6
    7.5%
    95% 2.6%–20%
  • GLM 5.3 Flash
    0%
    95% 0%–16%
  • Muse Spark 1.3
    0%
    95% 0%–43%
  • 2.7%
    95% 0.5%–14%
  • 5Claude Sonnet 5
    5.1%
    95% 1.4%–17%
  • 6Gemini 3.5 Flash
    5.9%
    95% 1.6%–19%
  • 7Gemini 3.7 Flash
    12%
    95% 8.7%–15%
  • 8Gemini 3.8 Flash
    12%
    95% 8.0%–18%
  • 9Muse Spark 1.2
    13%
    95% 6.3%–24%
  • 10Claude Haiku 4.5
    14%
    95% 6.6%–27%
  • 11Qwen3.8 Max
    15%
    95% 11%–19%
  • 12GLM 5.3 Flash
    15%
    95% 5.2%–36%
  • 13GPT-5.6 Sol
    18%
    95% 14%–24%
  • 14Kimi K3
    19%
    95% 10.0%–33%
  • 15Claude Fable 5
    21%
    95% 13%–34%
  • 16Claude Fable 5.1
    24%
    95% 18%–32%
  • 17GPT-6 Astra
    25%
    95% 19%–34%
  • 18Grok 4.6
    38%
    95% 24%–53%
  • 19Claude Opus 5
    38%
    95% 22%–57%
  • 0%
    95% 0%–13%
  • 1Claude Haiku 4.5
    0%
    95% 0%–8.2%
  • 1Claude Sonnet 5
    0%
    95% 0%–9.0%
  • 7GPT-5.6 Sol
    0.4%
    95% 0.1%–2.3%
  • 8Gemini 3.7 Flash
    2.2%
    95% 1.2%–4.1%
  • 9Grok 4.6
    2.3%
    95% 0.4%–12%
  • 10Gemini 3.5 Flash
    2.8%
    95% 0.5%–14%
  • 10GPT-5.6 Terra
    2.8%
    95% 0.5%–14%
  • 12Claude Fable 5.1
    3.2%
    95% 1.4%–7.2%
  • 13GPT-6 Astra
    3.2%
    95% 1.2%–7.9%
  • 14GPT-5.6 Luna
    5.1%
    95% 1.4%–17%
  • 15Gemini 3.8 Flash
    5.4%
    95% 2.9%–9.9%
  • 16Claude Fable 5
    9.7%
    95% 4.5%–20%
  • 17Qwen3.8 Max
    14%
    95% 11%–18%
  • 18GLM 5.3 Flash
    62%
    95% 49%–74%
  • 19Muse Spark 1.3
    67%
    95% 42%–85%
  • 17%
    95% 3.0%–56%
  • 4Qwen3.8 Max
    20%
    95% 3.6%–62%
  • 5Gemini 3.7 Flash
    36%
    95% 15%–65%
  • 6GPT-5.6 Sol
    38%
    95% 14%–69%
  • 7Gemini 3.8 Flash
    50%
    95% 9.5%–91%
  • 7Gemini 3.5 Flash
    50%
    95% 9.5%–91%
  • 7Claude Fable 5.1
    50%
    95% 15%–85%
  • 7Grok 4.6
    50%
    95% 9.5%–91%
  • 11Claude Sonnet 5
    83%
    95% 44%–97%
  • 12Claude Fable 5
    100%
    95% 34%–100%
  • 12Claude Haiku 4.5
    100%
    95% 21%–100%
  • Claimed success, judged wrong
    lower is better

    Runs that reported success and then collected a human 'wrong or made up' or 'didn't finish the job' label.

    Over runs that reported success AND received at least one flaw annotation

    Interpretation: Only defined on annotated runs, which are a minority and not a random sample of them. It is evidence of overclaiming where it fires, not an unbiased estimate of how often it happens.

    1. 1Claude Fable 5
      0%
      95% 0%–43%
    2. 1Claude Sonnet 5
      0%
      95% 0%–28%
    3. 1Grok 4.6
      0%
      95% 0%–66%
    4. 4Claude Opus 5
      25%
      95% 4.6%–70%
    5. 5Qwen3.8 Max
      32%
      95% 19%–50%
    6. 6Muse Spark 1.2
      33%
      95% 12%–65%
    7. 6Kimi K3
      33%
      95% 9.7%–70%
    8. 8Gemini 3.8 Flash
      39%
      95% 22%–59%
    9. 9Gemini 3.7 Flash
      42%
      95% 29%–56%
    10. 10GPT-5.6 Terra
      43%
      95% 16%–75%
    11. 10Claude Fable 5.1
      43%
      95% 24%–63%
    12. 12GPT-6 Astra
      48%
      95% 32%–65%
    13. 13GPT-5.6 Sol
      50%
      95% 33%–67%
    14. 13Gemini 3.5 Flash
      50%
      95% 24%–76%
    15. 13Claude Haiku 4.5
      50%
      95% 19%–81%
    16. 13GLM 5.3 Flash
      50%
      95% 9.5%–91%
    17. 17GPT-5.6 Luna
      67%
      95% 35%–88%
    18. 18Inkling
      80%
      95% 38%–96%
    3GPT-5.6 Terra
    88%
    95% 73%–95%
  • 5Gemini 3.7 Flash
    88%
    95% 84%–91%
  • 6Inkling
    87%
    95% 75%–94%
  • 7GPT-5.6 Sol
    81%
    95% 76%–86%
  • 8Gemini 3.8 Flash
    81%
    95% 74%–87%
  • 9Qwen3.8 Max
    80%
    95% 75%–84%
  • 10Muse Spark 1.2
    76%
    95% 64%–86%
  • 11Claude Fable 5.1
    76%
    95% 68%–82%
  • 12Claude Fable 5
    75%
    95% 62%–84%
  • 13GPT-6 Astra
    74%
    95% 65%–81%
  • 14Kimi K3
    71%
    95% 56%–83%
  • 15Claude Opus 5
    62%
    95% 43%–78%
  • 16Claude Haiku 4.5
    60%
    95% 46%–74%
  • 17GLM 5.3 Flash
    60%
    95% 39%–78%
  • 18Grok 4.6
    55%
    95% 40%–69%
  • 19Muse Spark 1.3
    20%
    95% 3.6%–62%
  • Used every action
    lower is better

    Runs that consumed their entire step budget.

    Over runs that produced a trajectory and recorded a budget

    Interpretation: A run granted a continuation ran under a larger budget; the budget is recorded per run, and battles with unequal budgets are excluded from the ranking.

    1. 1GPT-5.6 Luna
      0%
      95% 0%–9.4%
    2. 1Claude Fable 5.1
      0%
      95% 0%–2.4%
    3. 1Claude Opus 5
      0%
      95% 0%–13%
    4. 1Muse Spark 1.3
      0%
      95% 0%–43%
    5. 5GPT-5.6 Sol
      0.4%
      95% 0.1%–2.4%
    6. 6Gemini 3.7 Flash
      0.8%
      95% 0.3%–2.2%
    7. 7GPT-6 Astra
      0.8%
      95% 0.1%–4.5%
    8. 8Muse Spark 1.2
      1.8%
      95% 0.3%–9.6%
    9. 9Claude Fable 5
      3.6%
      95% 1.0%–12%
    10. 10Inkling
      4.3%
      95% 1.2%–14%
    11. 11Grok 4.6
      5.0%
      95% 1.4%–17%
    12. 11GLM 5.3 Flash
      5.0%
      95% 0.9%–24%
    13. 13Claude Sonnet 5
      5.1%
      95% 1.4%–17%
    14. 14Qwen3.8 Max
      5.6%
      95% 3.5%–8.7%
    15. 15Gemini 3.5 Flash
      5.9%
      95% 1.6%–19%
    16. 16Gemini 3.8 Flash
      6.5%
      95% 3.5%–11%
    17. 17GPT-5.6 Terra
      8.8%
      95% 3.0%–23%
    18. 18Claude Haiku 4.5
      9.3%
      95% 3.7%–22%
    19. 19Kimi K3
      12%
      95% 5.2%–25%
    Asked for more
    fingerprint

    Runs that were granted at least one continuation by the battle's creator.

    Over runs that produced a trajectory

    Interpretation: Granted by a human, not earned. It measures what the submitter chose to extend, which is not a property of the agent.

    1. Gemini 3.7 Flash
      0%
      95% 0%–1.0%
    2. Gemini 3.8 Flash
      0%
      95% 0%–2.4%
    3. Qwen3.8 Max
      0%
      95% 0%–1.2%
    4. GPT-5.6 Sol
      0%
      95% 0%–1.6%
    5. GPT-5.6 Luna
      0%
      95% 0%–9.4%
    6. Claude Fable 5
      0%
      95% 0%–6.4%
    7. Gemini 3.5 Flash
      0%
      95% 0%–10%
    8. Muse Spark 1.2
      0%
      95% 0%–6.5%
    9. Kimi K3
      0%
      95% 0%–8.4%
    10. GPT-5.6 Terra
      0%
      95% 0%–10%
    11. Inkling
      2.1%
      95% 0.4%–11%
    12. Claude Fable 5.1
      0%
      95% 0%–2.4%
    13. Claude Opus 5
      0%
      95% 0%–13%
    14. GPT-6 Astra
      0%
      95% 0%–3.1%
    15. Claude Haiku 4.5
      0%
      95% 0%–8.2%
    16. Claude Sonnet 5
      0%
      95% 0%–9.0%
    17. Grok 4.6
      0%
      95% 0%–8.8%
    18. GLM 5.3 Flash
      0%
      95% 0%–16%
    19. Muse Spark 1.3
      0%
      95% 0%–43%
    3
    Gemini 3.7 Flash
    5
    p10–p90 2–34
  • 3Claude Fable 5.1
    5
    p10–p90 1–29.8
  • 3Claude Opus 5
    5
    p10–p90 1.50–56
  • 6GPT-5.6 Luna
    6
    p10–p90 2–23
  • 6GPT-5.6 Terra
    6
    p10–p90 2–67.2
  • 8GPT-5.6 Sol
    8
    p10–p90 2–31.4
  • 9GPT-6 Astra
    9
    p10–p90 2–81.7
  • 10Kimi K3
    10
    p10–p90 2–97.4
  • 11Gemini 3.8 Flash
    16
    p10–p90 2–84
  • 12Muse Spark 1.2
    18
    p10–p90 2.80–75.4
  • 13Qwen3.8 Max
    18.5
    p10–p90 2–78.5
  • 14Claude Haiku 4.5
    20
    p10–p90 2.20–85.0
  • 15Gemini 3.5 Flash
    22.5
    p10–p90 6.90–66.5
  • 16Inkling
    23
    p10–p90 5–58
  • 17GLM 5.3 Flash
    25
    p10–p90 4.90–74.7
  • 18Grok 4.6
    26
    p10–p90 1.80–66.4
  • 19Muse Spark 1.3
    32
    p10–p90 27–41.6
  • Actions to finish
    lower is better

    Actions executed, counting only runs that reported success.

    Over runs that reported success

    Interpretation: The honest version of 'actions taken'. Still not comparable across tasks of different difficulty — `detour_factor` is.

    1. 1Claude Fable 5
      4
      p10–p90 2–12
    2. 1Claude Sonnet 5
      4
      p10–p90 2–20.8
    3. 3Gemini 3.7 Flash
      5
      p10–p90 2–33.3
    4. 3Claude Opus 5
      5
      p10–p90 2.40–74.4
    5. 5GPT-5.6 Luna
      6
      p10–p90 2–23
    6. 5GPT-5.6 Terra
      6
      p10–p90 2–18.4
    7. 5Claude Fable 5.1
      6
      p10–p90 2–26.5
    8. 8GPT-6 Astra
      7
      p10–p90 2–38.4
    9. 9GPT-5.6 Sol
      8
      p10–p90 2–31.8
    10. 10Kimi K3
      10
      p10–p90 2–27.0
    11. 11Gemini 3.8 Flash
      13
      p10–p90 2–60.6
    12. 12GLM 5.3 Flash
      14.5
      p10–p90 3.70–41.4
    13. 13Qwen3.8 Max
      17
      p10–p90 2.80–63
    14. 14Muse Spark 1.2
      18
      p10–p90 4–41
    15. 15Gemini 3.5 Flash
      21
      p10–p90 5.70–55.3
    16. 16Inkling
      22.5
      p10–p90 5.30–40.4
    17. 17Claude Haiku 4.5
      23
      p10–p90 3.40–55
    18. 18Grok 4.6
      39
      p10–p90 8–58.5
    19. 19Muse Spark 1.3
      46
      p10–p90 46–46
    Detour factor
    lower is better

    Geometric mean of (this agent's actions / the opponent's actions) across battles where BOTH sides completed the identical task.

    Over battles where both sides completed

    Interpretation: 1.0 is parity, 2.0 is twice the actions for the same finished task. Paired, so task difficulty cancels — which is exactly what raw step counts cannot do. GEOMETRIC mean because it is a ratio: taking twice as many and taking half as many must cancel to 1.0, and an arithmetic mean would report 1.25.

    1. 1Claude Sonnet 5
      0.46x
    2. 2Claude Opus 5
      0.47x
    3. 3Gemini 3.7 Flash
      0.52x
    4. 4Claude Fable 5.1
      0.64x
    5. 5GPT-5.6 Terra
      0.66x
    6. 6Claude Fable 5
      0.67x
    7. 7Kimi K3
      0.74x
    8. 8GPT-5.6 Luna
      0.79x
    9. 9GPT-6 Astra
      0.87x
    10. 10GPT-5.6 Sol
      0.88x
    11. 11GLM 5.3 Flash
      1.59x
    12. 12Claude Haiku 4.5
      1.63x
    13. 13Gemini 3.8 Flash
      1.70x
    14. 14Muse Spark 1.3
      1.84x
    15. 15Muse Spark 1.2
      1.89x
    16. 16Qwen3.8 Max
      2.19x
    17. 17Inkling
      2.23x
    18. 18Grok 4.6
      2.49x
    19. 19Gemini 3.5 Flash
      3.12x

    Wasted motion

    Actions that repeated, changed nothing, or went back where it came from.

    Repeated itself
    lower is better

    Actions byte-identical to the action immediately before them.

    Over actions after the first, per run

    Interpretation: EXACT match only. A click one pixel away is a different attempt, possibly a deliberate correction, and counting it here would score careful re-aiming as a loop.

    1. 1GPT-5.6 Luna
      0%
      95% 0%–1.3%
    2. 2Inkling
      0.6%
      95% 0.3%–1.2%
    3. 3GPT-5.6 Sol
      0.8%
      95% 0.5%–1.2%
    4. 4Muse Spark 1.2
      0.9%
      95% 0.5%–1.5%
    5. 5GPT-6 Astra
      1.0%
      95% 0.7%–1.4%
    6. 6Gemini 3.5 Flash
      1.0%
      95% 0.6%–1.8%
    7. 7Gemini 3.8 Flash
      1.1%
      95% 0.9%–1.5%
    8. 8GPT-5.6 Terra
      1.2%
      95% 0.6%–2.4%
    9. 9Gemini 3.7 Flash
      1.6%
      95% 1.3%–2.0%
    10. 10Grok 4.6
      1.8%
      95% 1.2%–2.7%
    11. 11Qwen3.8 Max
      2.3%
      95% 2.0%–2.6%
    12. 12Claude Fable 5.1
      3.2%
      95% 2.4%–4.2%
    13. 13Kimi K3
      3.2%
      95% 2.3%–4.6%
    14. 14Claude Opus 5
      3.9%
      95% 2.4%–6.4%
    15. 15Muse Spark 1.3
      4.3%
      95% 2.1%–8.6%
    16. 16GLM 5.3 Flash
      5.8%
      95% 4.2%–7.9%
    17. 17Claude Sonnet 5
      9.0%
      95% 6.8%–12%
    18. 18Claude Haiku 4.5
      11%
      95% 9.1%–12%
    19. 19Claude Fable 5
      18%
      95% 15%–22%
    Changed nothing
    lower is better

    Actions whose observation was identical to the previous step's, once the tab marker is stripped.

    Over actions after the first

    Interpretation: Two successful, genuinely repeated actions on a static page look the same as two no-ops. It measures the RECORD not changing, which is the strongest available proxy for the page not changing.

    1. 1GPT-5.6 Luna
      0%
      95% 0%–1.3%
    2. 2Inkling
      0.7%
      95% 0.4%–1.3%
    3. 3Muse Spark 1.2
      1.5%
      95% 1.0%–2.3%
    4. 4GPT-5.6 Sol
    Explicitly inert
    lower is better

    Actions the browser explicitly reported as having moved nothing — scrolling past the end of a document, scrolling a panel that does not scroll.

    Over actions

    Interpretation: Only scrolls currently report inertness explicitly, so this is a lower bound on doing-nothing across the whole action surface.

    1. 1Gemini 3.7 Flash
      0%
      95% 0%–0.1%
    2. 1Gemini 3.8 Flash
      0%
      95% 0%–0.1%
    3. 1Qwen3.8 Max
      0%
      95% 0%–0.0%
    4. 1GPT-5.6 Sol
      0%
      95% 0%–0.1%
    5. 1GPT-5.6 Luna
    Stayed put
    fingerprint

    Actions after which the page URL was unchanged from the previous step.

    Over actions after the first, on steps where both URLs were recorded

    Interpretation: NOT a defect. Filling a form is a long stall by design. It is here to be read against the task category — a high stall rate on a navigation task means something a high stall rate on form filling does not.

    1. Gemini 3.7 Flash
      93%
      95% 92%–94%
    2. Gemini 3.8 Flash
      90%
      95% 89%–91%
    3. Qwen3.8 Max
      90%
      95% 89%–90%
    4. GPT-5.6 Sol
      84%
      95% 82%–85%
    5. GPT-5.6 Luna
      68%
      95% 63%–73%
    Went back
    lower is better

    Steps landing on a URL the run had already visited and since left.

    Over steps with a recorded URL

    Interpretation: Backtracking is sometimes correct — returning to a search results page to try the next result is good behaviour. High revisit with high repeat_action is the combination that means lost.

    1. 1Claude Fable 5
      0%
      95% 0%–0.7%
    2. 1Muse Spark 1.3
      0%
      95% 0%–2.2%
    3. 3GPT-5.6 Terra
      0.2%
      95% 0.0%–0.9%
    4. 4Gemini 3.7 Flash

    Exploration

    How much of the page and the site it touched, and how it split looking from doing.

    Looking vs doing
    fingerprint

    Read-only actions (screenshot, read_page, hover) as a share of all actions that were either read-only or world-changing.

    Over read-only plus world-changing actions (terminal actions excluded)

    Interpretation: A FINGERPRINT, not a score. High is an agent that checks before it acts; it costs actions from a fixed budget. Which is right depends entirely on how expensive a mistake is on that page.

    1. Gemini 3.7 Flash
      27%
      95% 23%–31%
    2. Gemini 3.8 Flash
      31%
      95% 28%–35%
    3. Qwen3.8 Max
      36%
      95% 33%–38%
    4. GPT-5.6 Sol
      38%
      95% 34%–41%
    5. GPT-5.6 Luna
      38%
      95% 31%–45%
    6. Claude Fable 5
      35%
      95% 26%–44%
    7. Gemini 3.5 Flash
      33%
      95% 27%–39%
    8. Muse Spark 1.2
      24%
      95% 18%–30%
    9. Kimi K3
      33%
      95% 26%–41%
    10. GPT-5.6 Terra
      35%
      95% 28%–42%
    11. Inkling
      27%
      95% 22%–31%
    12. Claude Fable 5.1
      36%
      95% 31%–41%
    13. Claude Opus 5
      22%
      95% 12%–36%
    14. GPT-6 Astra
      39%
      95% 36%–42%
    15. Claude Haiku 4.5
      55%
      95% 50%–61%
    16. Claude Sonnet 5
      27%
      95% 19%–37%
    17. Grok 4.6
      18%
      95% 15%–22%
    18. GLM 5.3 Flash
      24%
      95% 18%–32%
    19. Muse Spark 1.3
      50%
      95% 22%–78%
    Pages visited
    fingerprint

    Distinct page URLs observed across a run.

    Over runs with at least one recorded URL

    Interpretation: Driven mostly by the task. Comparable only within a category.

    1. Gemini 3.7 Flash
      1
      p10–p90 1–4
    2. Gemini 3.8 Flash
      2
      p10–p90 1–8
    3. Qwen3.8 Max
      2
      p10–p90 1–8
    4. GPT-5.6 Sol
      1
      p10–p90 1–7
    5. GPT-5.6 Luna
      1.50
      p10–p90 1–9
    6. Claude Fable 5
    Left the site
    fingerprint

    Navigations to an origin other than the first one the run observed.

    Over navigate actions

    Interpretation: Leaving is correct for a research task and usually wrong for a form-filling one. Neutral by itself; informative per category.

    1. Gemini 3.7 Flash
      54%
      95% 49%–59%
    2. Gemini 3.8 Flash
      60%
      95% 55%–64%
    3. Qwen3.8 Max
      61%
      95% 58%–64%
    4. GPT-5.6 Sol
      27%
      95% 23%–31%
    5. GPT-5.6 Luna
      12%
      95% 6.9%–19%
    Action variety
    fingerprint

    Shannon entropy of the run's verb distribution, normalized by log2 of the twelve verbs AVAILABLE.

    Over runs using at least two distinct verbs

    Interpretation: A FINGERPRINT. It combines how much of the action surface was used with how evenly — 1.0 needs all twelve verbs used equally, which nothing real does. Normalizing by verbs USED instead would measure evenness alone and score an agent that alternates two verbs the same as one that spreads across eight; on real trajectories that form sits at 1.00 for everybody.

    1. Gemini 3.7 Flash
      0.28
      p10–p90 0.28–0.50
    2. Gemini 3.8 Flash
      0.38
      p10–p90 0.28–0.54
    3. Qwen3.8 Max
      0.42
      p10–p90 0.26–0.54
    4. GPT-5.6 Sol
      0.28
      p10–p90 0.26–0.54
    5. GPT-5.6 Luna
    Action mix
    fingerprint

    Share of each of the twelve verbs across all of the agent's actions.

    Over all actions

    Interpretation: The most legible difference between two agents with identical win rates. It ranks nothing.

    Gemini 3.7 FlashComposition
    navigate
    30.2%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.5%
    read page
    10.9%
    screenshot
    0.3%
    deliver
    30.8%
    done
    27.3%
    Gemini 3.8 FlashComposition
    navigate
    46.7%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.9%
    read page
    20.6%
    screenshot
    0.9%
    deliver
    17.6%
    done
    13.4%
    Qwen3.8 MaxComposition
    3Kimi K3
    1.8%
    95% 1.1%–2.8%
  • 4Grok 4.6
    1.9%
    95% 1.3%–2.8%
  • 5GPT-5.6 Sol
    2.2%
    95% 1.7%–2.7%
  • 6Claude Fable 5
    2.2%
    95% 1.3%–3.8%
  • 7Gemini 3.7 Flash
    2.3%
    95% 1.9%–2.7%
  • 8GPT-6 Astra
    2.3%
    95% 1.8%–2.9%
  • 9Claude Fable 5.1
    2.5%
    95% 1.8%–3.3%
  • 10GPT-5.6 Luna
    2.5%
    95% 1.3%–4.8%
  • 11Claude Sonnet 5
    2.7%
    95% 1.6%–4.4%
  • 12GLM 5.3 Flash
    2.7%
    95% 1.7%–4.3%
  • 13Gemini 3.8 Flash
    3.4%
    95% 2.9%–3.9%
  • 14Qwen3.8 Max
    3.4%
    95% 3.1%–3.8%
  • 15Muse Spark 1.3
    3.6%
    95% 1.7%–7.6%
  • 16Gemini 3.5 Flash
    4.1%
    95% 3.1%–5.5%
  • 17Muse Spark 1.2
    4.6%
    95% 3.7%–5.8%
  • 18Inkling
    5.3%
    95% 4.2%–6.6%
  • 19Claude Haiku 4.5
    10%
    95% 8.7%–12%
  • Worst error streak
    lower is better

    Longest run of consecutive failing actions within a run.

    Over runs that produced a trajectory

    Interpretation: Bounded above by the step budget, so it compresses at the top for agents that spend the whole budget failing.

    1. 1Gemini 3.7 Flash
      0
      p10–p90 0–1
    2. 1Gemini 3.8 Flash
      0
      p10–p90 0–1
    3. 1Qwen3.8 Max
      0
      p10–p90 0–1
    4. 1GPT-5.6 Sol
      0
      p10–p90 0–1
    5. 1GPT-5.6 Luna
      0
      p10–p90 0–0
    6. 1Claude Fable 5
      0
      p10–p90 0–1
    7. 1Muse Spark 1.2
      0
      p10–p90 0–2
    8. 1Kimi K3
      0
      p10–p90 0–1
    9. 1GPT-5.6 Terra
      0
      p10–p90 0–0.70
    10. 1Claude Fable 5.1
      0
      p10–p90 0–1
    11. 1Claude Opus 5
      0
      p10–p90 0–1
    12. 1GPT-6 Astra
      0
      p10–p90 0–1
    13. 1Claude Sonnet 5
      0
      p10–p90 0–1
    14. 1Grok 4.6
      0
      p10–p90 0–1
    15. 1GLM 5.3 Flash
      0
      p10–p90 0–1
    16. 16Gemini 3.5 Flash
      1
      p10–p90 0–2
    17. 16Inkling
      1
      p10–p90 0–1.40
    18. 16Claude Haiku 4.5
      1
      p10–p90 0–3
    19. 16Muse Spark 1.3
      1
      p10–p90 0–1.60

    Resilience

    Whether the next action was different, and whether it worked.

    Recovered
    higher is better

    Failing actions whose NEXT action did not also fail.

    Over failing actions that had a next action

    Interpretation: An error on the last step of a run has no successor and is excluded — counting it as a failure to recover would penalise an agent for where the step budget happened to end.

    1. 1GPT-5.6 Luna
      100%
      95% 68%–100%
    2. 1Claude Fable 5
      100%
      95% 76%–100%
    3. 1Kimi K3
      100%
      95% 82%–100%
    4. 1GPT-5.6 Terra
      100%
      95% 65%–100%
    5. 1Claude Fable 5.1
      100%
      95% 92%–100%
    6. 1Claude Opus 5
      100%
      95% 51%–100%
    7. 7Gemini 3.8 Flash
      97%
      95% 93%–99%
    8. 8Gemini 3.7 Flash
      94%
      95% 88%–97%
    9. 8GLM 5.3 Flash
      94%
      95% 74%–99%
    10. 10GPT-6 Astra
      94%
      95% 85%–97%
    11. 11Qwen3.8 Max
      92%
      95% 88%–94%
    12. 12Inkling
      90%
      95% 81%–95%
    13. 13Gemini 3.5 Flash
      89%
      95% 77%–95%
    14. 14Muse Spark 1.3
      83%
      95% 44%–97%
    15. 15Muse Spark 1.2
      83%
      95% 72%–90%
    16. 16GPT-5.6 Sol
      83%
      95% 72%–90%
    17. 17Grok 4.6
      83%
      95% 63%–93%
    18. 18Claude Haiku 4.5
      77%
      95% 69%–83%
    19. 19Claude Sonnet 5
      73%
      95% 48%–89%
    Retried the same thing
    lower is better

    Failing actions immediately followed by the byte-identical action — the same call, expecting a different answer.

    Over failing actions that had a next action

    Interpretation: One identical retry can be correct: a transient timeout is worth trying twice. The signal is the rate, and the pathology is when it approaches 1.

    1. 1Gemini 3.7 Flash
      0%
      95% 0%–3.4%
    2. 1GPT-5.6 Sol
      0%
      95% 0%–5.7%
    3. 1GPT-5.6 Luna
      0%
      95% 0%–32%
    4. 1Claude Fable 5
    1. 1Grok 4.6
      0%
      95% 0%–79%
    Clicked the same pixel again
    lower is better

    Clicks at an (x, y) already clicked earlier in the same run.

    Over click actions

    Interpretation: Legitimate on a control you must click twice. Read beside click_dispersion: repeating AND concentrated is the combination that means stuck.

    1. 1Grok 4.6
      0%
      95% 0%–79%
    Click spread
    fingerprint

    Mean pairwise distance between a run's clicks, divided by the viewport diagonal.

    Over runs with at least two clicks

    Interpretation: A FINGERPRINT. Near 0 works one region; toward 1 works the whole frame. Neither is better on its own — it is the shape of the task plus the shape of the agent, and it separates two agents that look identical on every rate.

    Click spread figures on request

    Get in touch and we will send this measure over.

    Contact the team→

    founders@coasty.ai

    Gesture

    Drags, modifiers, and whether a path is a path or two endpoints.

    Points per drag
    higher is better

    Coordinate pairs supplied per drag action.

    Over drag actions

    Interpretation: The intermediate points ARE the gesture for a slider or a canvas stroke. Two points is the endpoints-only failure mode that a corpus cannot detect after the fact — which is why the arena records the whole path.

    Points per drag figures on request

    Get in touch and we will send this measure over.

    Contact the team→

    founders@coasty.ai

    Endpoint-only drags
    lower is better

    Drags supplied with two points or fewer.

    Over drag actions

    Interpretation: A straight two-point drag is correct for some gestures. High rates mean the agent does not model the path at all.

    Endpoint-only drags figures on request

    Get in touch and we will send this measure over.

    Contact the team→

    founders@coasty.ai

    Modified clicks
    fingerprint

    Clicks holding Alt, Control, Meta or Shift.

    Over click actions

    Interpretation: Rare by nature. Presence indicates an agent that reaches for multi-select and open-in-new-tab; absence indicates nothing.

    1. Grok 4.6
      0%
      95% 0%–79%

    Text entry

    How it types: bursts or characters, keys or strings.

    Characters per typing action
    fingerprint

    Characters supplied per `type` action.

    Over type actions

    Interpretation: Distinguishes field-at-a-time entry from one long paste. Neither is better; a large burst on a field with live validation behaves very differently from the same burst on a plain input.

    Characters per typing action figures on request

    Get in touch and we will send this measure over.

    Contact the team→

    founders@coasty.ai

    Keys vs strings
    fingerprint

    `press_key` actions as a share of all text-entry actions (type plus press_key).

    Over type plus press_key actions

    Interpretation: A keyboard-driven agent (Tab, Enter, Escape) versus a mouse-and-string one. Style, not skill.

    Keys vs strings figures on request

    Get in touch and we will send this measure over.

    Contact the team→

    founders@coasty.ai

    3GPT-5.6 Terra
    1m 06s
    p10–p90 26s–7m 15s
  • 4Inkling
    1m 18s
    p10–p90 26s–3m 33s
  • 5Gemini 3.7 Flash
    1m 21s
    p10–p90 31s–7m 45s
  • 6Claude Fable 5
    1m 45s
    p10–p90 34s–5m 46s
  • 7GPT-5.6 Sol
    1m 49s
    p10–p90 20s–5m 35s
  • 8GPT-6 Astra
    2m 19s
    p10–p90 26s–13m 01s
  • 9Claude Fable 5.1
    2m 24s
    p10–p90 41s–8m 58s
  • 10Kimi K3
    3m 10s
    p10–p90 45s–13m 16s
  • 11Qwen3.8 Max
    4m 13s
    p10–p90 46s–16m 50s
  • 12Gemini 3.8 Flash
    4m 25s
    p10–p90 36s–16m 08s
  • 13Claude Opus 5
    4m 30s
    p10–p90 1m 09s–8m 43s
  • 14Claude Haiku 4.5
    4m 38s
    p10–p90 30s–10m 00s
  • 15Muse Spark 1.2
    4m 50s
    p10–p90 1m 23s–19m 09s
  • 16Gemini 3.5 Flash
    5m 37s
    p10–p90 1m 24s–15m 09s
  • 17Muse Spark 1.3
    6m 57s
    p10–p90 4m 29s–13m 31s
  • 18Grok 4.6
    7m 21s
    p10–p90 32s–17m 40s
  • 19GLM 5.3 Flash
    7m 26s
    p10–p90 2m 19s–20m 10s
  • Time to first action
    lower is better

    Elapsed milliseconds when the first action resolved.

    Over runs with at least one step

    Interpretation: Model think time and the first action's execution are not separable in the record; this is their sum, plus sandbox warm-up on the first step only.

    1. 1Inkling
      10s
      p10–p90 8.7s–13s
    2. 2Claude Haiku 4.5
      11s
      p10–p90 9.0s–18s
    3. 3Muse Spark 1.2
      14s
      p10–p90 11s–50s
    4. 4GPT-5.6 Luna
      14s
      p10–p90 12s–34s
    5. 5GPT-6 Astra
      14s
      p10–p90 11s–29s
    6. 6Gemini 3.5 Flash
      14s
      p10–p90 12s–20s
    7. 7GPT-5.6 Terra
      14s
      p10–p90 11s–43s
    8. 8GPT-5.6 Sol
      14s
      p10–p90 11s–45s
    9. 9Claude Sonnet 5
      16s
      p10–p90 11s–40s
    10. 10Grok 4.6
      16s
      p10–p90 11s–28s
    11. 11Gemini 3.8 Flash
      16s
      p10–p90 12s–29s
    12. 12Claude Opus 5
      17s
      p10–p90 12s–50s
    13. 13Gemini 3.7 Flash
      17s
      p10–p90 12s–41s
    14. 14Muse Spark 1.3
      18s
      p10–p90 13s–20s
    15. 15Claude Fable 5.1
      18s
      p10–p90 14s–57s
    16. 16Kimi K3
      20s
      p10–p90 11s–1m 54s
    17. 17Claude Fable 5
      20s
      p10–p90 14s–1m 02s
    18. 18Qwen3.8 Max
      20s
      p10–p90 10s–58s
    19. 19GLM 5.3 Flash
      33s
      p10–p90 14s–1m 56s
    Milliseconds per action
    lower is better

    Median gap between consecutive step timestamps within a run.

    Over runs with at least two steps

    Interpretation: Combines model response time with action execution and waiting. Page and provider delays can affect the two runs differently.

    1. 1Inkling
      2.3s
      p10–p90 2.0s–3.7s
    2. 2Claude Haiku 4.5
      4.7s
      p10–p90 3.5s–13s
    3. 3GPT-5.6 Sol
      5.9s
      p10–p90 3.7s–18s
    4. 4GPT-5.6 Terra
      6.3s
      p10–p90 4.3s–21s
    5. 5GPT-5.6 Luna
      6.6s
      p10–p90 4.4s–17s
    6. 6Qwen3.8 Max
      6.6s
      p10–p90 3.7s–16s
    7. 7Gemini 3.5 Flash
      7.1s
      p10–p90 5.8s–11s
    8. 8GPT-6 Astra
      7.5s
      p10–p90 4.7s–18s
    9. 9Kimi K3
      7.6s
      p10–p90 4.3s–32s
    10. 10Claude Sonnet 5
      9.0s
      p10–p90 4.7s–25s
    11. 11Gemini 3.8 Flash
      9.2s
      p10–p90 6.4s–20s
    12. 12Muse Spark 1.2
      9.2s
      p10–p90 6.8s–16s
    13. 13Grok 4.6
      9.3s
      p10–p90 4.3s–17s
    14. 14Muse Spark 1.3
      10s
      p10–p90 8.3s–20s
    15. 15GLM 5.3 Flash
      10s
      p10–p90 5.7s–20s
    16. 16Claude Opus 5
      11s
      p10–p90 5.5s–34s
    17. 17Gemini 3.7 Flash
      12s
      p10–p90 6.6s–28s
    18. 18Claude Fable 5.1
      12s
      p10–p90 7.6s–34s
    19. 19Claude Fable 5
      13s
      p10–p90 6.7s–32s

    Steadiness

    Whether the pace is even or spiky.

    Pace steadiness
    lower is better

    Interquartile range of run durations divided by their median.

    Over runs that produced a trajectory

    Interpretation: Relative spread, so a 2s spread on a 4s task and on a 200s task do not read as equally steady. Limited observations can make this estimate unstable; numerical sample counts are not published.

    1. 1GPT-5.6 Terra
      0.86
    2. 2Claude Opus 5
      0.87
    3. 3GPT-5.6 Luna
      0.94
    4. 4Muse Spark 1.3
      1.00
    5. 5Grok 4.6
      1.01
    6. 6Muse Spark 1.2
      1.04
    7. 7Claude Haiku 4.5
      1.08
    8. 8Inkling
      1.22
    9. 9GLM 5.3 Flash
      1.26
    10. 10Gemini 3.5 Flash
      1.29
    11. 11Kimi K3
      1.48
    12. 12Qwen3.8 Max
      1.57
    13. 13GPT-5.6 Sol
      1.64
    14. 14Claude Fable 5.1
      1.66
    15. 15Claude Sonnet 5
      1.91
    16. 16Claude Fable 5
      2.02
    17. 17Gemini 3.8 Flash
      2.42
    18. 18Gemini 3.7 Flash
      2.74
    19. 19GPT-6 Astra
      2.93
    3
    GPT-5.6 Luna
    36k
    p10–p90 10k–184k
  • 4Claude Fable 5
    37k
    p10–p90 7.2k–251k
  • 5Claude Opus 5
    43k
    p10–p90 12k–690k
  • 6GPT-5.6 Sol
    44k
    p10–p90 9.7k–306k
  • 7Claude Fable 5.1
    52k
    p10–p90 6.7k–471k
  • 8GPT-5.6 Terra
    54k
    p10–p90 13k–701k
  • 9GPT-6 Astra
    64k
    p10–p90 11k–940k
  • 10Kimi K3
    86k
    p10–p90 11k–1.3M
  • 11Qwen3.8 Max
    174k
    p10–p90 12k–1.3M
  • 12Inkling
    194k
    p10–p90 37k–702k
  • 13Gemini 3.8 Flash
    195k
    p10–p90 12k–2.2M
  • 14GLM 5.3 Flash
    248k
    p10–p90 28k–1.2M
  • 15Grok 4.6
    272k
    p10–p90 11k–901k
  • 16Muse Spark 1.2
    303k
    p10–p90 28k–1.3M
  • 17Claude Haiku 4.5
    338k
    p10–p90 16k–1.5M
  • 18Muse Spark 1.3
    378k
    p10–p90 274k–509k
  • 19Gemini 3.5 Flash
    396k
    p10–p90 54k–2.5M
  • Output tokens
    lower is better

    Tokens generated over a run, including reasoning where the provider bills it.

    Over runs that produced a trajectory

    Interpretation: Whether hidden reasoning is included differs by provider, so cross-vendor comparison overstates the ones that report it.

    1. 1GPT-5.6 Sol
      2.1k
      p10–p90 145–6.1k
    2. 2Inkling
      2.1k
      p10–p90 682–14k
    3. 3GPT-6 Astra
      2.2k
      p10–p90 413–8.4k
    4. 4GPT-5.6 Luna
      2.6k
      p10–p90 859–4.8k
    5. 5Claude Fable 5.1
      2.8k
      p10–p90 256–12k
    6. 6GPT-5.6 Terra
      2.9k
      p10–p90 919–7.8k
    7. 7Claude Fable 5
      3.2k
      p10–p90 418–8.4k
    8. 8Claude Sonnet 5
      3.4k
      p10–p90 647–10k
    9. 9Kimi K3
      3.6k
      p10–p90 447–12k
    10. 10Gemini 3.7 Flash
      3.8k
      p10–p90 577–17k
    11. 11Claude Opus 5
      4.2k
      p10–p90 483–15k
    12. 12GLM 5.3 Flash
      7.2k
      p10–p90 2.3k–24k
    13. 13Grok 4.6
      8.3k
      p10–p90 419–29k
    14. 14Gemini 3.8 Flash
      11k
      p10–p90 1.5k–43k
    15. 15Claude Haiku 4.5
      12k
      p10–p90 610–29k
    16. 16Gemini 3.5 Flash
      21k
      p10–p90 4.3k–53k
    17. 17Qwen3.8 Max
      21k
      p10–p90 2.1k–125k
    18. 18Muse Spark 1.2
      25k
      p10–p90 3.7k–92k
    19. 19Muse Spark 1.3
      28k
      p10–p90 14k–43k

    Efficiency

    Spend per action, and spend per finished task — the number that decides whether you can run it.

    Output tokens per action
    lower is better

    A run's output tokens divided by its actions.

    Over runs with at least one step

    Interpretation: How much the model says per thing it does. Not quality — a terse agent that fails is still terse.

    1. 1Inkling
      102
      p10–p90 38.0–634
    2. 2GPT-5.6 Sol
      201
      p10–p90 62–1.2k
    3. 3GPT-6 Astra
      229
      p10–p90 53.9–664
    4. 4GPT-5.6 Terra
      249
      p10–p90 123–1.3k
    5. 5Kimi K3
      284
      p10–p90 123–1.1k
    6. 6Claude Haiku 4.5
      341
      p10–p90 95.2–906
    7. 7GPT-5.6 Luna
      356
      p10–p90 132–1.3k
    8. 8Claude Fable 5.1
      375
      p10–p90 131–1.4k
    9. 9GLM 5.3 Flash
      382
      p10–p90 204–608
    10. 10Grok 4.6
      390
      p10–p90 136–658
    11. 11Claude Sonnet 5
      510
      p10–p90 165–1.9k
    12. 12Claude Fable 5
      544
      p10–p90 159–1.7k
    13. 13Gemini 3.8 Flash
      682
      p10–p90 233–2.2k
    14. 14Muse Spark 1.3
      718
      p10–p90 493–1.4k
    15. 15Gemini 3.7 Flash
      728
      p10–p90 151–2.3k
    16. 16Gemini 3.5 Flash
      753
      p10–p90 307–1.7k
    17. 17Claude Opus 5
      787
      p10–p90 118–1.9k
    18. 18Muse Spark 1.2
      1.2k
      p10–p90 767–2.7k
    19. 19Qwen3.8 Max
      1.3k
      p10–p90 469–3.8k
    Input tokens per action
    lower is better

    A run's input tokens divided by its actions.

    Over runs with at least one step

    Interpretation: How fast the context balloons. Sensitive to how much of each tool result a lane echoes forward, which is capped uniformly EXCEPT on the Gemini lane, where the upstream contract forbids truncation. That asymmetry is declared, not hidden.

    1. 1GPT-5.6 Luna
      6.6k
      p10–p90 5.3k–9.6k
    2. 2GPT-5.6 Sol
      6.8k
      p10–p90 4.9k–11k
    3. 3Gemini 3.7 Flash
      7.3k
      p10–p90 5.2k–17k
    4. 4GPT-6 Astra
    Tokens per finished task
    lower is better

    Total tokens across all of an agent's runs divided by the number that reported success.

    Over runs that reported success

    Interpretation: THE COMMERCIAL NUMBER: the cost of a success, not the cost of an attempt, so failed runs are correctly charged to the successes they did not produce. Undefined with zero completions, where it is null rather than infinite.

    1. 1GPT-5.6 Luna
      79k
    2. 2GPT-5.6 Sol
      154k
    3. 3GPT-5.6 Terra
      227k
    4. 4Claude Fable 5
      235k
  • GPT-5.6 Sol
    207
    p10–p90 64.7–365
  • GPT-5.6 Luna
    234
    p10–p90 99.4–743
  • Claude Fable 5
    184
    p10–p90 73–440
  • Gemini 3.5 Flash
    1.9k
    p10–p90 1.6k–2.4k
  • Muse Spark 1.2
    106
    p10–p90 82.4–122
  • Kimi K3
    341
    p10–p90 162–855
  • GPT-5.6 Terra
    210
    p10–p90 6.76–426
  • Inkling
    69.2
    p10–p90 11.4–248
  • Claude Fable 5.1
    110
    p10–p90 38.8–246
  • Claude Opus 5
    222
    p10–p90 99.5–771
  • GPT-6 Astra
    159
    p10–p90 0–420
  • Claude Haiku 4.5
    211
    p10–p90 116–624
  • Claude Sonnet 5
    147
    p10–p90 66.0–278
  • Grok 4.6
    244
    p10–p90 150–356
  • GLM 5.3 Flash
    1.1k
    p10–p90 501–1.7k
  • Muse Spark 1.3
    76.1
    p10–p90 68.5–84.6
  • Silent actions
    fingerprint

    Actions recorded with no reasoning text at all.

    Over actions

    Interpretation: Partly a lane property: providers differ in whether they return reasoning alongside a tool call. Compare within a vendor before comparing across.

    1. Gemini 3.7 Flash
      0.5%
      95% 0.3%–0.7%
    2. Gemini 3.8 Flash
      0.1%
      95% 0.1%–0.3%
    3. Qwen3.8 Max
      11%
      95% 10%–11%
    4. GPT-5.6 Sol
      67%
      95% 65%–68%
    5. GPT-5.6 Luna
      59%
      95% 54%–64%
    6. Claude Fable 5
      49%
      95% 45%–53%
    7. Gemini 3.5 Flash
      0.5%
      95% 0.3%–1.2%
    8. Muse Spark 1.2
      3.6%
      95% 2.8%–4.7%
    9. Kimi K3
      52%
      95% 49%–55%
    10. GPT-5.6 Terra
      56%
      95% 52%–60%
    11. Inkling
      70%
      95% 68%–73%
    12. Claude Fable 5.1
      43%
      95% 41%–46%
    13. Claude Opus 5
      43%
      95% 38%–48%
    14. GPT-6 Astra
      61%
      95% 59%–62%
    15. Claude Haiku 4.5
      0%
      95% 0%–0.3%
    16. Claude Sonnet 5
      21%
      95% 18%–24%
    17. Grok 4.6
      0.8%
      95% 0.4%–1.5%
    18. GLM 5.3 Flash
      34%
      95% 31%–38%
    19. Muse Spark 1.3
      4.2%
      95% 2.0%–8.4%
    GPT-5.6 Sol
    77%
    95% 71%–82%
  • GPT-5.6 Luna
    97%
    95% 86%–100%
  • Claude Fable 5
    71%
    95% 59%–82%
  • Gemini 3.5 Flash
    88%
    95% 73%–95%
  • Muse Spark 1.2
    85%
    95% 74%–92%
  • Kimi K3
    67%
    95% 52%–79%
  • GPT-5.6 Terra
    82%
    95% 66%–92%
  • Inkling
    68%
    95% 54%–80%
  • Claude Fable 5.1
    65%
    95% 58%–72%
  • Claude Opus 5
    54%
    95% 35%–71%
  • GPT-6 Astra
    75%
    95% 67%–82%
  • Claude Haiku 4.5
    56%
    95% 41%–70%
  • Claude Sonnet 5
    69%
    95% 54%–81%
  • Grok 4.6
    43%
    95% 29%–58%
  • GLM 5.3 Flash
    50%
    95% 30%–70%
  • Muse Spark 1.3
    80%
    95% 38%–96%
  • Form

    Shape of the final answer — the live confound on every human preference label.

    Answer length
    fingerprint

    Characters in the final answer handed to the judge.

    Over runs with a final answer

    Interpretation: Answer length can influence human preference independently of correctness. This is a descriptive measure of the final answer, not evidence that a longer answer is better. Published ratings do not adjust for answer length.

    1. Gemini 3.7 Flash
      878
      p10–p90 246–2.2k
    2. Gemini 3.8 Flash
      1.7k
      p10–p90 332–3.8k
    3. Qwen3.8 Max
      1.3k
      p10–p90 602–2.3k
    4. GPT-5.6 Sol
      383
      p10–p90 180–567
    5. GPT-5.6 Luna
      318
      p10–p90 236–491
    6. Claude Fable 5
      699
      p10–p90 398–886
    7. Gemini 3.5 Flash
      1.1k
      p10–p90 305–2.1k
    8. Muse Spark 1.2
      997
      p10–p90 523–2.6k
    9. Kimi K3
      656
      p10–p90 382–1.0k
    10. GPT-5.6 Terra
      284
      p10–p90 208–461
    11. Inkling
      512
      p10–p90 303–1.1k
    12. Claude Fable 5.1
      739
      p10–p90 390–1.2k
    13. Claude Opus 5
      843
      p10–p90 473–1.2k
    14. GPT-6 Astra
      598
      p10–p90 190–1.2k
    15. Claude Haiku 4.5
      1.4k
      p10–p90 468–2.5k
    16. Claude Sonnet 5
      829
      p10–p90 423–1.3k
    17. Grok 4.6
      349
      p10–p90 220–545
    18. GLM 5.3 Flash
      1.1k
      p10–p90 614–1.9k
    19. Muse Spark 1.3
      634
      p10–p90 255–901
    Empty answers
    lower is better

    Runs that reported success with a final answer under 20 characters.

    Over runs that reported success

    Interpretation: A short answer can be the correct one ('42'). It is a flag to look at the trajectory, not a verdict.

    1. 1Gemini 3.7 Flash
      0%
      95% 0%–1.1%
    2. 1Gemini 3.8 Flash
      0%
      95% 0%–3.1%
    3. 1Qwen3.8 Max
      0%
      95% 0%–1.6%
    4. 1GPT-5.6 Sol
      0%
      95% 0%–2.1%
    5. 1GPT-5.6 Luna
    84.4%
    Claude Opus 5
    52.1%
    Qwen3.8 Max
    52.1%
    Claude Sonnet 5
    62.3%
    Gemini 3.8 FlashBy opponent
    Claude Fable 5
    80.0%
    Claude Fable 5.1
    42.3%
    Gemini 3.5 Flash
    42.9%
    Gemini 3.7 Flash
    48.2%
    GLM 5.3 Flash
    100.0%
    GPT-5.6 Luna
    87.5%
    GPT-5.6 Sol
    48.2%
    GPT-5.6 Terra
    75.0%
    GPT-6 Astra
    85.7%
    Grok 4.6
    68.2%
    Claude Haiku 4.5
    66.7%
    Inkling
    28.6%
    Kimi K3
    41.7%
    Muse Spark 1.2
    50.0%
    Claude Opus 5
    83.3%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    66.7%
    Qwen3.8 MaxBy opponent
    Claude Fable 5
    45.2%
    Claude Fable 5.1
    57.1%
    Gemini 3.5 Flash
    50.0%
    Gemini 3.7 Flash
    47.9%
    Gemini 3.8 Flash
    50.0%
    GLM 5.3 Flash
    83.3%
    GPT-5.6 Luna
    36.2%
    GPT-5.6 Sol
    50.0%
    GPT-5.6 Terra
    35.3%
    GPT-6 Astra
    80.0%
    Grok 4.6
    50.0%
    Claude Haiku 4.5
    47.2%
    Inkling
    54.8%
    Kimi K3
    52.9%
    Muse Spark 1.2
    50.0%
    Claude Opus 5
    50.0%
    Claude Sonnet 5
    47.7%
    GPT-5.6 SolBy opponent
    Claude Fable 5
    43.8%
    Claude Fable 5.1
    59.5%
    Gemini 3.5 Flash
    67.4%
    Gemini 3.7 Flash
    43.0%
    Gemini 3.8 Flash
    51.8%
    GLM 5.3 Flash
    56.3%
    GPT-5.6 Luna
    43.2%
    GPT-5.6 Terra
    55.7%
    GPT-6 Astra
    66.7%
    Grok 4.6
    72.2%
    Claude Haiku 4.5
    58.1%
    Inkling
    61.7%
    Kimi K3
    59.4%
    Muse Spark 1.2
    65.0%
    Claude Opus 5
    52.4%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    56.7%
    GPT-5.6 LunaBy opponent
    Claude Fable 5
    47.6%
    Claude Fable 5.1
    35.7%
    Gemini 3.5 Flash
    59.5%
    Gemini 3.7 Flash
    38.0%
    Gemini 3.8 Flash
    12.5%
    GLM 5.3 Flash
    56.3%
    GPT-5.6 Sol
    56.8%
    GPT-5.6 Terra
    50.0%
    GPT-6 Astra
    55.6%
    Grok 4.6
    46.8%
    Claude Haiku 4.5
    55.1%
    Inkling
    66.7%
    Kimi K3
    50.0%
    Muse Spark 1.2
    64.8%
    Claude Opus 5
    42.1%
    Qwen3.8 Max
    63.8%
    Claude Sonnet 5
    49.0%
    Claude Fable 5By opponent
    Claude Fable 5.1
    64.3%
    Gemini 3.5 Flash
    64.3%
    Gemini 3.7 Flash
    46.5%
    Gemini 3.8 Flash
    20.0%
    GLM 5.3 Flash
    45.8%
    GPT-5.6 Luna
    52.4%
    GPT-5.6 Sol
    56.3%
    GPT-5.6 Terra
    66.0%
    GPT-6 Astra
    50.0%
    Grok 4.6
    70.8%
    Claude Haiku 4.5
    61.1%
    Inkling
    56.9%
    Kimi K3
    48.1%
    Muse Spark 1.2
    76.1%
    Muse Spark 1.3
    100.0%
    Claude Opus 5
    52.3%
    Qwen3.8 Max
    54.8%
    Claude Sonnet 5
    65.0%
    Gemini 3.5 FlashBy opponent
    Claude Fable 5
    35.7%
    Claude Fable 5.1
    60.0%
    Gemini 3.7 Flash
    34.7%
    Gemini 3.8 Flash
    57.1%
    GLM 5.3 Flash
    54.5%
    GPT-5.6 Luna
    40.5%
    GPT-5.6 Sol
    32.6%
    GPT-5.6 Terra
    55.6%
    GPT-6 Astra
    62.5%
    Grok 4.6
    73.4%
    Claude Haiku 4.5
    45.3%
    Inkling
    40.3%
    Kimi K3
    44.4%
    Muse Spark 1.2
    55.3%
    Claude Opus 5
    42.9%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    56.7%
    Muse Spark 1.2By opponent
    Claude Fable 5
    23.9%
    Claude Fable 5.1
    33.3%
    Gemini 3.5 Flash
    44.7%
    Gemini 3.7 Flash
    15.6%
    Gemini 3.8 Flash
    50.0%
    GLM 5.3 Flash
    44.4%
    GPT-5.6 Luna
    35.2%
    GPT-5.6 Sol
    35.0%
    GPT-5.6 Terra
    34.4%
    GPT-6 Astra
    77.3%
    Grok 4.6
    40.6%
    Claude Haiku 4.5
    65.8%
    Inkling
    34.8%
    Kimi K3
    18.4%
    Claude Opus 5
    50.0%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    25.0%
    Kimi K3By opponent
    Claude Fable 5
    51.9%
    Claude Fable 5.1
    54.5%
    Gemini 3.5 Flash
    55.6%
    Gemini 3.7 Flash
    50.0%
    Gemini 3.8 Flash
    58.3%
    GLM 5.3 Flash
    56.3%
    GPT-5.6 Luna
    50.0%
    GPT-5.6 Sol
    40.6%
    GPT-5.6 Terra
    53.7%
    GPT-6 Astra
    70.0%
    Grok 4.6
    53.6%
    Claude Haiku 4.5
    45.2%
    Inkling
    55.9%
    Muse Spark 1.2
    81.6%
    Claude Opus 5
    47.4%
    Qwen3.8 Max
    47.1%
    Claude Sonnet 5
    36.0%
    GPT-5.6 TerraBy opponent
    Claude Fable 5
    34.0%
    Claude Fable 5.1
    33.3%
    Gemini 3.5 Flash
    44.4%
    Gemini 3.7 Flash
    50.0%
    Gemini 3.8 Flash
    25.0%
    GLM 5.3 Flash
    78.6%
    GPT-5.6 Luna
    50.0%
    GPT-5.6 Sol
    44.3%
    GPT-6 Astra
    100.0%
    Grok 4.6
    39.6%
    Claude Haiku 4.5
    63.0%
    Inkling
    50.0%
    Kimi K3
    46.3%
    Muse Spark 1.2
    65.6%
    Muse Spark 1.3
    100.0%
    Claude Opus 5
    46.6%
    Qwen3.8 Max
    64.7%
    Claude Sonnet 5
    51.9%
    InklingBy opponent
    Claude Fable 5
    43.1%
    Claude Fable 5.1
    50.0%
    Gemini 3.5 Flash
    59.7%
    Gemini 3.7 Flash
    48.4%
    Gemini 3.8 Flash
    71.4%
    GLM 5.3 Flash
    43.8%
    GPT-5.6 Luna
    33.3%
    GPT-5.6 Sol
    38.3%
    GPT-5.6 Terra
    50.0%
    GPT-6 Astra
    33.3%
    Grok 4.6
    50.0%
    Claude Haiku 4.5
    44.8%
    Kimi K3
    44.1%
    Muse Spark 1.2
    65.2%
    Claude Opus 5
    53.3%
    Qwen3.8 Max
    45.2%
    Claude Sonnet 5
    28.3%
    Claude Fable 5.1By opponent
    Claude Fable 5
    35.7%
    Gemini 3.5 Flash
    40.0%
    Gemini 3.7 Flash
    37.5%
    Gemini 3.8 Flash
    57.7%
    GLM 5.3 Flash
    66.7%
    GPT-5.6 Luna
    64.3%
    GPT-5.6 Sol
    40.5%
    GPT-5.6 Terra
    66.7%
    GPT-6 Astra
    25.0%
    Grok 4.6
    55.0%
    Claude Haiku 4.5
    50.0%
    Inkling
    50.0%
    Kimi K3
    45.5%
    Muse Spark 1.2
    66.7%
    Claude Opus 5
    70.0%
    Qwen3.8 Max
    42.9%
    Claude Sonnet 5
    63.6%
    Claude Opus 5By opponent
    Claude Fable 5
    47.7%
    Claude Fable 5.1
    30.0%
    Gemini 3.5 Flash
    57.1%
    Gemini 3.7 Flash
    47.9%
    Gemini 3.8 Flash
    16.7%
    GLM 5.3 Flash
    70.0%
    GPT-5.6 Luna
    57.9%
    GPT-5.6 Sol
    47.6%
    GPT-5.6 Terra
    53.4%
    GPT-6 Astra
    60.0%
    Grok 4.6
    62.5%
    Claude Haiku 4.5
    53.9%
    Inkling
    46.7%
    Kimi K3
    52.6%
    Muse Spark 1.2
    50.0%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    63.9%
    GPT-6 AstraBy opponent
    Claude Fable 5
    50.0%
    Claude Fable 5.1
    75.0%
    Gemini 3.5 Flash
    37.5%
    Gemini 3.7 Flash
    0.0%
    Gemini 3.8 Flash
    14.3%
    GLM 5.3 Flash
    75.0%
    GPT-5.6 Luna
    44.4%
    GPT-5.6 Sol
    33.3%
    GPT-5.6 Terra
    0.0%
    Grok 4.6
    57.1%
    Claude Haiku 4.5
    68.8%
    Inkling
    66.7%
    Kimi K3
    30.0%
    Muse Spark 1.2
    22.7%
    Claude Opus 5
    40.0%
    Qwen3.8 Max
    20.0%
    Claude Sonnet 5
    41.7%
    Claude Haiku 4.5By opponent
    Claude Fable 5
    38.9%
    Claude Fable 5.1
    50.0%
    Gemini 3.5 Flash
    54.7%
    Gemini 3.7 Flash
    50.0%
    Gemini 3.8 Flash
    33.3%
    GLM 5.3 Flash
    50.0%
    GPT-5.6 Luna
    44.9%
    GPT-5.6 Sol
    41.9%
    GPT-5.6 Terra
    37.0%
    GPT-6 Astra
    31.3%
    Grok 4.6
    58.1%
    Inkling
    55.2%
    Kimi K3
    54.8%
    Muse Spark 1.2
    34.2%
    Claude Opus 5
    46.1%
    Qwen3.8 Max
    52.8%
    Claude Sonnet 5
    36.5%
    Claude Sonnet 5By opponent
    Claude Fable 5
    35.0%
    Claude Fable 5.1
    36.4%
    Gemini 3.5 Flash
    43.3%
    Gemini 3.7 Flash
    37.7%
    Gemini 3.8 Flash
    33.3%
    GLM 5.3 Flash
    72.2%
    GPT-5.6 Luna
    51.0%
    GPT-5.6 Sol
    43.3%
    GPT-5.6 Terra
    48.1%
    GPT-6 Astra
    58.3%
    Grok 4.6
    50.0%
    Claude Haiku 4.5
    63.5%
    Inkling
    71.7%
    Kimi K3
    64.0%
    Muse Spark 1.2
    75.0%
    Claude Opus 5
    36.1%
    Qwen3.8 Max
    52.3%
    Grok 4.6By opponent
    Claude Fable 5
    29.2%
    Claude Fable 5.1
    45.0%
    Gemini 3.5 Flash
    26.6%
    Gemini 3.7 Flash
    37.5%
    Gemini 3.8 Flash
    31.8%
    GLM 5.3 Flash
    59.1%
    GPT-5.6 Luna
    53.2%
    GPT-5.6 Sol
    27.8%
    GPT-5.6 Terra
    60.4%
    GPT-6 Astra
    42.9%
    Claude Haiku 4.5
    41.9%
    Inkling
    50.0%
    Kimi K3
    46.4%
    Muse Spark 1.2
    59.4%
    Claude Opus 5
    37.5%
    Qwen3.8 Max
    50.0%
    Claude Sonnet 5
    50.0%
    GLM 5.3 FlashBy opponent
    Claude Fable 5
    54.2%
    Claude Fable 5.1
    33.3%
    Gemini 3.5 Flash
    45.5%
    Gemini 3.7 Flash
    50.0%
    Gemini 3.8 Flash
    0.0%
    GPT-5.6 Luna
    43.8%
    GPT-5.6 Sol
    43.8%
    GPT-5.6 Terra
    21.4%
    GPT-6 Astra
    25.0%
    Grok 4.6
    40.9%
    Claude Haiku 4.5
    50.0%
    Inkling
    56.3%
    Kimi K3
    43.8%
    Muse Spark 1.2
    55.6%
    Muse Spark 1.3
    100.0%
    Claude Opus 5
    30.0%
    Qwen3.8 Max
    16.7%
    Claude Sonnet 5
    27.8%
    Muse Spark 1.3By opponent
    Claude Fable 5
    0.0%
    GLM 5.3 Flash
    0.0%
    GPT-5.6 Terra
    0.0%

    Variance

    Upsets, ties, and how separable its battles are.

    Beat a higher rating
    higher is better

    Wins against opponents whose published rating was higher at the time of the fit.

    Over ranked battles against higher-rated opponents

    Interpretation: Computed against the CURRENT fit, not the rating at the moment of the battle — a retrospective label on a past result, which is the honest way to do it with a fit that has no history.

    1. 1Gemini 3.7 Flash
      54%
      95% 48%–61%
    2. 2Gemini 3.8 Flash
      49%
      95% 40%–59%
    3. 3GPT-5.6 Luna
      47%
      95% 41%–53%
    4. 4Claude Opus 5
      47%
      95% 42%–51%
    5. 5Qwen3.8 Max
      47%
      95% 40%–53%
    6. 6Claude Fable 5.1
      47%
      95% 36%–58%
    7. 7Kimi K3
      46%
      95% 39%–54%
    8. 8Claude Haiku 4.5
      44%
      95% 39%–49%
    9. 9GPT-5.6 Terra
      43%
      95% 39%–48%
    10. 10Inkling
      43%
      95% 37%–49%
    11. 11Gemini 3.5 Flash
      43%
      95% 37%–49%
    12. 12Claude Sonnet 5
      42%
      95% 37%–46%
    13. 13GPT-5.6 Sol
      42%
      95% 36%–47%
    14. 14GPT-6 Astra
      42%
      95% 30%–54%
    15. 15Grok 4.6
      39%
      95% 34%–45%
    16. 16GLM 5.3 Flash
      38%
      95% 29%–48%
    17. 17Muse Spark 1.2
      35%
      95% 29%–42%
    18. 18Muse Spark 1.3
      0%
      95% 0%–56%
    Inseparable battles
    fingerprint

    Ranked battles judged a tie or both-bad.

    Over ranked battles

    Interpretation: As much a property of the task as of the agent: easy tasks tie because both sides succeed, impossible ones because both fail.

    1. Gemini 3.7 Flash
      22%
      95% 19%–25%
    2. Gemini 3.8 Flash
      16%
      95% 12%–21%
    3. Qwen3.8 Max
      26%
      95% 22%–30%
    4. GPT-5.6 Sol
      25%
      95% 23%–27%
    5. GPT-5.6 Luna
      20%
      95% 18%–24%
    output contamination
    0.0%
    wrong answer
    11.8%
    incomplete
    52.9%
    unnecessary navigation
    14.7%
    repeated action
    0.0%
    misclick
    20.6%
    slow recovery
    0.0%
    Qwen3.8 MaxComposition
    output contamination
    0.0%
    wrong answer
    13.9%
    incomplete
    44.4%
    unnecessary navigation
    30.6%
    repeated action
    0.0%
    misclick
    11.1%
    slow recovery
    0.0%
    GPT-5.6 SolComposition
    output contamination
    0.0%
    wrong answer
    10.3%
    incomplete
    61.5%
    unnecessary navigation
    17.9%
    repeated action
    0.0%
    misclick
    10.3%
    slow recovery
    0.0%
    GPT-5.6 LunaComposition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    66.7%
    unnecessary navigation
    11.1%
    repeated action
    0.0%
    misclick
    22.2%
    slow recovery
    0.0%
    Claude Fable 5Composition
    output contamination
    0.0%
    wrong answer
    12.5%
    incomplete
    62.5%
    unnecessary navigation
    25.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Gemini 3.5 FlashComposition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    77.8%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    22.2%
    slow recovery
    0.0%
    Muse Spark 1.2Composition
    output contamination
    0.0%
    wrong answer
    6.3%
    incomplete
    50.0%
    unnecessary navigation
    31.3%
    repeated action
    0.0%
    misclick
    12.5%
    slow recovery
    0.0%
    Kimi K3Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    72.7%
    unnecessary navigation
    27.3%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    GPT-5.6 TerraComposition
    output contamination
    0.0%
    wrong answer
    12.5%
    incomplete
    37.5%
    unnecessary navigation
    37.5%
    repeated action
    0.0%
    misclick
    12.5%
    slow recovery
    0.0%
    InklingComposition
    output contamination
    0.0%
    wrong answer
    22.2%
    incomplete
    55.6%
    unnecessary navigation
    11.1%
    repeated action
    0.0%
    misclick
    11.1%
    slow recovery
    0.0%
    Claude Fable 5.1Composition
    output contamination
    0.0%
    wrong answer
    8.3%
    incomplete
    61.1%
    unnecessary navigation
    25.0%
    repeated action
    0.0%
    misclick
    5.6%
    slow recovery
    0.0%
    Claude Opus 5Composition
    output contamination
    0.0%
    wrong answer
    10.0%
    incomplete
    70.0%
    unnecessary navigation
    10.0%
    repeated action
    0.0%
    misclick
    10.0%
    slow recovery
    0.0%
    GPT-6 AstraComposition
    output contamination
    0.0%
    wrong answer
    1.8%
    incomplete
    61.8%
    unnecessary navigation
    23.6%
    repeated action
    0.0%
    misclick
    12.7%
    slow recovery
    0.0%
    Claude Haiku 4.5Composition
    output contamination
    0.0%
    wrong answer
    15.8%
    incomplete
    52.6%
    unnecessary navigation
    15.8%
    repeated action
    0.0%
    misclick
    15.8%
    slow recovery
    0.0%
    Claude Sonnet 5Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    33.3%
    unnecessary navigation
    40.0%
    repeated action
    0.0%
    misclick
    26.7%
    slow recovery
    0.0%
    Grok 4.6Composition
    output contamination
    0.0%
    wrong answer
    5.9%
    incomplete
    76.5%
    unnecessary navigation
    17.6%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    GLM 5.3 FlashComposition
    output contamination
    0.0%
    wrong answer
    12.5%
    incomplete
    62.5%
    unnecessary navigation
    25.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Muse Spark 1.3Composition
    output contamination
    0.0%
    wrong answer
    0.0%
    incomplete
    100.0%
    unnecessary navigation
    0.0%
    repeated action
    0.0%
    misclick
    0.0%
    slow recovery
    0.0%
    Flagged runs
    lower is better

    Runs that collected at least one flaw annotation.

    Over runs that were annotated at all

    Interpretation: Annotation is voluntary and judges annotate what looks interesting, so the annotated set is not a random sample. This is the incidence the profile is missing, with the selection bias stated.

    1. 1Claude Fable 5
      73%
      95% 43%–90%
    2. 2Gemini 3.5 Flash
      75%
      95% 47%–91%
    3. 3Gemini 3.7 Flash
      75%
      95% 63%–84%
    4. 4Qwen3.8 Max
      84%
      95% 70%–92%
    5. 5Claude Sonnet 5
      88%
      95% 66%–97%
    6. 6GPT-5.6 Terra
      89%
      95% 56%–98%
    7. 6GLM 5.3 Flash
      89%
      95% 56%–98%
    8. 8Gemini 3.8 Flash
      89%
      95% 76%–96%
    9. 9Inkling
      90%
      95% 60%–98%
    10. 10GPT-5.6 Sol
      90%
      95% 78%–96%
    11. 11Claude Opus 5
      91%
      95% 62%–98%
    12. 12Kimi K3
      92%
      95% 65%–99%
    13. 13Claude Fable 5.1
      95%
      95% 82%–99%
    14. 14GPT-5.6 Luna
      100%
      95% 70%–100%
    15. 14Muse Spark 1.2
      100%
      95% 81%–100%
    16. 14GPT-6 Astra
      100%
      95% 93%–100%
    17. 14Claude Haiku 4.5
      100%
      95% 83%–100%
    18. 14Grok 4.6
      100%
      95% 82%–100%
    19. 14Muse Spark 1.3
      100%
      95% 21%–100%

    Character

    The low-effort labels judges reach for.

    Character stamps
    fingerprint

    Share of the agent's stamps across the five run-level character labels.

    Over stamps on this agent's runs

    Interpretation: The cheapest label a judge can leave, and therefore the most plentiful and the least considered.

    Gemini 3.7 FlashComposition
    wandered
    25.0%
    recovered
    25.0%
    speedran
    25.0%
    made it up
    25.0%
    nailed it
    0.0%
    Claude Fable 5Composition
    wandered
    0.0%
    recovered
    0.0%
    speedran
    100.0%
    made it up
    0.0%
    nailed it
    0.0%
    Claude Sonnet 5Composition
    wandered
    0.0%
    recovered
    0.0%
    speedran
    100.0%
    made it up
    0.0%
    nailed it
    0.0%
    GLM 5.3 FlashComposition
    wandered
    0.0%
    recovered
    0.0%
    speedran
    0.0%
    made it up
    0.0%
    nailed it
    100.0%

    Preference

    Strength and separability of the labels its battles produced.

    Margin when it won
    higher is better

    Mean judge-reported strength (barely / clearly / not close, as 1-3) on battles this agent won.

    Over wins where the judge tapped the optional strength chip

    Interpretation: The chip is optional and skippable by design, so this is measured on a self-selected minority of wins — judges are likelier to record a margin when it was decisive.

    1. 1GPT-5.6 Luna
      2.08
    2. 2GPT-5.6 Terra
      1.92
    3. 3GPT-5.6 Sol
      1.85
    4. 4Claude Opus 5
      1.77
    5. 5Claude Fable 5
      1.72
    6. 6Gemini 3.5 Flash
      1.71
    7. 7Claude Sonnet 5
      1.58
    8. 8Claude Haiku 4.5
      1.55
    1.5%
    95% 1.1%–2.1%
  • 5GPT-5.6 Terra
    1.8%
    95% 1.0%–3.2%
  • 6Gemini 3.5 Flash
    2.1%
    95% 1.4%–3.1%
  • 7Gemini 3.8 Flash
    2.2%
    95% 1.8%–2.7%
  • 8GPT-6 Astra
    2.3%
    95% 1.8%–2.9%
  • 9Gemini 3.7 Flash
    2.4%
    95% 1.9%–2.8%
  • 10GLM 5.3 Flash
    2.5%
    95% 1.5%–4.0%
  • 11Grok 4.6
    3.0%
    95% 2.2%–4.1%
  • 12Kimi K3
    3.8%
    95% 2.7%–5.2%
  • 13Claude Opus 5
    3.9%
    95% 2.4%–6.4%
  • 14Claude Fable 5.1
    4.2%
    95% 3.3%–5.3%
  • 15Qwen3.8 Max
    4.7%
    95% 4.2%–5.1%
  • 16Muse Spark 1.3
    8.6%
    95% 5.2%–14%
  • 17Claude Sonnet 5
    9.4%
    95% 7.2%–12%
  • 18Claude Haiku 4.5
    10%
    95% 8.7%–12%
  • 19Claude Fable 5
    20%
    95% 16%–23%
  • 0%
    95% 0%–1.2%
  • 1Claude Fable 5
    0%
    95% 0%–0.7%
  • 1Gemini 3.5 Flash
    0%
    95% 0%–0.3%
  • 1Muse Spark 1.2
    0%
    95% 0%–0.3%
  • 1Kimi K3
    0%
    95% 0%–0.4%
  • 1GPT-5.6 Terra
    0%
    95% 0%–0.6%
  • 1Inkling
    0%
    95% 0%–0.3%
  • 1Claude Fable 5.1
    0%
    95% 0%–0.2%
  • 1Claude Opus 5
    0%
    95% 0%–0.9%
  • 1GPT-6 Astra
    0%
    95% 0%–0.1%
  • 1Claude Haiku 4.5
    0%
    95% 0%–0.3%
  • 1Claude Sonnet 5
    0%
    95% 0.0%–0.7%
  • 1Grok 4.6
    0%
    95% 0%–0.3%
  • 1GLM 5.3 Flash
    0%
    95% 0%–0.6%
  • 1Muse Spark 1.3
    0%
    95% 0%–2.2%
  • Claude Fable 5
    92%
    95% 89%–94%
  • Gemini 3.5 Flash
    89%
    95% 87%–91%
  • Muse Spark 1.2
    91%
    95% 89%–92%
  • Kimi K3
    90%
    95% 88%–92%
  • GPT-5.6 Terra
    85%
    95% 82%–87%
  • Inkling
    82%
    95% 80%–84%
  • Claude Fable 5.1
    87%
    95% 86%–89%
  • Claude Opus 5
    94%
    95% 92%–96%
  • GPT-6 Astra
    73%
    95% 71%–74%
  • Claude Haiku 4.5
    88%
    95% 86%–89%
  • Claude Sonnet 5
    87%
    95% 84%–90%
  • Grok 4.6
    72%
    95% 69%–74%
  • GLM 5.3 Flash
    83%
    95% 80%–86%
  • Muse Spark 1.3
    99%
    95% 97%–100%
  • 0.3%
    95% 0.2%–0.5%
  • 5GPT-5.6 Sol
    0.4%
    95% 0.2%–0.7%
  • 6Claude Fable 5.1
    0.4%
    95% 0.2%–0.8%
  • 7Kimi K3
    0.4%
    95% 0.2%–1.1%
  • 8GPT-5.6 Luna
    0.6%
    95% 0.2%–2.2%
  • 9Muse Spark 1.2
    0.7%
    95% 0.4%–1.2%
  • 10Claude Opus 5
    0.7%
    95% 0.3%–2.1%
  • 11Gemini 3.5 Flash
    0.8%
    95% 0.4%–1.6%
  • 12Gemini 3.8 Flash
    1.0%
    95% 0.8%–1.3%
  • 13Qwen3.8 Max
    1.4%
    95% 1.2%–1.7%
  • 14Claude Sonnet 5
    1.8%
    95% 1.0%–3.2%
  • 15Grok 4.6
    2.0%
    95% 1.3%–2.9%
  • 16GLM 5.3 Flash
    2.0%
    95% 1.2%–3.3%
  • 17Claude Haiku 4.5
    2.1%
    95% 1.5%–3.0%
  • 18Inkling
    2.2%
    95% 1.5%–3.1%
  • 19GPT-6 Astra
    5.3%
    95% 4.5%–6.2%
  • 1
    p10–p90 1–3.60
  • Gemini 3.5 Flash
    3
    p10–p90 1–7.40
  • Muse Spark 1.2
    1
    p10–p90 1–9.40
  • Kimi K3
    1
    p10–p90 1–5.90
  • GPT-5.6 Terra
    2
    p10–p90 1–7
  • Inkling
    2
    p10–p90 1–14.5
  • Claude Fable 5.1
    1
    p10–p90 1–4.10
  • Claude Opus 5
    1
    p10–p90 1–3
  • GPT-6 Astra
    2
    p10–p90 1–17
  • Claude Haiku 4.5
    2
    p10–p90 1–11
  • Claude Sonnet 5
    1
    p10–p90 1–4
  • Grok 4.6
    5.50
    p10–p90 1–27.5
  • GLM 5.3 Flash
    2
    p10–p90 1–19.2
  • Muse Spark 1.3
    1
    p10–p90 1–1.60
  • Claude Fable 5
    40%
    95% 30%–52%
  • Gemini 3.5 Flash
    61%
    95% 53%–69%
  • Muse Spark 1.2
    69%
    95% 62%–76%
  • Kimi K3
    42%
    95% 33%–51%
  • GPT-5.6 Terra
    18%
    95% 12%–26%
  • Inkling
    75%
    95% 69%–80%
  • Claude Fable 5.1
    56%
    95% 50%–62%
  • Claude Opus 5
    33%
    95% 20%–50%
  • GPT-6 Astra
    48%
    95% 44%–52%
  • Claude Haiku 4.5
    51%
    95% 43%–59%
  • Claude Sonnet 5
    40%
    95% 29%–52%
  • Grok 4.6
    71%
    95% 66%–76%
  • GLM 5.3 Flash
    57%
    95% 47%–67%
  • Muse Spark 1.3
    0%
    95% 0%–49%
  • 0.38
    p10–p90 0.28–0.51
  • Claude Fable 5
    0.42
    p10–p90 0.28–0.56
  • Gemini 3.5 Flash
    0.44
    p10–p90 0.28–0.54
  • Muse Spark 1.2
    0.35
    p10–p90 0.24–0.52
  • Kimi K3
    0.30
    p10–p90 0.28–0.47
  • GPT-5.6 Terra
    0.42
    p10–p90 0.28–0.54
  • Inkling
    0.44
    p10–p90 0.26–0.57
  • Claude Fable 5.1
    0.42
    p10–p90 0.28–0.56
  • Claude Opus 5
    0.38
    p10–p90 0.28–0.52
  • GPT-6 Astra
    0.32
    p10–p90 0.26–0.54
  • Claude Haiku 4.5
    0.39
    p10–p90 0.26–0.54
  • Claude Sonnet 5
    0.28
    p10–p90 0.28–0.56
  • Grok 4.6
    0.29
    p10–p90 0.17–0.48
  • GLM 5.3 Flash
    0.28
    p10–p90 0.24–0.49
  • Muse Spark 1.3
    0.44
    p10–p90 0.43–0.44
  • navigate
    45.8%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.6%
    read page
    24.8%
    screenshot
    0.8%
    deliver
    17.1%
    done
    11.0%
    GPT-5.6 SolComposition
    navigate
    42.0%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.1%
    read page
    25.0%
    screenshot
    0.4%
    deliver
    17.0%
    done
    15.5%
    GPT-5.6 LunaComposition
    navigate
    43.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    26.4%
    screenshot
    0.0%
    deliver
    16.3%
    done
    14.0%
    Claude Fable 5Composition
    navigate
    36.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    19.2%
    screenshot
    0.0%
    deliver
    23.2%
    done
    21.2%
    Gemini 3.5 FlashComposition
    navigate
    49.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    24.0%
    screenshot
    0.0%
    deliver
    15.5%
    done
    11.1%
    Muse Spark 1.2Composition
    navigate
    47.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    9.7%
    screenshot
    5.1%
    deliver
    25.1%
    done
    12.7%
    Kimi K3Composition
    navigate
    47.3%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    22.8%
    screenshot
    0.9%
    deliver
    15.6%
    done
    13.4%
    GPT-5.6 TerraComposition
    navigate
    48.0%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    25.1%
    screenshot
    0.4%
    deliver
    13.0%
    done
    13.5%
    InklingComposition
    navigate
    56.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    3.7%
    read page
    15.2%
    screenshot
    6.5%
    deliver
    8.8%
    done
    9.5%
    Claude Fable 5.1Composition
    navigate
    40.8%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    21.9%
    screenshot
    0.8%
    deliver
    18.1%
    done
    18.4%
    Claude Opus 5Composition
    navigate
    45.6%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    12.7%
    screenshot
    0.0%
    deliver
    21.5%
    done
    20.3%
    GPT-6 AstraComposition
    navigate
    51.3%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    31.8%
    screenshot
    0.7%
    deliver
    8.7%
    done
    7.6%
    Claude Haiku 4.5Composition
    navigate
    34.4%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    1.0%
    read page
    16.8%
    screenshot
    26.8%
    deliver
    14.3%
    done
    6.6%
    Claude Sonnet 5Composition
    navigate
    39.6%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    4.4%
    read page
    11.9%
    screenshot
    4.4%
    deliver
    17.6%
    done
    22.0%
    Grok 4.6Composition
    navigate
    70.9%
    click
    0.2%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    3.5%
    read page
    16.3%
    screenshot
    0.0%
    deliver
    4.2%
    done
    4.8%
    GLM 5.3 FlashComposition
    navigate
    63.1%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    20.1%
    screenshot
    0.0%
    deliver
    8.7%
    done
    8.1%
    Muse Spark 1.3Composition
    navigate
    26.7%
    click
    0.0%
    hover
    0.0%
    drag
    0.0%
    type
    0.0%
    press key
    0.0%
    scroll
    0.0%
    switch tab
    0.0%
    read page
    20.0%
    screenshot
    6.7%
    deliver
    40.0%
    done
    6.7%
    0%
    95% 0%–24%
  • 1Gemini 3.5 Flash
    0%
    95% 0%–7.9%
  • 1Kimi K3
    0%
    95% 0%–18%
  • 1GPT-5.6 Terra
    0%
    95% 0%–35%
  • 1Claude Opus 5
    0%
    95% 0%–49%
  • 1Grok 4.6
    0%
    95% 0%–14%
  • 1GLM 5.3 Flash
    0%
    95% 0%–18%
  • 11Gemini 3.8 Flash
    0.7%
    95% 0.1%–3.6%
  • 12Qwen3.8 Max
    1.3%
    95% 0.5%–3.2%
  • 13Claude Fable 5.1
    2.4%
    95% 0.4%–12%
  • 14Muse Spark 1.2
    4.6%
    95% 1.6%–13%
  • 15GPT-6 Astra
    4.8%
    95% 1.7%–13%
  • 16Inkling
    5.8%
    95% 2.3%–14%
  • 17Claude Sonnet 5
    6.7%
    95% 1.2%–30%
  • 18Muse Spark 1.3
    17%
    95% 3.0%–56%
  • 19Claude Haiku 4.5
    19%
    95% 13%–27%
  • 7.7k
    p10–p90 5.1k–16k
  • 5GPT-5.6 Terra
    8.6k
    p10–p90 5.6k–11k
  • 6GLM 5.3 Flash
    8.8k
    p10–p90 5.8k–14k
  • 7Inkling
    8.9k
    p10–p90 7.2k–13k
  • 8Kimi K3
    9.0k
    p10–p90 5.6k–15k
  • 9Claude Fable 5
    9.4k
    p10–p90 7.0k–15k
  • 10Claude Fable 5.1
    9.5k
    p10–p90 6.7k–16k
  • 11Claude Sonnet 5
    9.5k
    p10–p90 7.0k–15k
  • 12Grok 4.6
    9.7k
    p10–p90 6.4k–17k
  • 13Claude Opus 5
    9.8k
    p10–p90 6.7k–18k
  • 14Qwen3.8 Max
    9.9k
    p10–p90 5.9k–20k
  • 15Muse Spark 1.3
    12k
    p10–p90 9.7k–14k
  • 16Gemini 3.8 Flash
    13k
    p10–p90 5.6k–30k
  • 17Claude Haiku 4.5
    14k
    p10–p90 7.4k–22k
  • 18Muse Spark 1.2
    15k
    p10–p90 8.0k–24k
  • 19Gemini 3.5 Flash
    18k
    p10–p90 7.3k–30k
  • 5Claude Fable 5.1
    248k
  • 6Gemini 3.7 Flash
    259k
  • 7Claude Sonnet 5
    301k
  • 8Claude Opus 5
    380k
  • 9Inkling
    443k
  • 10GPT-6 Astra
    503k
  • 11Kimi K3
    594k
  • 12Qwen3.8 Max
    704k
  • 13Muse Spark 1.2
    776k
  • 14GLM 5.3 Flash
    815k
  • 15Gemini 3.8 Flash
    937k
  • 16Grok 4.6
    990k
  • 17Claude Haiku 4.5
    1.0M
  • 18Gemini 3.5 Flash
    1.1M
  • 19Muse Spark 1.3
    2.1M
  • 0%
    95% 0%–9.6%
  • 1Claude Fable 5
    0%
    95% 0%–8.8%
  • 1Gemini 3.5 Flash
    0%
    95% 0%–12%
  • 1Muse Spark 1.2
    0%
    95% 0%–8.6%
  • 1Kimi K3
    0%
    95% 0%–12%
  • 1GPT-5.6 Terra
    0%
    95% 0%–12%
  • 1Inkling
    0%
    95% 0%–10%
  • 1Claude Fable 5.1
    0%
    95% 0%–3.4%
  • 1Claude Opus 5
    0%
    95% 0%–20%
  • 1GPT-6 Astra
    0%
    95% 0%–4.8%
  • 1Claude Haiku 4.5
    0%
    95% 0%–13%
  • 1Claude Sonnet 5
    0%
    95% 0%–12%
  • 1Grok 4.6
    0%
    95% 0%–19%
  • 1GLM 5.3 Flash
    0%
    95% 0%–28%
  • 1Muse Spark 1.3
    0%
    95% 0%–79%
  • Claude Fable 5
    22%
    95% 20%–24%
  • Gemini 3.5 Flash
    23%
    95% 19%–27%
  • Muse Spark 1.2
    22%
    95% 18%–27%
  • Kimi K3
    21%
    95% 18%–25%
  • GPT-5.6 Terra
    26%
    95% 23%–29%
  • Inkling
    21%
    95% 17%–25%
  • Claude Fable 5.1
    27%
    95% 21%–34%
  • Claude Opus 5
    23%
    95% 21%–25%
  • GPT-6 Astra
    17%
    95% 10%–26%
  • Claude Haiku 4.5
    19%
    95% 16%–23%
  • Claude Sonnet 5
    21%
    95% 18%–24%
  • Grok 4.6
    26%
    95% 21%–30%
  • GLM 5.3 Flash
    23%
    95% 16%–31%
  • Muse Spark 1.3
    0%
    95% 0%–56%