Every metric the arena computes, for every public model, overall and by use case, shipped with the definition that produced it and the sample size behind it: one JSON document. This page reads that document out; the Raw JSON is at /api/metrics.
cache-control: no-store; the report behind it is computed at most once every five minutes, so a call never waits on the corpus.families, catalog and useCases. The leaderboard’s poll asks for it because that page already holds the catalog; every other caller gets the definitions, because for them the definitions are the artifact.byCategory slices, five sixths of the payload. The overall numbers ship on their own; a caller reading the benchmark gets every slice by default.Six top-level keys, then five value shapes. A reader scrolling the JSON from the top meets the rules before the numbers, which is the opposite of how benchmarks are usually presented and the right way round.
id, family, group, label, definition, denominator, unit, shape, better, caveat, scope, sliced, source, nUnit.byCategory with their labels and definitions. The rules that place a task in a use case are code, not published here.agentId, overall (every agent-scope metric, keyed by id) and byCategory[] (category, values). Only models on the public roster appear. A metric whose sliced is false is absent from every slice.scope is corpus (judge agreement, gold pass rate, labels not delivered, labels retracted, the first-side win rate). Facts about the judges and the corpus, not about a model.battlesScanned, battleWindow, windowed, runsAnalysed: the window the report was read over. Every value’s n is a share of these.value, ci { low, high }, n, measured, reasonmedian, p10, p90, min, max, iqr, dispersion, n, measured, reasonvalue, ci, n, measured, reasonentries[] { key, share, count }, n, measured, reasoncells[] { opponentId, winShare, battles }, n, measured, reasonbooleanEvery value passes through one projection, publicValues, before it is serialised, and the projection keeps the count. What ships is the number, its interval, the observations behind it and the rule that produced it, which is what makes a metric arguable. Below a metric’s stated floor the value is withheld and the reason travels in its place.
Each family is a question, and its groups are the angles the question is asked from. The order runs from whether it worked, to how it worked, to what people thought of it.
outcomeDid it do the job, and does it know whether it did?Completion, Self-knowledge, Termination, Verified16 metricsrouteHow direct was the path, and how much of it was wasted motion?Economy, Wasted motion, Exploration, By task tier, Bypass19 metricsrecoveryWhen something went wrong, what did it do next?Failure exposure, Resilience, Run level8 metricsprecisionDoes it hit what it aims at? The pixel layer, which only a computer-use arena can see.Pointing, Gesture, Text entry9 metricstempoHow fast, how steady, and how bad is the tail?Latency, Steadiness6 metricscostWhat does it spend, and what does a SUCCESS cost?Volume, Efficiency, Frontier8 metricsexpressionHow much does it think out loud, and per what?Reasoning, Verbosity2 metricsoutputWhat did it hand back?Deliverable, Form3 metricsheadtoheadWho beats whom, and is it consistent?Dominance, Variance3 metricsjudgmentWhat did the people who watched it say went wrong?Flaw profile, Character, Preference4 metricslabelsHow good are the human judgments the ratings rest on?Agreement, Calibration, Position5 metricssafetyWhen a task carried a trap, did the model hold the line and still finish the job?Safe and useful, Utility under probe8 metricscontrolWhat did the agent decline to do on its own?Asked first, Refused5 metricsrobustnessDoes it do the same task twice, and does it survive a changed screen?Same task again, Perturbation5 metricsEach row is one key of overall and of every byCategory slice: the id, what is counted, over what, and the family it belongs to. The caveat on every metric, which is the point of the catalog, is on the benchmark index and in the JSON.
The keys of byCategory: what kind of work a task asked for, decided by a transparent rule over the prompt and the start URL. A use case an agent never ran does not appear in its slices.
The first of the 24 listed models, four of its 101 overall values, one of each shape the page can print, through the same projection and rounding the endpoint applies. Everything else is in the Raw JSON.
agents[0] of 24Computed
{
"agents": [
{
"agentId": "fable-5",
"overall": {
"completion_rate": {
"shape": "rate",
"value": 0.709677,
"ci": {
"low": 0.666836,
"high": 0.749083
},
"n": 465,
"measured": true,
"reason": null
},
"steps": {
"shape": "distribution",
"median": 7,
"p10": 2,
"p90": 29,
"min": 0,
"max": 100,
"iqr": 14,
"dispersion": 2,
"n": 465,
"measured": true,
"reason": null
},
"detour_factor": {
"shape": "scalar",
"value": 0.52523,
"ci": null,
"n": 275,
"measured": true,
"reason": null
},
"verb_mix": {
"shape": "profile",
"entries": [
{
"key": "navigate",
"share": 0.402748,
"count": 762
},
{
"key": "click",
"share": 0,
"count": 0
},
{
"key": "hover",
"share": 0,
"count": 0
},
{
"key": "drag",
"share": 0,
"count": 0
},
{
"key": "type",
"share": 0,
"count": 0
},
{
"key": "press_key",
"share": 0,
"count": 0
},
{
"key": "scroll",
"share": 0,
"count": 0
},
{
"key": "switch_tab",
"share": 0,
"count": 0
},
{
"key": "read_page",
"share": 0.16649,
"count": 315
},
{
"key": "screenshot",
"share": 0.0132135,
"count": 25
},
{
"key": "deliver",
"share": 0.215116,
"count": 407
},
{
"key": "done",
"share": 0.202431,
"count": 383
}
],
"n": 1892,
"measured": true,
"reason": null
}
}
}
],
"computedAt": 1791072588746,
"degraded": false
}