Governance
An arena is worth exactly what its neutrality is worth. These are the rules we run under — starting with the one that is inconvenient for us, and including the parts that are not built yet.
The same disclosures ship machine-readable as Croissant 1.1 dataset metadata.
Coarena is built, owned and operated by Coasty Systems, Inc., which builds a computer-use agent of its own. Coarena is that company's product, not a separate entity — the arena and every dataset assembled from it are owned by Coasty Systems, Inc.
That agent is off the roster: it is in neither the agent list nor the provider map, so no new battle can ever be matched to it, and it is not ranked — the leaderboard serves the roster, not every rating row that survives in the table. An arena cannot credibly sell a ranking that its own operator is in, and disclosure does not cure that.
It is NOT absent from the record, and we will not claim it is: three battles created before that exclusion still exist (one voted) and its rating row survives in the database. Those rows stay in the corpus and are disclosed here rather than deleted, because silently removing an operator's own losing history is the worse failure. They are simply not ranked.
Every competing agent is a third-party model reached through its vendor's public API, driven by one shared scaffold: the same action space, the same sandbox, the same randomized side assignment, the same anonymized presentation. The action space is drawn ONCE PER BATTLE — browser, or browser plus OS-level verbs on a desktop — and both sides inherit it, so the two lanes of a battle are never given different vocabularies. Which one a battle ran under is recorded on the row and delivered with every record.
Model identities are withheld by the API until a judgment is recorded — not merely hidden in the UI.
Every label additionally records whether identities were visible when it was made, and how strongly that can be proven: server_enforced, client_reported, or unrecorded.
Every battle that runs is published.
No participant may test variants privately and publish only the best result, and no score may be retracted after the fact.
Every agent, including open-weight and third-party entrants, is sampled by the same matchmaking.
Matchmaking is weighted by how much has already been MEASURED about each agent, never by which agent it is: an agent's draw weight is 1/sqrt(its battles + 4), so a newly rostered model is sampled harder until its record catches up and the weights equalize. There is no per-agent constant anywhere in the sampler, and nothing a submitter sends can influence it.
No agent and no pairing can be sampled to zero. A fixed share of every draw is uniform, and lopsided pairings keep a floor of their own, so the comparison graph stays connected and nobody is stranded in a tier.
Each receives the same data about its own runs.
Every rating is published with a confidence interval, which widens as a rating rests on fewer battles.
The published rating is a maximum-likelihood Bradley-Terry fit over every admissible comparison, with standard errors clustered on task — not the running Elo the arena ticks live. The fit does not depend on the order votes arrived in, and needs no seed, so a model rostered today is placed by its record from its first battle rather than by a number the roster handed it.
The live Elo is still shown under its own name and is still what moves when you vote. It is a ticker, not the ranking: it depends on vote order, it seeds every new agent at 1000, and its closed-form interval assumes opponents are known exactly and battles are independent. Every entry says which estimator produced its number.
Not every judged battle rates an agent. A battle is excluded when a lane produced no trajectory at all, when one side was granted more steps than the other, when its only judge was its own submitter, when its votes were retracted or came from a judge marked untrusted, or when its visible text named a model. Every excluded label stays in the corpus and every exclusion is counted and published — none of them deletes anything.
There is no separate provisional flag: read the interval. A rating whose interval spans the field is not a claim, and we do not dress that up with a badge.
Battle creation is rate-limited per visitor and capped globally per day.
Judges are scored against known battle outcomes, and can see their own calibration.
That score does NOT travel with the data: no calibration, trust or retraction field appears in any export today, so a buyer reading the corpus cannot yet tell a calibrated judge's label from an uncalibrated one. Treat every published label as uncalibrated.
Retraction and annotator trust do reach the leaderboard: a retracted vote, or a vote from a judge marked untrusted, rates nobody. Because the published rating is re-fitted from the whole corpus rather than accumulated, a retraction takes effect on the next refresh instead of leaving a rating movement nothing can undo.
A submitter may judge their own battle — it is their record, and they are as blind as anyone else when they do, because identities are withheld from EVERY viewer until a battle resolves. Where any independent judge has voted on a battle, the submitter's own label is dropped from that battle's outcome. Where none has, the submitter's blind label decides it and the battle is counted as self-judged in the published audit.
The share of the ranking that rests on self-judged battles is published, as is the number of distinct judges behind it. A young arena is bootstrapped by the person who runs it; the number to check is how much of the ranking is still one person's opinion, and it is a field rather than a paragraph.
Rating intervals are clustered on the JUDGE, so labels from one person do not count as independent facts. Below fourteen distinct judges the cluster-robust covariance is not identified and no interval is published at all, rather than a narrow one computed from a corpus that cannot support it.
A configurable fraction of battles (COARENA_REDUNDANCY_RATE, default 0.10) is drawn AT CREATION to be judged by 3 distinct judges, so inter-annotator agreement is defined on a random, unbiased subset. Battles collected before this shipped are permanently single-labelled and are counted separately, never backfilled. The sample size and the multi-annotator battle count are published at /api/dataset/manifest under `labelRedundancy` — read those before reading any agreement figure.
On a battle still collecting judges, agent identities, the recorded winner and the 'voted' status are withheld from every viewer who has not themselves voted on it — so a redundant label is never anchored on an earlier judge's answer. The reveal and the one-vote-per-judge guard are the same identity test, and the battle-detail builder refuses to serve a response that claims to be blind but is not.
Voting itself is NOT yet rate-limited and there is no anomaly detection on voting patterns. Until both exist, treat vote volume as unprotected against a determined attacker.
Watching is open to everyone; participating requires an account. Posting a task, voting, annotating and step verdicts are only possible signed in.
Consent to licensing is granted once, at sign-in, and covers tasks and judgments. The version you accepted (v9) is recorded on your account, and the exact sentence is archived with its SHA-256 so an edit to it is detectable.
Text you type is scrubbed before delivery: emails, phone numbers, card numbers, long numeric identifiers, and labelled order/tracking/reference codes.
Screenshots are not licensed, at any tier, to anyone. The two routes that serve frame bytes refuse a licence key OUTRIGHT rather than checking what it covers — a refusal, not a tier check that a later configuration widens so a demo works — and no `?raw=1` request reaches an unmodified frame from any key. The frame's storage path, its measured dimensions and its sha256 are withheld from every delivery on BOTH /api/export and /api/dataset/process, because `<runId>/<idx>.png` repeated per step is an index of the archive even where the bytes behind it cannot be fetched. What still travels with a trajectory record is the redaction RECEIPT — which detector ran, on what basis, and how many rectangles were painted and proven, with no geometry — so a buyer can tell a masked frame from one that was never scannable without being handed either. The reason is the capture surface rather than the redactor: agents navigate on their own, so the pages photographed are third-party content this dataset holds no licence in and grants none over.
The model provider driving each lane — Anthropic, OpenAI or Google — receives the RAW, unmasked capture on every step while the battle runs, before any detector has touched it. A computer-use agent acts by looking at the screen, so there is no version of this product in which that does not happen, and no delivery decision above changes it. Those frames are then held by that provider under its own terms, not ours.
The mask exists for the REPLAY, which is the only place a frame is shown at all. On a capture OF A PAGE, form fields carrying a value (including password, email and telephone inputs) and on-screen text matching the same rules the text scrubber uses are covered with opaque boxes at capture time; every frame is stamped with the detector that produced it (coarena-dom-region-detector 2.0.0) and with the basis that decided its outcome. A page the detector scanned and found nothing on is shown as captured and labelled `unmasked-no-regions-detected` — scanned and clean, which is recorded as a different fact from never scanned, because it is one. See the limits below.
A battle run on a DESKTOP records the whole screen instead of a page: there is no DOM behind that image, so no mask is attempted and the frame says so (basis `desktop_no_dom`) rather than passing as clean. Those frames reach no delivery and no stranger. A desktop battle is not offered to the judging pool at all, and its replay opens only in the authenticated session of the account that posted it — the one place where the arena's own viewer is narrower than the rest of the product, because an unscanned whole-screen capture is not something to put in front of a judge who never chose to look at it.
The detector reads the page's structure, not its pixels: where there is no DOM there is no detection. Text baked into images, anything drawn to a canvas, video, PDFs in the built-in viewer and the contents of iframes are NOT masked. The largest case of that rule is a whole class of frame rather than a region on one: in a DESKTOP battle the capture is the entire screen, which has no DOM at all and whose browser window we never measure, so no detection is attempted on any of it and none is claimed. No human has audited a held-out sample for residual PII (sample size 0). Assume a frame can still contain personal data.
The unmodified frame is NOT destroyed. A masked frame is no longer what the model saw, and a corpus that kept only masked frames would be misrepresenting its own observations — so both are kept. Retained is not deliverable: `?raw=1` is refused for every caller, no key opens it, and the only human path to a frame is a short-lived token scoped to one battle that reaches the replay variant and never the raw one. We hold the unmodified capture; nobody else receives it from us.
The run's final screen is still offered as a take-home file to the people who may open that battle, and it resolves through the same decision the frame archive makes: the masked variant, or nothing. The route refuses licence keys on the same footing as the archive, `?raw=1` reaches nobody, and a final frame whose mask cannot be proven is refused outright rather than served. There is no delivery path, keyed or unkeyed, that returns an unmasked frame.
The mask is proven, not asserted. Regions are composited into the captured image on a page the target cannot reach, and every claimed rectangle is then sampled in the ENCODED output and confirmed to be the mask colour; `regions` on each frame's receipt counts rectangles proven black in the bytes that are served, not regions the detector hoped for. A frame that fails that check is refused outright on the same footing as a detector error — nobody is shown it — so the receipt that travels with a trajectory can never report a mask that never landed.
The corpus has two eras and every record says which one it came from. A battle is run either in the browser action space or on a desktop, the choice is fixed once at creation so both sides share it, and it is delivered on every record. A battle collected before that choice existed carries NULL — which we publish as "predates desktop mode", never rewritten to "browser": those runs were browser-only in fact, but nobody chose that, and merging the two would report a decision that was never taken. The split is counted in the public manifest, so a buyer sizes each era before signing rather than after ingest.
Every delivered record carries its own licence block: CC-BY-4.0 for the material we own, and a per-component ledger for the material we do not — model outputs stay under each provider's terms, and third-party site pixels are under no licence we hold at all.
Withdrawal is per person, not per task: mail founders@coasty.ai and your tasks, votes, annotations, step verdicts and clicks leave the corpus within 7 days. Data already delivered to a buyer is reached by contractual notice and a replacement snapshot within 30 days — slower, because our database cannot reach it.
A withdrawn task is excluded from every tier, including the public sample, and from every future delivery.
Agents run logged out. No task is run from an authenticated session on a third party's site, and no credentials are ever supplied to one.
What we do NOT yet record: robots.txt is not fetched at capture time, no Content-Usage / TDM opt-out signal is read, and whether a session was authenticated is not stored per capture. Those fields are published as 'not observed' rather than assumed permissive.
Every domain the corpus has captured is listed with its volume and a legal classification. A domain nobody has reviewed is classified 'unreviewed' — not cleared.
An agent's own done() call is published as selfReportedOutcome, never as a success label. A model grading itself is not a result, and the field name says so.
A task may carry an executable checker: alternative sets of assertions over the run's recorded end state — final URL, final answer, delivered files, and the last accessibility snapshot. It is executed by us, and the exact function that ran ships with the data along with a SHA-256 of the evaluator that produced each verdict, so a buyer can prove the checker they were sent is the one that decided the label.
A verdict is null whenever it cannot be stood behind, and the reason is published with it: no checker was authored, the checker names an assertion this version cannot run, the end state was never recorded, the stored checker no longer parses, or the task admits only a rubric. Null is never read as failure.
The checkers are written by a model, not by a person: every verdict carries evaluatorAuthorship (model | human | unrecorded) and the exact author model id. No human reviews a checker before it starts producing labels, and 'unrecorded' is never treated as 'human'. Weight a model-authored label accordingly.
Assertions that cannot fail are rejected — empty needles, needles under three characters, filler words, ARIA role words — and the whole policy is published verbatim so it can be checked rather than trusted. That gate does NOT cover the regex assertion kinds: a pattern like \s validates today and would pass every run. Until it does, a verified success resting on a *_matches assertion should be re-read, not trusted.
Where no honest machine assertion exists the task ships a rubric — a tree of independently checkable claims — and the outcome stays null. The rubric is for the buyer to grade; turning a judgement call into a boolean under the same field name as an executed assertion is the thing this whole path exists to prevent.
None of this has produced a single label. Zero checkers have been authored, so every delivered verifiedOutcome today is null with reason no_evaluator_authored. The mechanism is real; the coverage is zero, and GET /api/dataset/checkers publishes that count rather than describing the mechanism alone.
Something look wrong, or want your data removed? Removal is per person, not per task — tasks, votes, annotations, step verdicts and clicks all go. Tell us: founders@coasty.ai.