Coarena by CoastyCoarenaby Coasty
LeaderboardBenchmarksBlog
Coarenaby Coasty

Real-world evals for computer-use agents. Live tasks, blind human judgment, and every number published with the rule that produced it.

Arena

  • Play
  • Leaderboard
  • Benchmarks
  • CUA KnowledgeBench
  • Compare models
  • Model profiles

Evidence

  • Dataset
  • Methodology
  • Metrics API
  • Cite us

About

  • Mission
  • Governance
  • Blog

Legal

  • Terms
  • Privacy
  • Security

© 2026 Coasty Systems, Inc.

Every claim on this site cites the file that keeps it
Benchmarks / KnowledgeBench
Private evaluation

CUA KnowledgeBench (12-Hours & 1-Day)

Real computer-use capability across more than 100 applications. Built around reasoning, current information, and the way people and companies actually work.

Discuss an evaluation Explore the editions
Inside the working environment
Illustrative fictional workspace with an operations spreadsheet, team inbox, and project board. This is not an actual benchmark capture.

One piece of work. Several connected tools.

An inbox, an operational spreadsheet, and a project board show how information and responsibilities can be distributed across a working environment.

View full size ↗ (opens image in a new tab)

Illustrative workspaces · AI-generated images with fictional data. Actual benchmark environments and materials remain private.

100+Applications across the benchmark
12-Hours & 1-DayPrivate evaluations by arrangement
3-Days & 7-DaysComing soon
EditionsEnvironmentCapabilitiesPrivate evaluationResults
01 / Editions

A longer view of capability.

Two editions for private evaluation today. Two longer editions on the horizon.

Available by arrangement

12-Hours

CUA KnowledgeBench (12-Hours)

A 12-hour private evaluation of reasoning and computer use in realistic, connected work environments.

Discuss 12-Hours
Available by arrangement

1-Day

CUA KnowledgeBench (1-Day)

A full-day (24-hour) private evaluation across applications, evolving information, and practical work.

Discuss 1-Day

Contact us to discuss the edition, execution setup, and timing for your agent.

Coming soon

CUA KnowledgeBench (3-Days & 7-Days)

Two longer editions for sustained knowledge work. Availability and evaluation details will be announced.

02 / The environment

The whole working context.

Over 100 applications across the benchmark, with the interconnected context of real-world work. The emphasis is on how tools, information, and human decisions fit together.

Software is only part of the environment.

The harder part is understanding what is happening inside it. People leave decisions in conversations, maintain information in different places, and depend on work completed by others. A realistic evaluation needs to reflect that surrounding context.

Information has a history—and a present.

Work draws on what was known before and what has changed since. Logic and current information matter together: finding a recent fact is useful only when the agent understands how it affects the work already in progress.

01

Communication

Conversations, inboxes, and updates carry intent, decisions, and changes in direction.

02

Documents & knowledge

Working documents and reference material hold the context that decisions depend on.

03

Spreadsheets & analysis

Structured information needs interpretation, reconciliation, and consistent use elsewhere.

04

Projects & operations

Work has owners, dependencies, status, and consequences for the next person in the process.

05

Customer & business context

Records connect people and organisations to the history and details of ongoing work.

06

Research & coordination

Current sources, schedules, and shared commitments connect information to action.

These are the kinds of workplace context the benchmark is designed to reflect. The 100+ application figure describes the benchmark as a whole; it does not mean every application is used in every evaluation. The detailed application roster and evaluation cases are private.

03 / Capabilities

From operating software
to carrying work forward.

CUA KnowledgeBench is designed to examine practical computer-use capability in context. These are the questions that make that capability meaningful.

01Reason through connected decisionsInterpret context, reconcile information, and understand what depends on what.+

Useful computer work involves choosing a sensible course of action from incomplete or conflicting information. An agent needs to connect facts across sources, recognise constraints, and understand how an early decision affects later work. The reasoning has to remain consistent as it becomes actions in software.

02Work with information that changesFind current information and recognise when earlier assumptions need revisiting.+

Real-time information is part of the work itself. A source can change, an update can supersede an earlier message, and a previously reasonable assumption can become outdated. Capable agents need to find relevant information, consider its freshness, and use it in context rather than treating an earlier observation as permanently true.

03Carry context across applicationsKeep the meaning of the work intact as it moves between tools.+

Information takes different forms in different applications: a conversation, a document, a table, a calendar event, or a record. Moving between those interfaces requires recognising the same people, entities, dates, and decisions. The capability of interest is whether the agent can use those connections accurately while operating the computer.

04Maintain continuity over timeKeep track of progress, open questions, and commitments as work develops.+

Extended work puts pressure on consistency. An agent needs to distinguish completed work from pending work, retain relevant context, and avoid repeating or contradicting earlier actions. This is a question of continuity: whether decisions and actions remain coherent as the surrounding work develops.

05Respond to uncertainty and changeRecognise a changed situation and make a considered next decision.+

The way people work includes interruptions, ambiguity, missing information, and competing priorities. A capable agent needs to notice when the situation differs from its assumptions, reassess what is known, and determine what can reasonably happen next. Continuing an outdated plan can be as consequential as choosing the wrong action initially.

06Produce work that people can useBring accuracy, consistency, and context into the resulting work.+

The purpose of computer use is a useful outcome. Documents, records, analyses, and handoffs need to make sense to the people who rely on them. This means attention to the information behind the work, consistency between related artifacts, and clarity about what has been established and what remains unresolved.

These descriptions explain the evaluation focus. They are not a published scoring rubric or a set of benchmark tasks.

04 / Private by design

The benchmark stays private.
The ambition is clear.

CUA KnowledgeBench is one of our most closely held benchmarks. Its purpose is to measure real capability through realistic computer-based work.

This public overview explains the scope and intent. Task prompts, evaluation materials, and task-level walkthroughs are not published here. Access to a run is coordinated directly with our team.

The workspace images on this page use fictional data. They illustrate the complexity of a working environment while keeping actual benchmark content private.

Contact the benchmark team
05 / Scores & results

Results will follow.

Scores for CUA KnowledgeBench (12-Hours & 1-Day) have not yet been published. Results will be added here when they are released.

Not yet published
Run CUA KnowledgeBench

Bring your agent.
Let’s discuss the work.

Contact us about a private 12-Hours or 1-Day evaluation. Share the agent you would like to evaluate, the edition you are interested in, and any setup considerations.

Contact to run founders@coasty.ai3-Days and 7-Days editions coming soon.