CUA KnowledgeBench (12-Hours & 1-Day)
Real computer-use capability across more than 100 applications. Built around reasoning, current information, and the way people and companies actually work.

One piece of work. Several connected tools.
An inbox, an operational spreadsheet, and a project board show how information and responsibilities can be distributed across a working environment.
Illustrative workspaces · AI-generated images with fictional data. Actual benchmark environments and materials remain private.
A longer view of capability.
Two editions for private evaluation today. Two longer editions on the horizon.
12-Hours
CUA KnowledgeBench (12-Hours)
A 12-hour private evaluation of reasoning and computer use in realistic, connected work environments.
Discuss 12-Hours1-Day
CUA KnowledgeBench (1-Day)
A full-day (24-hour) private evaluation across applications, evolving information, and practical work.
Discuss 1-DayContact us to discuss the edition, execution setup, and timing for your agent.
The whole working context.
Over 100 applications across the benchmark, with the interconnected context of real-world work. The emphasis is on how tools, information, and human decisions fit together.
Software is only part of the environment.
The harder part is understanding what is happening inside it. People leave decisions in conversations, maintain information in different places, and depend on work completed by others. A realistic evaluation needs to reflect that surrounding context.
Information has a history—and a present.
Work draws on what was known before and what has changed since. Logic and current information matter together: finding a recent fact is useful only when the agent understands how it affects the work already in progress.
Communication
Conversations, inboxes, and updates carry intent, decisions, and changes in direction.
Documents & knowledge
Working documents and reference material hold the context that decisions depend on.
Spreadsheets & analysis
Structured information needs interpretation, reconciliation, and consistent use elsewhere.
Projects & operations
Work has owners, dependencies, status, and consequences for the next person in the process.
Customer & business context
Records connect people and organisations to the history and details of ongoing work.
Research & coordination
Current sources, schedules, and shared commitments connect information to action.
These are the kinds of workplace context the benchmark is designed to reflect. The 100+ application figure describes the benchmark as a whole; it does not mean every application is used in every evaluation. The detailed application roster and evaluation cases are private.
From operating software
to carrying work forward.
CUA KnowledgeBench is designed to examine practical computer-use capability in context. These are the questions that make that capability meaningful.
Reason through connected decisionsInterpret context, reconcile information, and understand what depends on what.
Useful computer work involves choosing a sensible course of action from incomplete or conflicting information. An agent needs to connect facts across sources, recognise constraints, and understand how an early decision affects later work. The reasoning has to remain consistent as it becomes actions in software.
Work with information that changesFind current information and recognise when earlier assumptions need revisiting.
Real-time information is part of the work itself. A source can change, an update can supersede an earlier message, and a previously reasonable assumption can become outdated. Capable agents need to find relevant information, consider its freshness, and use it in context rather than treating an earlier observation as permanently true.
Carry context across applicationsKeep the meaning of the work intact as it moves between tools.
Information takes different forms in different applications: a conversation, a document, a table, a calendar event, or a record. Moving between those interfaces requires recognising the same people, entities, dates, and decisions. The capability of interest is whether the agent can use those connections accurately while operating the computer.
Maintain continuity over timeKeep track of progress, open questions, and commitments as work develops.
Extended work puts pressure on consistency. An agent needs to distinguish completed work from pending work, retain relevant context, and avoid repeating or contradicting earlier actions. This is a question of continuity: whether decisions and actions remain coherent as the surrounding work develops.
Respond to uncertainty and changeRecognise a changed situation and make a considered next decision.
The way people work includes interruptions, ambiguity, missing information, and competing priorities. A capable agent needs to notice when the situation differs from its assumptions, reassess what is known, and determine what can reasonably happen next. Continuing an outdated plan can be as consequential as choosing the wrong action initially.
Produce work that people can useBring accuracy, consistency, and context into the resulting work.
The purpose of computer use is a useful outcome. Documents, records, analyses, and handoffs need to make sense to the people who rely on them. This means attention to the information behind the work, consistency between related artifacts, and clarity about what has been established and what remains unresolved.
These descriptions explain the evaluation focus. They are not a published scoring rubric or a set of benchmark tasks.
The benchmark stays private.
The ambition is clear.
CUA KnowledgeBench is one of our most closely held benchmarks. Its purpose is to measure real capability through realistic computer-based work.
This public overview explains the scope and intent. Task prompts, evaluation materials, and task-level walkthroughs are not published here. Access to a run is coordinated directly with our team.
The workspace images on this page use fictional data. They illustrate the complexity of a working environment while keeping actual benchmark content private.
Contact the benchmark teamResults will follow.
Scores for CUA KnowledgeBench (12-Hours & 1-Day) have not yet been published. Results will be added here when they are released.
Bring your agent.
Let’s discuss the work.
Contact us about a private 12-Hours or 1-Day evaluation. Share the agent you would like to evaluate, the edition you are interested in, and any setup considerations.