ERPBENCH: A STATE-GROUNDED EVALUATION PARADIGM FOR COMPUTER-USE AGENTS IN ENTERPRISE SOFTWARE Kratika Bhagtani∗ , Kusha Sridhar∗ , Maziyar Baran Pouyan, Yuying Zhao, Eugene Siow Agentic AI Center of Excellence, Accenture {kratika.bhagtani, k.sridhara.murthy, maziyar.baran.pouyan, yuying.d.zhao, eugene.siow}@accenture.com arXiv:2609.17885v2 [cs.AI] 23 Sep 2026 ABSTRACT sarily translate to enterprise software. Business systems such Computer-use agents that operate through screenshots and as Enterprise Resource Planning (ERP) and Customer Re- simulated actions are advancing rapidly, yet their evaluation lationship Management (CRM), spanning finance, procure- remains anchored to general desktop and web tasks. Enter- ment, and inventory, differ fundamentally from consumer ap- prise Resource Planning systems run the finance, procure- plications in both workflow structure and failure cost. Here ment, inventory, and customer operations of organizations a visually plausible interaction is not enough: an agent may worldwide, and pose distinct challenges for computer-use navigate to the correct screen, click the expected controls, and agents: dense interfaces, coordinated multi-step interactions, observe a success confirmation, yet leave the business record and errors that alter persistent business records rather than unchanged or wrong. The confirmation reflects that the inter- surfacing on screen. Existing enterprise computer-use bench- face accepted an action, not that the intended value reached marks rely on proprietary platforms or on simulated ap- the database: the value entered on screen may never have proximations of such software. We introduce ERPBench, a been bound to the field it appears to fill, so what persists is benchmark that evaluates screenshot-only agents on a live stale, empty, or wrong. Such failures can be invisible from and reproducible system and scores each task against ground- screenshots alone, yet they propagate into downstream op- truth values in its database. Beyond the benchmark, we erational and financial processes. Existing GUI benchmarks present a production-grade harness that gates agent actions primarily target general desktop and web environments [5, 6, behind human approval for safe deployment. Evaluating six 7], while recent enterprise benchmarks use proprietary plat- closed and open-source agents, we demonstrate that strong forms or simulated enterprise environments [8, 9, 10]. Nei- general performance does not transfer to enterprise reliability. ther decomposes a run into stages, so a failure to reach the Even when an agent reaches the right form and saves it, the form is scored like one where the agent saved and still stored stored record is often wrong: some agents save in up to 85% the wrong value. These limitations motivate an evaluation of runs but write the correct value in as few as 3%. We further paradigm that measures business state and localizes failures characterize failure modes specific to enterprise workflows. within a workflow. Index Terms— Computer-use agents, agent safety, enter- We call this paradigm state-grounded evaluation, and prise software, human-in-the-loop, ERP instantiate it in ERPBench. ERPBench is built on a live, self-hosted instance of ERPNext [11], an open-source ERP 1. INTRODUCTION spanning finance, procurement, inventory, and customer man- agement. Agents operate from screenshots alone, and each Advances in multimodal large language models (LLMs) have task is graded against ground-truth database fields, extending enabled computer-use agents (CUAs) that operate graphical execution-based state verification [5, 7] to business records. user interfaces (GUIs) directly, through screenshots and sim- We impose the screenshot-only constraint because enterprise ulated mouse and keyboard actions [1, 2, 3, 4]. On gen- software is routinely reached through remote-desktop and eral desktop and web benchmarks such as OSWorld [5], Win- Virtual Desktop Infrastructure (VDI) sessions that expose dows Agent Arena (WAA) [6], and WebArena [7], CUAs no Document Object Model (DOM), accessibility, or Appli- have improved rapidly (an example of a typical task is editing cation Programming Interface (API) hooks: pixel input is a spreadsheet in LibreOffice), suggesting that general GUI in- the only channel that generalizes across such deployments. teraction is becoming tractable. This progress does not neces- Field-level grading exposes enterprise-specific failure modes © 2026 Accenture. All rights reserved. that metrics scored at the task level can conflate, including ∗ These authors contributed equally. agents trapped in unrecoverable interaction loops. Table 1. Comparison with existing CUA benchmarks. ITSM = IT Service Management; CRM = Customer Relationship Man- agement; ERP = Enterprise Resource Planning; A11y = Accessibility tree; DOM = Document Object Model; API = Application Programming Interface. ∼ = partial (SCUBA: difficulty labels, milestone rewards; not structural tiers or fixed-stage grading). Open / reproducible = the environment is self-hostable without a proprietary account or license. A ✓ under Needs structured access is a limitation, not a feature: such inputs ease grounding but assume A11y/DOM/API hooks that legacy, pixel-only systems do not expose. No prior benchmark satisfies every row; ERPBench is the only one to combine all of them. Property OSWorld WAA WorkArena SCUBA CRMArena- UI-CUBE ERPBench Pro (ours) Domain General General Enterprise Enterprise Enterprise Enterprise Enterprise Subdomain OS OS ITSM CRM CRM Mixed ERP Platform Ubuntu Windows ServiceNow Salesforce Salesforce Web mocks ERPNext Distinct tasks 369 154 33 60 19 226 30 (150 runs) Pixel-only input × × × ✓ × × ✓ Needs structured access ✓(A11y) ✓(A11y) ✓(A11y+DOM) × ✓(Text/API) ✓(DOM) × Live, non-simulated app ✓ ✓ ✓ ✓ ✓ × ✓ Open / reproducible ✓ × × × × ✓ ✓ Database-verified ground truth × × ✓ ✓ ✓ × ✓ Human baseline ✓ ✓ × × × ✓ ✓ Task complexity tiers × × × ∼ × ✓ ✓ Stage-wise grading × × × ∼ × × ✓ Formal failure taxonomy × × × × × × ✓ ERPBench runs on a production-grade harness that we desktop and web applications. OSWorld [5] evaluates mul- present. In deployment, it gates irreversible agent actions timodal agents on open-ended tasks in real desktop environ- behind human approval for safe enterprise use [2]. For ments, while Windows Agent Arena [6] extends evaluation benchmarking, the same harness runs autonomously, with to Windows. WebArena [7] provides realistic web environ- an auto-approver in place of the human reviewer. We evalu- ments and evaluates task completion against the underlying ate six closed and open-source CUAs [2, 1, 4, 12, 13] on tasks environment state rather than the rendered page. Collectively, spanning single-field form edits, multi-field business record these benchmarks establish interaction and execution success creation, and multi-screen chained workflows. We demon- as the dominant evaluation paradigm for general-purpose GUI strate that high apparent GUI competence does not imply agents, but do not model the domain-specific business state enterprise reliability: agents that navigate successfully often characteristic of enterprise applications. ERPBench builds on fail at the interaction or commit stage, and even the strongest the execution-based verification principle of WebArena [7] proprietary model exhibits enterprise-specific failures that but extends it to relational business records verified at the existing methodologies do not capture. Our contributions database-field level. are threefold: (1) ERPBench, the first benchmark to opera- Enterprise Computer-Use Evaluation: Several recent tionalize pixel-only, state-grounded evaluation for enterprise benchmarks extend computer-use evaluation to enterprise CUAs on a live, openly reproducible enterprise system, grad- workflows. WorkArena [14] evaluates web agents on knowl- ing screenshot-only agents at the database-field level on an edge work tasks in ServiceNow, while Salesforce Computer open-source ERP [11]. (2) A production-grade harness Use Benchmark (SCUBA) [8] evaluates computer-use agents with human-in-the-loop safety gating, run autonomously to on Salesforce CRM workflows using application APIs for benchmark the same system deployed in production. (3) binary and milestone-based grading. CRMArena-Pro [9] The first systematic analysis of enterprise-specific CUA evaluates LLM agents on diverse CRM workflows, but uses failures, showing that interaction-level evaluation overstates text-based interaction rather than screenshot-driven GUI con- enterprise reliability, and localizing the stage at which agents trol. EntWorld [15] evaluates enterprise GUI agents across break down. six dockerized open-source business applications with SQL- based deterministic verification and schema-driven task gen- 2. RELATED WORK eration. However, EntWorld exposes an accessibility tree alongside screenshots and does not decompose runs into fixed Computer-use agent evaluation has developed along four stages or define a named failure taxonomy. UI-CUBE [10] threads: general GUI benchmarks, enterprise computer-use is particularly close to our work, arguing, as we do, that evaluation, agent safety, and computer-use agent systems. task-level accuracy alone is insufficient to assess enterprise General GUI benchmarks: Recent benchmarks have estab- readiness and introducing operational-reliability evaluation lished standardized environments for evaluating CUAs across across enterprise workflows. However, UI-CUBE evaluates agents on simulated web applications, exposes DOM infor- (A) Production CUA Harness (B) ERPBench Evaluation Mode mation alongside screenshots, and verifies outcomes against ERPNext (Docker) Task YAML Fixture ERPNext (Docker) MariaDB · Redis · noVNC MariaDB · Redis · noVNC application-state snapshots. ERPBench instead evaluates Screenshot prompt fixture Screenshot expected screenshot-only agents on a live, self-hosted ERP deploy- CUA Brain CUA Brain Claude · Qwen · OpenCUA · UI-TARS Claude · Qwen · OpenCUA · UI-TARS ment and grounds correctness in database fields. Moreover, Action Proposal Action Proposal ERPBench decomposes execution into fixed stages and intro- Flask Observer (phase FSM) Flask Observer (phase FSM) idle → task active → proposed → executing idle → task active → proposed → executing duces a named failure taxonomy to localize where enterprise agents fail. Risk Classifier Auto Approver safe / commit / irreversible Agent Safety and Human Oversight: Deploying CUAs in Irreversible (200 ms poll) deflects WAITING_HUMAN enterprise systems also raises safety concerns because GUI Human Reviewer Approved Action approve / actions may modify or irreversibly commit business records. reject / DB Grader Partial Grader escalate (Frappe API) (4-stage) Prior work has studied human-in-the-loop (HITL) approval, Approved Action Log Grader automated intent verification, and agent safety awareness as (browser execution) mechanisms for controlling risky actions [16, 17, 18, 19]. These works primarily evaluate safety mechanisms or agent- side risk awareness. ERPBench instead incorporates a risk- Fig. 1. ERPBench system overview that compares the tiered approval gate into a production-oriented execution har- two harness configurations: (A) Production CUA harness. ness and runs the same infrastructure autonomously during Demonstrates how proposed agent actions from different benchmarking, enabling enterprise reliability to be evaluated CUA brains transition through a Flask Finite State Machine under the controls intended for deployment. (FSM) observer into a Risk Classifier. Risky or irreversible Computer-Use Agent Systems: Recent systems including actions (such as deleting records) require explicit human ap- Claude [2] and OpenAI’s [3] computer-use capability, Open- proval before browser execution. (B) Evaluation mode. Re- CUA [4], UI-TARS [1], Holo3-35B-A3B [13], and Qwen3- places the human reviewer with an Auto Approver polling at VL [12] paired with OmniParser [20] demonstrate different 200 ms intervals. Post-execution state is processed by the DB approaches to screenshot-based computer interaction, ranging Grader, Partial (Stage) Grader, and Log Grader. from native computer-use capabilities and GUI-specialized training to explicit visual element parsing [2, 3, 4, 1, 12, 20, 13]. ERPBench evaluates representative systems from these DOM or accessibility tree is exposed for perception or for approaches as benchmark subjects; its contribution is com- control. This pixel-in, coordinate-out channel mirrors how plementary to these model-development efforts, focusing on an agent must operate enterprise software behind remote- how reliably such agents perform enterprise operations rather desktop or VDI deployments, where no API or accessibility than on improving GUI grounding itself. hooks are available, the setting ERPBench is built to evaluate. The rest of this section describes ERPBench’s four compo- Table 1 positions ERPBench against existing benchmarks nents: the task suite (Sec. 3.1), the harness that runs agents along observation modality, deployment, grounding, and against ERPNext (Sec. 3.2), the graders that score each run grading. Of the four threads above, the benchmarks compared (Sec. 3.3), and a human baseline (Sec. 3.4). here fall into two: the general GUI thread (OSWorld [5], Windows Agent Arena [6]) and the enterprise computer-use evaluation thread (WorkArena [14], SCUBA [8], CRMArena- 3.1. Task Design Pro [9], UI-CUBE [10]); the agent safety and computer-use ERPBench organizes tasks into three complexity tiers (T1– agent systems threads comprise safety studies and agent mod- T3). T1 covers 20 single-record tasks in three types: text-edit els rather than benchmarks. ERPBench advances the enter- (typing into a text or long-text field, e.g., a customer’s tax prise computer-use evaluation thread. ERPBench is the only ID), select/toggle (choosing via dropdown, autocomplete, benchmark that satisfies every row; each prior benchmark checkbox, or date picker, e.g., a disabled flag), and create satisfies only a subset. (adding a simple record with a few fields, e.g., a new cus- tomer). T2 introduces multi-field coordination and relational field selection, with 6-9 fields per record (e.g., creating a new 3. ERPBENCH Customer or Supplier from its full field set). T3 introduces ERPBench evaluates CUAs on ERPNext, an open-source multi-screen chained workflows: the agent creates and links ERP deployed locally in Docker (Fig. 1). The agent drives 2–4 dependent ERPNext documents in sequence (e.g., Lead ERPNext through a browser as a person would: Chromium → Opportunity → status advance, or Project → Task → renders the application on a virtual display (Xvfb), streamed Timesheet), and one variant additionally submits a document, over noVNC. At each step the agent observes only a 1280 × ERPNext’s stronger and often irreversible commit, rather than 720-pixel screenshot of this display and acts through sim- only saving it. Three T3 tasks are OSWorld-analog variants ulated mouse and keyboard input at pixel coordinates. No that start from a blank page, requiring the agent to navigate human. Before each warm-start run, the browser is initialized at a task-specific start URL to remove URL-level navigation as a confound. Navigation within the ERP interface, from the start state to the target record or field, remains part of the agent’s task and is scored as the first grading stage. Each run has a turn limit; a run that hits the limit is recorded as a re- covery failure (see Sec. 4.3 for additional details). 3.3. Grading Each run is scored by three graders: Database grader, Stage grader, and Log grader. The database grader reads the tar- get field values from ERPNext through the Frappe REST API Fig. 2. ERPBench harness User Interface: Shows the live after the run. For a task with target fields F , let yi∗ and ŷi de- operator (human) interface divided into four functional pan- note the expected and observed post-execution values of field ⃝ Live ERPNext App: The streaming screenshot-only els: 1 i ∈ F . A per-field predicate g(ŷi , yi∗ ) ∈ {0, 1} marks field i workspace. ⃝ 2 Planned Workflow Action Log: Displays correct under type-specific equivalence: numeric values agree planned step sequences. ⃝ 3 Risk-Gated Approval Panel: within a tolerance τ (default 0.01); Boolean-equivalent val- Pauses execution on sensitive actions (e.g., irreversible dele- ues agree (e.g., 1 and True, or “Yes” and True); and strings tion) pending operator review. ⃝ 4 Chat Panel: Provides a di- are compared after HTML stripping and whitespace normal- rect communication for operator intervention. ization. A task succeeds iff every target field is correct, i.e., ∗ Q i∈F g(ŷi , yi ) = 1. to ERPNext itself before any form work. Because the harness For multi-field T2 records, credit is field-weighted across and graders are specialized to ERPNext, running them un- the target fields rather than all-or-nothing. For multi-document changed on another benchmark is not straightforward; these T3 tasks, the grader is applied per document in the chain, and variants instead reproduce the operating conditions of general partial credit is the chain-depth: the fraction of chain docu- GUI benchmarks inside ERPBench, providing a comparison ments left correct in the database. The stage grader awards point without a harness port. Together, the tiers progres- partial credit over ordered stages, agent actions followed sively increase interaction complexity, from a single field by a verification of state: Navigation (the target record or (T1) to a full multi-field record (T2) to a chain of linked field is reached), Interaction (the required GUI manipulation records across screens (T3), while preserving a common cor- is completed), Commit (a save or submit is invoked), and rectness criterion: the final database state. Each task is a Database (the target fields match ground truth). Navigation, YAML file specifying the prompt, a start URL, a fixture of Commit, and Database apply to every tier; Interaction is spe- pre-seeded fields, the expected field values, and grader con- cific to single-field edits, and chained workflows substitute figuration. Before each run, the fixture is seeded through the chain-depth. Frappe REST API under a UUID-suffixed record name (e.g., Each stage is evaluated independently from observable Acme-Corp-RUN-edba5c7f). This isolates runs without evidence, while the final database state remains the definitive a full database reset. criterion for task success. This localizes where a run failed, not just whether it failed. The log grader reads the action 3.2. Harness and turn logs for action count, agent/API turns, latency, and, At each turn, the agent sees the current screenshot and pro- where available, token cost. Together, the three graders give poses one action: move, scroll, click (field, button, link, or each run a pass/fail result, a per-stage breakdown, and its cost. dropdown), type, select, set date, clear, press a key, or nav- igate. Each action carries target coordinates, a short reason, and a risk category (safe, commit, or irreversible). If it is 3.4. Human Baseline approved, the harness runs it as a real mouse or keyboard ac- As a human reference, three annotators independently com- tion and takes a new screenshot. The agent then proposes pleted all ERPBench tasks on the same ERPNext instance the next action. The same harness runs in two modes. A and were scored by the same database grader as the agents. Flask server tracks the task’s state (idle → active → proposed Because humans operate the browser directly rather than → executing). In deployment, actions tagged commit or ir- through the screenshot-only agent interface, we report their reversible (e.g., the save/submit bar) are surfaced for human task success and completion time and omit agent-specific approval before execution, while actions tagged safe execute metrics such as token usage. The reference deliberately spans directly (Fig. 2). In evaluation, an auto-approver approves ev- different levels of platform familiarity: human annotators #1 ery action across all risk tiers, so a full task runs without a and #2 are experienced ERPNext users and serve as an expert ceiling, while human annotator #3 provides first-time-user Verified (77.8 against 72.5) yet reaches 34% on T1 where reference, having never used ERPNext before. Claude reaches 94%. Table 3 localizes where this competence leaks away. 4. EXPERIMENTS AND RESULTS Open models successfully navigate to target records in 90– 95% of T1 runs, but fail at interaction and database per- We first describe the experimental setup, then report task suc- sistence. Their interaction rate falls to 0–53% and their cess and efficiency, and finally analyze how and where the database-correct rate to 0–34%: they arrive at the right screen agents fail. yet cannot reliably enter and persist the intended value. UI- TARS is the clearest case: it reaches the target in 95% of runs 4.1. Experimental Setup and fires a save in 68%, yet only 9% leave the correct value in the database, so nearly every save it commits is wrong. We evaluate six representative CUAs spanning proprietary The pattern sharpens with tier depth: on T2 the open mod- and open-weight systems: Claude Sonnet 4.6, Qwen3-VL- els reach the record in every run yet are database-correct in 32B, OpenCUA-32B, UI-TARS-1.5-7B, OpenCUA-7B, and at most 5%, with UI-TARS committing in 85% of runs but Holo3-35B-A3B. Claude is accessed through the Anthropic database-correct in 3%, and on T3 the open models complete API, while the open-weight models are self-hosted; Qwen3- no chain end to end. Claude, by contrast, stays near ceiling VL-32B is paired with OmniParser for GUI grounding. Each across every stage (T1 navigate 99%, interact 97%, commit task is executed for five independent runs under the same con- 98%, database 94%). The open models’ collapse with tier figuration and grader, yielding 100 T1, 20 T2, and 30 T3 depth mirrors the capability cliff UI-CUBE [10] reports as runs per model. Three human annotators complete the same enterprise tasks grow more complex, and the broader enter- tasks through the same database grader as a human refer- prise gap EntWorld [15] observes. ERPBench’s stage-wise, ence. On OSWorld-Verified these agents report 77.8 (Holo3- database-level grading provides a finer resolution on this gap 35B-A3B), 72.5 (Claude Sonnet 4.6), 34.8 (OpenCUA-32B), by isolating runs where an agent navigates and even commits 27.4 (UI-TARS-1.5-7B), and 26.6 (OpenCUA-7B); Qwen3- yet leaves the wrong value in the record, a silent failure that VL-32B provides no native computer-use agent and is paired screenshot-level or interaction-level grading would misclas- with OmniParser, so no OSWorld score applies. sify as success. Table 2 reports task success and efficiency across the three Table 2 also shows the cost of the longer horizon. Claude tiers; Table 3 decomposes each run into ordered stages and completes all six T3 chained workflows (30/30), matching the illustrates where agents fail during execution. Success is the best human annotator, but at 36 actions and 323 s per run ver- share of runs in which every target field matches the database; sus 12 actions and 89 s on T1. The open models stay cheap Actions, Duration, and Input tokens are per-run means. Ac- only because they stop early: OpenCUA-7B averages 1.6 ac- tions counts executed GUI operations, while Turns (Table 4) tions on T1 and Qwen3-VL 4.9, reflecting premature loop ter- counts agent-model calls and exceeds Actions when a turn mination rather than task efficiency. yields no executed operation (e.g., a rejected or terminal step). The stage columns report the share of runs reaching each Table 4 reports the remaining per-run metrics; they track stage: Navigation (the target record or field is reached), In- the same pattern, Claude’s turn counts and output tokens grow teraction (the correct value is entered; T1 only), Commit (a with tier depth (13.8 turns in T1 to 38.2 turns in T3, 2.3K save or submit is invoked), Database (the target fields match output tokens in T1 to 6.5K output tokens in T3) while its ground truth; field-weighted for T2), and, for T3, Chain-depth per-turn thinking time grows more modestly, from about 4 s (the mean fraction of the document chain left correct). Human on T1–T2 to 6 s on T3. Open models exhibit low average rows are a human reference ceiling and omit token metrics, turns and low output token counts primarily because they hit which are agent-specific. loop-breaks or terminate early. Table 5 shows partial credit for all tiers and analyzes performance across specific sub- task types. With the exception of Holo3, open models per- 4.2. Results form worse on exact free-text field entry (22–47%) than on Strong general GUI performance does not translate into re- select/toggle choices (31–54%), i.e., typing an exact value is liable enterprise execution. Claude reaches 94%, 100%, and harder than picking a known option. In T3, removing the nav- 100% success across the three tiers, approaching the human igation crutch (starting URL prompts) causes chain comple- reference; however, its token requirements scale heavily, tion depth to drop significantly. For example, OpenCUA-32B from 231K tokens per run on T1 to roughly 2.0M on T3. drops from 22% warm-start depth to 9% blank-start depth. By contrast, the strongest open models, Holo3-35B-A3B For human annotators, familiarity has only a modest ef- and Qwen3-VL-32B, reach only 34% and 32% on T1 and fect: the experts reach 99–100% on T1, 95% on T2, and 97– fall to 0–3% on T2 and T3, while OpenCUA-7B fails every 100% on T3, and the first-time user still reaches 96%, 90%, task (Table 2). General GUI scores do not predict enterprise and 87% respectively. Even this first-time user far exceeds ev- reliability: Holo3-35B-A3B outscores Claude on OSWorld- ery open model (which top out at 34% on T1 and 3% on T3), Table 2. ERPBench Benchmark results: Task Success and Efficiency Across Tiers. Demonstrates success rate (%), mean GUI actions per run, mean duration per run (seconds), and mean input token usage per run across models. Human rows are a human reference ceiling. T1 T2 T3 Model Success Actions Duration (s) Input tokens (K) Success Actions Duration (s) Input tokens (K) Success Actions Duration (s) Input tokens (K) Proprietary Claude Sonnet 4.6 94 11.7 89.2 231 100 25.6 170.1 784 100 36.2 322.7 2009 Open-weight Holo3-35B-A3B 34 6.5 37.1 34 0 14.2 58.6 93 3 6.8 56.0 40 Qwen3-VL-32B 32 4.9 29.1 76 0 7.7 39.4 107 3 10.5 61.9 149 OpenCUA-32B 23 5.5 52.8 105 0 6.7 57.8 101 0 7.3 64.9 108 UI-TARS-1.5-7B 9 13.1 59.2 51 0 26.2 99.0 103 0 25.3 111.4 107 OpenCUA-7B 0 1.6 26.5 60 0 1.4 28.3 75 0 11.8 84.2 145 Human reference Human #1 100 7.5 41.4 N/A 95 25.1 85.7 N/A 100 27.1 99.9 N/A Human #2 99 14.8 89.1 N/A 95 26.8 150.5 N/A 97 40.6 395.5 N/A Human #3 96 8.9 42.3 N/A 90 24.6 93.2 N/A 87 28.7 100.7 N/A Table 3. Stage-wise localization within ERPBench illustrates Table 4. Additional efficiency metrics across tiers. Demon- where agents fail during execution. Demonstrates percent of strates agent/API turn counts, generated output tokens, and runs reaching each stage (Chain-depth as mean percent of the per-turn median thinking times (seconds). document chain completed). Interaction is a T1-only stage T1 T2 T3 (right value entered); Chain-depth is T3-only. Agents reach Output tokens Thinking (s) Output tokens Thinking (s) Output tokens Thinking (s) the page but fail to persist correct database state. For T2, Database is field-weighted across the target fields, so it can Turns Turns Turns exceed the all-or-nothing Success in Table 2. Model T1 T2 T3 Proprietary Claude Sonnet 4.6 13.8 2,255 4.320 27.6 4,455 4.400 38.2 6,468 6.217 Navigation Interaction Navigation Navigation Chain-depth Open-weight Commit Database Commit Database Commit Database Holo3-35B-A3B 8.5 1,634 2.095 15.2 2,866 2.000 9.7 2,054 2.067 Qwen3-VL-32B 8.4 543 1.212 10.9 687 1.536 14.5 1,068 1.554 Model OpenCUA-32B 11.9 751 2.660 11.1 684 3.009 11.4 775 2.527 Proprietary UI-TARS-1.5-7B 15.8 1,503 1.875 30.0 2,871 1.850 29.4 2,825 1.817 Claude Sonnet 4.6 99 97 98 94 100 100 100 100 100 100 100 OpenCUA-7B 6.9 495 1.471 8.3 548 1.492 14.8 1,072 1.623 Open-weight Holo3-35B-A3B 90 53 47 34 100 5 0 33 17 4 3 Qwen3-VL-32B 90 30 42 32 100 15 5 67 23 6 3 OpenCUA-32B 92 23 34 23 100 0 0 67 47 16 0 the target value correctly but fail to trigger save. Perception UI-TARS-1.5-7B 95 30 68 9 100 85 3 60 80 3 0 failures save successfully, but an incorrect value is written OpenCUA-7B 90 0 3 0 100 10 0 33 3 0 0 to the persistent database (silent failure). Recovery failures Human reference get trapped in interaction loops or exhaust the turn budget Human #1 100 100 100 100 100 100 95 100 100 100 100 Human #2 100 97 100 99 100 100 99 100 100 98 97 without making progress. Human #3 100 93 98 96 100 100 94 100 100 94 87 Open-weight models fail predominantly post-navigation rather than in planning. For instance, of Qwen3-VL-32B’s 68 failed T1 runs, 40 are Grounding, 10 Planning, 10 Perception, leaving Claude as the only agent that performs in the human and 8 Save-step. Of UI-TARS-1.5-7B’s 91 failures, 54 are range; ERPBench’s difficulty for agents is therefore not only Perception and 25 Grounding. a matter of ERP-specific expertise. Across the 24 T1 and T2 Each failed run is assigned to a single mode using the tasks, the three annotators reached identical outcomes on 16 priority order Recovery, Planning, Perception, Save-step, (67%) and agreed to within a single run on 23 (96%), with the then Grounding. Recovery is checked first because a run that first-time user accounting for most of the small differences. loops or times out can do so at any stage. Planning is next (never reaching the record), followed by the two silent-failure 4.3. Failure Analysis checks, Perception (a wrong value was saved) and Save-step Figure 3 indicates the failure taxonomy over failed T1 runs. (a correct value was never saved), leaving Grounding as the It categorizes non-successful T1 runs into five hierarchi- residual category for runs that reach the record but act on the cal modes. Planning failures never reach the target record. wrong element. Grounding failures reach the record but act on the wrong el- Two of these modes are exactly why enterprise evalua- ement, so the value is never entered. Save-step failures enter tion needs database-level grading. High visual competence it breaks. Table 5. Partial credit across tiers by sub-task type (%). Text-edit and Select/Toggle (T1) and Multi-field (T2) report 5. CONCLUSION the mean partial score, the fraction of that tier’s grading stages passed, averaged over tasks and runs. Create (T1) re- We introduced ERPBench, a benchmark for screenshot-only ports the pass rate for simple single-record creates (graded computer-use agents on a real, open, self-hosted ERP. Tasks pass/fail). T3 reports the mean chain-depth (fraction of the are graded against the database state the agents leave behind, document chain completed), split into warm-start and blank- and stage-wise evaluation identifies where failures occur start (OSWorld-analog) variants. across navigation, interaction, commit, and database correct- T1 T2 T3 ness. We also contribute the production-grade harness used Text- Select/ Multi- Warm- Blank- to run the agents, which gates actions behind human approval Model edit Toggle Create field start start in deployment, and runs the same harness autonomously for Claude Sonnet 4.6 96 97 100 100 100 100 evaluation. Holo3-35B-A3B 61 58 15 35 9 0 Qwen3-VL-32B 43 49 70 40 11 2 Across six agents, strong general GUI performance did OpenCUA-32B 37 49 30 33 22 9 not carry over to enterprise work. Models that navigate UI-TARS-1.5-7B 47 54 35 63 7 0 OpenCUA-7B 22 31 0 37 0 0 well still fail at interaction and commit, and some appear to succeed on screen while leaving the database wrong or unchanged. Interaction-level evaluation can therefore over- on screen does not correlate directly with database accuracy. estimate enterprise reliability; database-grounded evaluation In Save-step and Perception failures the agent acts, the screen reveals failures that remain invisible at the GUI level. ERP- shows a successful save, yet the persistent record is wrong or Bench currently covers single-field edits, multi-field record unchanged. Evaluating computer-use agents solely through creation and multi-screen chained workflows. We intend to screenshot comparisons or execution traces hides silent enter- release the benchmark and its evaluation artifacts. prise commit failures, underscoring the necessity of database- Because these failures localize to interaction, commit, and grounded verification. These silent failures do not surface persistence, they point targeted improvement, such as fine- on screen but propagate into the business records that down- tuning on state-grounded trajectories, at those stages, while stream processes depend on, which is precisely what state- the human-in-the-loop harness makes pixel-only agents de- grounded evaluation is designed to catch. ployable in the meantime. Claude Holo3-35B- Qwen3-VL- Sonnet 4.6 A3B 32B 6. ACKNOWLEDGMENTS 26% 15% 15% The authors are grateful to Eduardo Salamanca de Diego and 33% 11% 12% Jordan Ackerman for their guidance and continued support 67% 61% 59% of this work, and to Christopher Clarke, Shafiuddin Rehan 3% Ahmed, and Teja Kanchinadam for their generous assistance 6 failed / 100 66 failed / 100 68 failed / 100 with model deployment and debugging. The authors also OpenCUA- UI-TARS- OpenCUA- thank Shanka Subhra Mondal and Dhaval Potdar for their 32B 7B 7B helpful review of the manuscript. The authors used Claude (Anthropic) for language refine- 14% 27% 12% 10% 5% 10% ment and structural organization of this paper. Generative AI 7% 5% 78% tools were not used to generate experiments, perform data 57% 6% 59% 3% 2% 2% analysis, or produce scientific claims. All methodologies, re- sults, and conclusions were developed independently by the 77 failed / 100 91 failed / 100 100 failed / 100 authors. Grounding Planning Perception Save-step Recovery 7. COMPLIANCE WITH ETHICAL STANDARDS Fig. 3. Failure taxonomy over failed T1 runs, by model. Most failures occur after successful navigation, clustering in The human reference was produced by the authors them- grounding, save-step, and perception rather than planning: selves; no external participants and no sensitive personal data reaching the right screen does not imply completing the task were involved, and no institutional review board approval correctly. was required. Stage-wise grading and the failure taxonomy are comple- mentary lenses: the stages identify where in the execution pipeline a run breaks, while the taxonomy characterizes how 8. REFERENCES [10] Horia Cristescu, Charles Park, Trong Canh Nguyen, Sergiu Talmacel, Alexandru-Gabriel Ilie, and Stefan [1] Yujia Qin et al., “UI-TARS: Pioneering Automated GUI Adam, “UI-CUBE: Enterprise-Grade Computer Use Interaction with Native Agents,” arXiv:2501.12326, Agent Benchmarking Beyond Task Accuracy to Opera- January 2025. tional Reliability,” arXiv:2511.17131, November 2025. [2] Claude Platform Docs, Anthropic, “Computer [11] Frappe, “ERPNext: Free and Open Source Enter- use tool,” https://platform.claude.com/ prise Resource Planning,” https://github.com/ docs/en/agents-and-tools/tool-use/ frappe/erpnext, 2026. computer-use-tool, November 2024. [12] Shuai Bai, Yuxuan Cai, Ruizhe Chen, et al., “Qwen3- [3] OpenAI Developers, OpenAI, “Computer use,” VL technical report,” arXiv:2511.21631, November https://developers.openai.com/api/ 2025. docs/guides/tools-computer-use, 2025. [13] H Company, “Holo3 - Open Foundation Mod- [4] Xinyuan Wang et al., “Opencua: Open foundations els for Navigation and Computer Use Agents,” for computer-use agents,” arXiv:2508.09123, October https://huggingface.co/Hcompany/ 2025. Holo3-35B-A3B, 2026. [14] Alexandre Drouin, Maxime Gasse, Massimo Caccia, Is- [5] Tianbao Xie et al., “OSWorld: Benchmarking Multi- sam H. Laradji, Manuel Del Verme, Tom Marty, David modal Agents for Open-Ended Tasks in Real Computer Vazquez, Nicolas Chapados, and Alexandre Lacoste, Environments,” Proceedings of the Advances in Neu- “WorkArena: How Capable are Web Agents at Solv- ral Information Processing Systems, vol. 37, pp. 52040– ing Common Knowledge Work Tasks?,” Proceedings of 52094, December 2024, Vancouver, Canada. the International Conference on Machine Learning, vol. [6] Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon 235, pp. 11642–11662, July 2024, Vienna, Austria. Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin [15] Ying Mo, Yu Bai, Dapeng Sun, Yuqian Shi, Yukai Miao, Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Li Chen, and Dan Li, “EntWorld: A Holistic Envi- Jang, and Zheng Hui, “Windows Agent Arena: Eval- ronment and Benchmark for Verifiable Enterprise GUI uating Multi-Modal OS Agents at Scale,” Proceedings Agents,” arXiv:2601.17722, January 2026. of the International Conference on Machine Learning, pp. 4874–4910, July 2025, Vancouver, Canada. [16] Emre Turan, “Oversight Has a Capacity: Calibrat- ing Agent Guards to a Subjective, Fatiguing Human,” [7] Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, arXiv:2606.08919, June 2026. Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Gra- [17] Peiran Wang, Ying Li, and Yuan Tian, “Reframing LLM ham Neubig, “WebArena: A Realistic Web Environ- Agent Security as an Agent-Human Interaction Prob- ment for Building Autonomous Agents,” Proceedings lem,” arXiv:2605.24309, May 2026. of the International Conference on Learning Represen- [18] Tianyu Chen, Chujia Hu, Ge Gao, Dongrui Liu, Xia tations, pp. 15585–15606, May 2024, Vienna, Austria. Hu, and Wenjie Wang, “LPS-Bench: Benchmarking [8] Yutong Dai, Krithika Ramakrishnan, Jing Gu, Matthew Safety Awareness of Computer-Use Agents in Long- Fernandez, Yanqi Luo, Viraj Prabhu, Zhenyu Hu, Silvio Horizon Planning under Benign and Adversarial Scenar- Savarese, Caiming Xiong, Zeyuan Chen, and Ran Xu, ios,” arXiv:2602.03255, February 2026. “SCUBA: Salesforce Computer Use Benchmark,” Pro- [19] Jia-Chen Zhang, Ze-Yu Zhang, and Kai-Wei Zhang, ceedings of the International Conference on Learning “Invisible Ink Threats: Adversarial Goals Be- Representations, pp. 24363–24386, April 2026, Rio de hind Legitimate Tasks in Computer-Use Agents,” Janeiro, Brazil. arXiv:2608.02018, August 2026. [9] Kung-Hsiang Huang, Akshara Prabhakar, Onkar Tho- [20] Yadong Lu, Jianwei Yang, Yelong Shen, and Ahmed rat, Divyansh Agarwal, Prafulla Kumar Choubey, Yixin Awadallah, “OmniParser for Pure Vision Based GUI Mao, Silvio Savarese, Caiming Xiong, and Chien-Sheng Agent,” arXiv:2408.00203, August 2024. Wu, “CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Inter- actions,” Transactions on Machine Learning Research, p. 35, January 2026.