OSWorld 2.0 Benchmark Exposes How Far Computer-Use AI Still Has to Go

The Core · TL;DR
- OSWorld 2.0 introduces 108 long-horizon computer-use tasks that take humans a median of 1.6 hours, far exceeding the original benchmark's difficulty.
- Claude Opus 4.8 tops the leaderboard with just 20.6% full task completion (54.8% partial score) at 500 steps; GPT-5.5 is more token-efficient but plateaus near 13%.
- Tasks now require an average of 318 tool calls versus roughly 30 in OSWorld 1.0, reflecting deliberately harder streaming, dynamic-environment, and cross-source reasoning challenges.
- The 68-page paper, from lead authors Mengqi Yuan, Zilong Zhou, and Xinzhuang Xiong, underscores how far current AI agents remain from reliably automating real-world multi-step computer work.
Twenty percent. That is the best completion rate any frontier model manages on OSWorld 2.0, a new benchmark designed to test whether AI agents can actually finish the kind of long, messy computer tasks that real people handle every day.
The benchmark, described in a paper posted to arXiv on June 28, 2026 and revised on July 13, comes from researchers Mengqi Yuan, Zilong Zhou, and Xinzhuang Xiong, who share equal contribution across the 68-page study. It builds on the original OSWorld framework by dramatically raising the bar on task complexity. Where the first version asked agents to complete relatively short digital chores, OSWorld 2.0 packages 108 long-horizon workflows spanning both everyday and professional computer use, each one modeled on tasks that take a human a median of roughly 1.6 hours to finish.
That jump in difficulty shows up starkly in the numbers. Agents tackling OSWorld 2.0 tasks need an average of 318 tool calls when running on Claude Opus 4.7 with maximum thinking enabled, versus about 30 tool calls for equivalent tasks in the original OSWorld. The tenfold increase reflects the benchmark's deliberate design around five specific failure points for current agents: streaming interaction, environments that change while the agent works, reasoning across multiple information sources, inferring state that isn't explicitly shown on screen, and fine-grained visual-spatial precision when clicking or navigating interfaces.
Frontier Models Still Fall Short
On the benchmark's primary binary-completion metric, the strongest configuration tested, Claude Opus 4.8 running with maximum thinking and batched tool calls over 500 steps, completes only 20.6% of tasks outright, though it earns partial credit on 54.8% of them under OSWorld 2.0's scoring for incremental progress.
GPT-5.5 tells a different story about efficiency versus raw capability. The model uses fewer tokens per task than Claude Opus 4.8, but that efficiency comes at a cost: its completion rate plateaus around 13%, meaningfully behind Anthropic's top performer despite requiring less compute per attempt.
The gap between these two results matters for anyone evaluating agent products for production use. A model that is cheaper to run but caps out at 13% completion may still lose on total cost when factoring in retries and human intervention, while a costlier model with higher completion still fails on roughly four out of five tasks. Neither number suggests computer-use agents are close to reliably automating the kind of multi-hour, multi-application workflows that knowledge workers handle routinely.
With 42 figures documenting failure modes and task breakdowns, the paper reads less as a victory lap for any single lab and more as a diagnostic map. The five challenge categories the authors isolate, particularly implicit-state inference and visual-spatial precision, point to specific architectural gaps rather than a general capability shortfall, giving both model builders and enterprise buyers a clearer picture of exactly where today's agents break down.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
