Two years ago, AI agents that operate computers the way humans do, by looking at a screen and deciding where to click, were completing roughly one in eight assigned tasks correctly. Today, the best systems are clearing 85% on the same test.

The benchmark in question is OSWorld, currently the most widely cited evaluation platform for multimodal computer-use agents (CUAs). These are AI systems that navigate a desktop environment using screenshots, mouse clicks, and keystrokes rather than direct API calls, essentially watching a screen and acting on what they see.

The numbers behind the leap

In April 2024, leading agents were scoring around 12% on OSWorld. By mid-2025 that figure had climbed into the mid-30s. By June 2026, three Anthropic models sit at the top of the OSWorld-Verified leaderboard: Claude Mythos Preview at 85.4%, Fable 5 at 85.0%, and Opus 4.8 at 83.4%.

For context, the human baseline on OSWorld sits at approximately 72%. When Simular’s Agent S3 crossed that threshold in December 2025, reaching 72.6%, it was the first widely noted instance of a CUA surpassing average human performance on the test. Anthropic’s current crop has now pushed well past it.