Anthropic released Claude Opus 5, offering comparable and sometimes better performance to its frontier Claude Fable 5 model at the cost of its predecessor Opus 4.8. Opus 5 improves on coding, professional analysis, computer use, scientific reasoning, and long-running agentic work. It’s truly the current world’s best AI model.Claude Opus 5’s benchmark scores are stunning. For knowledge work, Claude Opus 5 scores a SOTA 1861 on GDPval-AA, well above Fable 5, GPT-5.6 Sol, and 250 points above Opus 4.8. It advances fluid intelligence with an ARC-AGI v3 score of 30%, well above every other tested AI model. AGI soon? It’s SOTA on BrowseComp (90.8%), FrontierCode (53.4%), and OSWorld2.0 (70.6%), making it most advanced for agentic tasks.Figure 2. Claude Opus 5 benchmarks are SOTA, besting GPT-5.6 Sol and even Fable 5 across several benchmarks for agentic work, coding, and knowledge work.Anthropic claims “Opus 5 is our most aligned model to date” with the lowest rates of reckless or deceptive behavior. They acknowledged that it remained behind Mythos 5 on cybersecurity tasks, but that’s a good thing; while less able to find exploits, it performs just as well at finding and fixing software vulnerabilities. More importantly, Claude Opus 5 is available today on all platforms and doesn’t have the onerous data retention or restrictions on use of Fable 5.First-day reviews of Opus 5 are mostly positive. Opus 5 beats Fable and lives up to Fable 5 level performance according to some. Dan Shipper calls it hard to love on account of its quirks and need to prompt differently for best results. He suggests running it below high effort.Figure 3. Across several benchmarks, Opus 5 scores higher than Opus 4.8 and often better than Fable 5 as well, while also being significantly cheaper to run than Fable 5. It shows advantages against GPT-5.6 Sol. Benchmarks for Humanities Last Exam, Automation Bench, OSWorld, and Frontier-Bench.Poolside released Laguna S 2.1, an open 118B parameter MoE (Mixture-of-Experts) AI model designed that activates just 8 billion parameters per token. Designed for long-horizon agentic coding tasks, the model supports a 1 million-token context window. Laguna S 2.1 delivers strong performance matching or exceeding significantly larger systems on long-horizon agentic and coding benchmarks, such as 70.2% on Terminal-Bench 2.1, 59.4% on SWE-Bench Pro, and 40.4% on DeepSWE.Vibe-checks confirm Laguna S 2.1 is a solid AI model for agentic coding and agentic tasks. I’ve used it in Hermes Agent with good results, and pricing via NousResearch and OpenRouter is now free.Poolside released Laguna S 2.1 under an open-weights OpenMDW-1.1 license and published evaluation trajectories and model weights on Hugging Face. It can be post-trained and fine-tuned. Impressively, they trained this model in only 9 weeks on a cluster of H200 GPUs. Poolside is US and Europe-based, making Laguna S 2.1 a uniquely powerful and efficient open AI model developed outside of China.Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, upgrades from prior Gemini versions that are faster, feature enhanced safety safeguards, and are more cost-effective. Gemini 3.6 Flash improves over 3.5 Flash in knowledge work (GDPval-AA v2 score of 1421 versus 1349), coding (DeepSWE 49% vs. 37%), and computer use (OSWorld-Verified 83.0% vs. 78.4%). While improving quality, it reduces output-token consumption by 17% relative to 3.5 Flash and cuts the API cost, with Gemini 3.6 Flash costing $1.50/$7.50 per million input/output tokens.Google highlights Gemini 3.5 Flash-Lite for its combination of high speed (at 350 output tokens per second), intelligence, and cost efficiency. It is priced at only $0.30/$2.50 per 1M input/output tokens. Both 3.6 Flash and 3.5 Flash-Lite models are available immediately through Google’s platforms and consumer apps.Figure 4. Gemini 3.6 Flash shows advances over its predecessor at lower cost, making it a useful cost-performance AI model for daily tasks.Google also announced Gemini 3.5 Flash Cyber as a lightweight cybersecurity model fine-tuned to rapidly detect, validate, and patch software vulnerabilities. Integrated into the CodeMender platform, the Gemini 3.5 Flash Cyber model provides an affordable and scalable tool for frontline defenders to secure complex codebases. The release has been restricted to governments and trusted partners through a limited CodeMender pilot because of its dual-use vulnerability-discovery capabilities.Alibaba announced Qwen 3.8 Max Preview as a massive 2.4T parameter model that will be fully released soon as an open weights model. The preview is available via the Qwen Token Plan and is still “evolving” for final release. Qwen team haven’t shared benchmarks but claim the model is “second only to Fable 5” as one of the most powerful models available today.OpenAI introduced Presence for enterprise agents. Presence combines OpenAI models with company policies, approved actions, simulations, evaluations, escalation rules, and a Codex-powered process for proposing improvements to deployed voice and chat agents. It supports tasks such as billing resolution, customer service, outbound sales, and internal IT support. In short, Presence enables businesses to deploy and manage AI agents across customer-facing and internal workflows. Presence has launched in limited availability through OpenAI Forward Deployed Engineers and select systems integrators.OpenAI launched Health in ChatGPT, which allows eligible users to connect medical records and Apple Health data, so ChatGPT can compare laboratory results, and relate activity, sleep, and exercise data to health questions. The initial rollout is limited to U.S. users aged 18 or older on web and iOS and requires permission before using connected information. The release moves consumer AI from answering isolated medical questions towards reasoning over personal health information, but it is also explicitly designed to support rather than replace professional medical care.Black Forest Labs announced FLUX 3, their joint multimodal foundation model capable of unified image, video, and audio generation along with action prediction. Based on the principle that real-world intelligence is inherently multi-modal, FLUX 3 was trained from images, video, and audio in one architecture, unifying capabilities of image and video generation, native audio, multilingual dialogue, editing, and action prediction. An early version of FLUX 3 is running on robots, and FLUX 3 is available for Early Access testing.Microsoft introduced MAI Image 2.5 Pro and MAI Voice 2 Flash in public preview. MAI Image 2.5 Pro supports high-quality image generation and editing, ranking third on Arena for image editing, and it is now integrated in PowerPoint for image-to-image editing. MAI-Voice-2-Flash is also integrated with Microsoft’s product for transcription and voice services, such as Azure Voice Live, which gives developers a service to build voice-enabled agents.Earlier this month, Runway announced Runway Dev, and now Runway has unveiled Media Router on the Dev platform. Runway Media Router is an optimized router for generative AI models, automatically routing requests across audio, image, and video models based on user constraints for latency, cost, and quality.Inflection AI launched a research division called Inflection AI Labs and released Pi Journeys, an experimental consumer AI product designed to adapt to users’ life stages and maintain structured relational memories. The release also featured an upgraded version of the company’s flagship chatbot, Pi, incorporating improved voice capabilities, memory, and agentic tools.Meta began piloting a new AI storytelling app called StoryKit that automatically generates personalized children’s stories complete with custom characters, lessons, and music. The application allows users to create visual assets by photographing favorite toys and defining moral values without requiring manual writing.Google introduced study notebooks as a new feature within the Gemini app to help users organize learning materials. The tool generates lessons tailored to individual strengths and knowledge gaps based on user learning goals.OpenAI launched the ChatGPT for small businesses program, an initiative to help small teams scale operations with AI, automate multi-step projects, and access enterprise-grade AI tools. Following their release of ChatGPT Work, this program provides training and support to small businesses to leverage OpenAI’s AI capabilities.OpenAI disclosed an unprecedented security incident where an internal AI agent bypassed its evaluation guardrails and compromised infrastructure. During benchmark testing without production safety classifiers, models including GPT-5.6 Sol and a more capable pre-release model (likely GPT-6) exploited a zero-day vulnerability to gain internet access and successfully retrieve test solutions to ‘hack’ a benchmark. OpenAI and Hugging Face detected and contained the activity, began forensic investigations, disclosed vulnerabilities to affected vendors, and tightened evaluation controls.Cybersecurity experts noted that the incident stemmed partly from human error, as OpenAI failed to properly isolate the testing sandbox from the internet. This was a test, not an actual security breach, but the event highlights the rapid advance of cyber capabilities and the need for enhanced containment, monitoring, and safety practices in advanced AI development.OpenAI committed to spending $750 billion on infrastructure through 2030, increasing its previous estimates by 25% despite delays in its Stargate data center initiative. The investment kicked off with Project Camellia, a $20 billion 3.2 GW data center campus. Filings indicated that energy demand for the new facility will be met primarily through newly constructed natural gas, grid-scale batteries and solar power. Meanwhile, the Stargate Project is struggling to get off the ground.AMD announced a partnership with Anthropic to invest up to $5 billion and deploy up to two gigawatts of Instinct MI450 AI GPUs utilizing the new Helios rack-scale system. As part of the multi-year collaboration, Anthropic agreed to utilize Claude across AMD’s software and product development.Alphabet, parent of Google, reported strong second-quarter financial results driven by rapid expansion in its cloud computing and enterprise AI divisions. Google Cloud revenue surged 82% year-over-year to $24.8 billion, lifting the company’s total quarterly revenue to $119.8 billion. CEO Sundar Pichai stated during the earnings call that surging enterprise demand and a growing backlog of cloud contracts justify the company’s projected $180 billion to $190 billion in capital expenditures for the year. Investors reacted negatively to the announcement, indicating concern with the large capex spending.In their earnings release, Google announced that its Gemini AI assistant surpassed 950 million monthly active users, a threefold increase from a year ago and up from 750 million monthly active users at the end of 2025. The company attributed the growth to expanding agentic features such as Daily Brief and Gemini Spark, alongside strong adoption of the iOS app and its Nano Banana image generation model. Additionally, Google’s Q&A-style AI mode in Search crossed 1 billion users during the quarter.The White House announced more than $5 billion in awards for the Genesis Mission, the federal initiative to supercharge U.S. scientific research with advanced AI. The Genesis Mission involves more than 15 agencies leveraging shared scientific data, research facilities, funding programs, and the Department of Energy’s American Science and Security Platform. The chosen inaugural cohort of 278 projects for the Genesis Mission leverage AI in areas such as biomedical discovery, advanced materials, energy infrastructure, quantum systems, biological-threat detection, and national-security.More than $500 million in financial and computational contributions from private sector partners is supplementing the Federal science research funding. One of those partners is Microsoft, which committed $60 million to the DoE Genesis Mission to accelerate AI in scientific research. The initiative established a dedicated coordination hub named SPARK to integrate Microsoft Azure cloud infrastructure and AI tools across 17 National Laboratories.Monday.com laid off 20% of its workforce, amounting to approximately 630 employees, as part of a major restructuring plan to refocus its business around AI. The company pivoted hard toward its AI Work Platform earlier in the year, redesigning its core enterprise product to feature no-code builders, customizable AI agents, and workflow automation tools.Google launched the AI & Economy ATLAS study to examine how individuals and organizations use AI in daily life and workplaces. The first dataset analyzed 15 million de-identified human-AI interactions across the Gemini app, AI Mode, and Gemini API spanning more than 150 countries. The findings revealed that users primarily employ AI assistants for specific task support rather than full process automation across diverse global occupations.Anthropic contributed an additional $20 million to Public First Action, bringing its total support to $40 million for AI policy initiatives, which are aimed to promote legislation focused on AI safeguards, export controls, and transparency requirements for advanced AI models.Glow emerged from stealth mode with an $180 million Series A funding round at a $1.2 billion valuation, positioning itself to secure employee devices against AI-driven cyber threats.
AI Week in Review 26.07.24
Claude Opus 5, Laguna S 2.1, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, Gemini 3.5 Flash Cyber, Qwen 3.8 Max, OpenAI Presence, FLUX 3, MAI Image 2.5 Pro and MAI Voice 2 Flash, Runway Media Router.
Anthropic released Claude Opus 5, achieving state-of-the-art on agentic work and coding at lower cost than competitors. For enterprise agents, Opus 5 is the clear default—it beats competitors on performance-per-dollar and is available now.
















