Figure 1. AI video generation model SeeDance 2.5 has arrived. It can generate up to 30 seconds of action-packed video in 720p .ByteDance’s latest video model SeeDance 2.5 has officially launched, upgrading its joint audio-video model to generate 30-second clips in a single pass for longer, more controllable AI video. The model can use up to 30 images, 10 videos and 10 audio clips as references. It has precise editing controls with timestamp-specific prompting and editing, camera-perspective controls, green-screen replacement and improved motion, lighting and audiovisual quality. SeeDance 2.5 can also extend videos across multiple rounds while preserving characters, environments and narrative pacing.User reviews suggest imperfections in complex action sequences, and ByteDance acknowledges that multi-subject interactions can still be unstable. Another limitation is 720p output, less than some video generation alternatives. Curious Refuge calls it “incredible, until it isn’t.” However, Seedance 2.5 is moving AI-generated video beyond isolated clips toward an end-to-end system for creating and revising coherent, multi-minute audiovisual stories.DeepSeek released DeepSeek V4 Flash Official API in public beta, and they are presenting incredible near-frontier benchmarks for this 284B parameter MoE model with only13B active parameters. DeepSeek V4 Flash 0731 benchmark scores are surpassing V4 Pro-Preview’s scores, including near-SOTA 82.7% on Terminal Bench 2.1 and 54.4% on DeepSWE. In addition, DeepSeek reports:We’ve massively upgraded its Agent capabilities …The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!A limitation of DeepSeek V4-Flash 0731 is that it is text-only, but it is priced at just $0.14 / $0.28 per million input/output tokens, making it the best current cost-performance AI model for many coding and agentic uses, as much as one tenth the cost of comparable closed AI models.Figure 2. DeepSeek-V4-Flash-0731 achieves near-frontier performance, beating GLM-5.2 and close to Opus 4.8, despite being much smaller and much cheaper than its competitors.OpenAI announced updates to GPT-5.6 model family, significantly reducing prices for the Luna and Terra models and improving the GPT-5.6 Sol Fast mode to deliver up to 2.5 times faster speeds for Sol without sacrificing intelligence. The most significant improvement is the 80% drop for GPT-5.6 Luna, now only $0.20 / $1.20 per million input / output tokens, making Luna a price-performance leader with DeepSeek V4 Flash.Figure 3. OpenAI’s GPT-5.6 Luna (max) leads in AI intelligence price-performance, beating out Gemini 3.6 Flash, Grok 4.5, and even Gemini 3.5 Flash-lite. Only just-released DeepSeek V4 Flash 0731 outdoes it on lowest cost for near-frontier intelligence.Open AI also dropped GPT-5.6 Terra prices to $2/$12, a 20% drop, and GPT-5.6 Sol Fast mode in the API now offers up to 2.5x the speed for 2x the price at the same intelligence. OpenAI claims these all came from the “recursive self-improvement” of GPT-5.6 Sol examining their inference infrastructure and implementing efficiency improvements that reduced serving costs. As Chubby said, this means “the speed of releases is increasing, and models are improving even faster.” AI improvements are not slowing down.Google DeepMind has launched Gemini Robotics 2, releasing three versatile foundational models to support embodied intelligence and robotics. Gemini Robotics ER 2 is a “high-level brain for robots,” an embodied reasoning model for robotic systems that understands the physical world and orchestrates multi-step tasks by planning quickly and handing off execution to a VLA model. The model can also collaborate by processing live video streams and interact with external tools like Google Search.Gemini Robotics 2 is their latest vision-language-action model (VLA) that converts vision and language input into motor control. It translates natural language into coordinated whole-body movements across diverse hardware like humanoid and robotic arms. The third embodied model is Gemini Robotics On-Device 2, built to run locally on robotic devices that can adapt to run on a broad range of robots.Demonstrations show a humanoid robot using these models to manipulate objects, perform dexterous tasks such as tying bags, and collaborate with another robot. It’s impressive, but not as impressive as Uniubi’s back-flipping robot dog and not yet production-ready. Google acknowledges that movement speed and multi-finger reliability still need improvement. Gemini Robotics 2 models are publicly accessible to developers via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform.Figure 4. Google’s demonstrates robotic dexterity with their Gemini Robotics 2 platform, having a humanoid robot change a lightbulb.SpaceXAI announced Grok Voice Think Fast 2.0, their next-generation intelligent voice model with improved conversational capabilities and state-of-the-art transcription accuracy, especially in noisy real-world environments. This model can reason while speaking, making the model substantially smarter than other speech-to-speech models with no impact on latency. It is priced at $0.08 per minute of audio and integrates reasoning tokens to execute tool calls more efficiently during live conversations.Microsoft launched MAI-Cyber-1-Flash inside MDASH, putting a model built to find challenging software vulnerabilities into Microsoft’s MDASH vulnerability identification and remediation harness. Microsoft’s first dedicated cybersecurity model is designed to handle roughly 90% of vulnerability-analysis tasks efficiently, while MDASH routes the hardest cases to larger models such as GPT-5.4. Microsoft reports that this multi-model system scored approximately 96% on CyberGym, 12% better than even Claude Mythos while costing 50% less than its previous MDASH configuration.Figure 5. MAI-Cyber-1-Flash and GPT-5.4 in the MDASH harness is SOTA on CyberGym benchmark of software vulnerability assessment.Last week, Block launched Buzz, an open-source AI Agent collaboration platform that combines Slack-like channels, threads, direct messages and voice with code repositories, workflows and first-class AI-agent participation. Since then, Buzz has hit 20k stars on GitHub and generated buzz as a truly AI-native collaboration layer, not just another agent harness.Built on the decentralized Nostr protocol, Buzz gives every human and agent a portable cryptographic identity, permissions and auditable activity record, while its ACP integration connects existing agent harnesses, including Codex, Claude Code and Goose, to the shared workspace. This could be a preview of a new AI orchestration paradigm, but it remains an early version, with mobile clients and huddle features still under development. Consider it for hobbyist exploration; it’s not a “Slack killer” yet.Thinking Machines released Inkling-Small, an efficient open-weights 276B parameter Mixture-of-Experts model with 12B active parameters that achieves comparable performance to their Inkling model at a quarter of its size. Inkling-Small is a natively multimodal model crafted for audio intelligence, making it suitable for real-world audio applications. As with Inkling, Inkling-Small is designed for enterprise developers who customize the AI model for specific uses.SpaceXAI introduced an integration that brings its Grok AI assistant directly into Google Workspace applications, including Sheets, Slides, and Docs. The tool enables users to query spreadsheets, generate presentations from outlines, and draft text documents using natural language commands. The add-on is available for free on the Google Workspace Marketplace alongside its existing Microsoft 365 integration.Google updated Gemini Spark with direct Google Chrome integration for direct web browsing use, allowing the AI assistant to use logged-in accounts and saved passwords to handle complex web tasks like scheduling appointments and researching flights The updated features include built-in safeguards against prompt injection and require user approval for sensitive actions such as payments. Additionally, Google expanded access to Gemini Spark for Google AI Pro subscribers, putting the Spark agentic interface inside the Gemini web and mobile apps for users in 160 countries.Google added natural voice interactions for the Gemini app on macOS. The voice interaction feature provides intelligent dictation that automatically cleans up spoken transcriptions, such as removing filler words and handling mid-sentence corrections. Users can also opt into advanced Gemini reasoning to let the app analyze screen context and execute complex desktop tasks via voice input.Friend re-launched its AI pendant device, featuring a newly added speaker that enables two-way voice conversations with users. The redesigned pendant includes capabilities to remember user conversations for extended periods. The updated wearable is now priced at $249 alongside a $10 monthly subscription.Nvidia expanded its Agent Toolkit with re-engineered PhysicsNeMo libraries and CUDA-X components that agents can invoke for physics modeling, sparse-system solving and quantum-chemistry calculations. The package is aimed at autonomous engineering workflows in chip design, verification, packaging and industrial simulation, with Cadence, Siemens and Synopsys among the early adopters.OpenAI discovered that retaining reasoning in its API harness tripled GPT-5.6 Sol’s score on ARC-AGI-3, increasing its score on the puzzle benchmark from 13.3% to 38.3%. Previous low scores were caused by API settings that discarded private reasoning and utilized rolling truncation. Adjusting the harness significantly improved the model’s memory retention and complex problem-solving efficiency during agentic evaluations.Moonshot AI released the Kimi K3 Technical Report. The report on the 2.8T parameter Kimi K3 open-weights model presents several architectural and training innovations: It combines efficient Kimi Delta Attention with periodic global-attention layers, using Attention Residuals to retrieve information from earlier network depths; it introduces Stable LatentMoE, which activates 16 of 896 experts while using SiTU-GLU and Quantile Balancing for stability in training. These and other innovations in Kimi K3 result in 2.5 times greater scaling efficiency than Kimi K2.Nvidia released a real-time world model called Cosmos-H-Dreams for surgical robotics. Cosmos-H-Dreams is an action-conditioned generative simulator that predicts surgical scenes from an initial image and a stream of robot movements. By distilling a larger bidirectional model into a causal student and serving it through Nvidia’s FlashDreams runtime, researchers increased generation from roughly 10 to about 160 frames per second on one RTX Pro 6000. Nvidia stresses that this remains a research platform, not a surgical tool.Claude discovers improved attacks on cryptographic algorithms. Anthropic researchers used Claude Mythos Preview to develop a key-recovery attack against the HAWK-256 experimental post-quantum signature scheme and a new technique that speeds up a known attack on seven-round AES-128 by 200-800 times. The results demonstrate meaningful AI-assisted cryptographic discovery, but neither breaks production versions of AES or other encryption systems.OpenAI launched the ChatGPT for Academic Researchers program to provide frontier AI model access to academic researchers. The program will start with providing free access to frontier AI models such as GPT-5.6 Sol Pro to its initial cohort of 10,000 researchers this summer. The initiative is backed by a commitment of more than $250 million through 2027 to support external scientific research, grant writing, and advanced computational workflows, and it will expand to support 100,000 scientists, mathematicians, and engineers at selected institutions.Stock market turbulence in AI and tech stocks in July led to the wipeout of a hedge fund betting on AGI. Situational Awareness, a hedge fund founded by former OpenAI employee Leopold Aschenbrenner, faced margin calls as bets on companies like Marvell went sour, and he was forced to unwind public equity holdings as financial losses mounted.Nvidia has invested in Ilya Sutskever’s Safe Superintelligence, announcing a long-term partnership that includes Nvidia investment in SSI and access to Vera Rubin systems that would increase SSI’s computing capacity by an order of magnitude for their AGI development work.CEO Dario Amodei clarified that Anthropic has never advocated for a ban on open-weights models. In a post on Anthropic’s website, he emphasized that “Open-weights models that don’t have dangerous capabilities are a public good,” and protectionist bans fail to address national security concerns regarding dangerous capabilities and industrial-scale distillation. Amodei instead reiterated support for targeted export controls on advanced chips, restrictions on massive model distillation, and mandatory pre-release safety testing.His statement is a response to commentary on Anthropic’s conspicuous absence on a joint statement “Open Weights and American AI Leadership” in support of open source AI models.In yet another joint statement, major AI research organization leaders signed an open letter at pacingthefrontier.com cautioning to pace AI development as recursive self-improvement works to accelerate AI progress:The world’s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems. … We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” - Pacing the Frontier statement.
AI Week in Review 26.07.31
SeeDance 2.5, DeepSeek V4 Flash 0731, GPT-5.6 Luna 80% price cut and faster GPT-5.6 Sol, Gemini Robotics 2 and Robotics ER 2, Grok Voice Think Fast 2.0, MAI-Cyber-1-Flash, Block's Buzz, Inkling-Small.
DeepSeek V4 Flash and OpenAI Luna spark price war on frontier models ($0.14–0.20/M) with 80% cost reduction, reshaping inference. LLM stack ROI rethinking urgent; competition shifts from capability to cost-per-inference while robotics and cybersecurity specialization emerge.











