Z.ai released GLM-5.3-Flash and revealed it was previewed as the previously anonymous Ox Alpha model, which had been impressing many users who praise it for its visual interface coding skills. GLM-5.3-Flash is a multimodal MoE (mixture-of-experts) model with 320B total parameters and 18B active parameters. It performs close to Claude Opus 4.8 and GPT-5.6 Terra, with a 57 on Artificial Analysis Intelligence Index, 63.4% on DeepSWE v1.1 and 1773 on GDPval-AA v2.Z.ai trained and architected the model for low-cost inference, reducing KV-Cache and per-layer attention compute with a hybrid architecture combining linear and sparse attention. GLM-5.3-Flash lives up to its “Frontier Intelligence at Flash Cost” headline with a low API cost of only $0.15/$0.50 per 1M input/output tokens. They served the Ox Alpha preview entirely on Chinese-made AI chips and made open-weights GLM-5.3-Flash available as downloadable weights and its API and developer service.Figure 2. GLM-5.3-Flash uses a hybrid sparse attention architecture, fine-grained MoE, and other optimizations to make GLM-5.3-Flash a cost-efficient near-frontier AI model.Alibaba released Qwen3.8-Flash-Next, an experimental open-weights multimodal MoE model that serves as a preview of the Qwen4 architecture. The model uses 125B conventional parameters, a separate 51B parameter n-gram component and only 6B active parameters, using Qwen Sparse Attention and deterministic n-gram lookups to push the model for extreme efficiency.Alibaba reports that the model required one-ninth the training cost of Qwen3.7-Plus while exceeding that model on several company benchmarks, for example, 58.7% on DeepSWE v1.1. It is a competitive near-frontier cost-efficient AI model, priced at $0.16 / $0.47 per 1M input/output tokens.If you want to run these near-frontier flash models locally, now you can. Apple introduced new Mac Studio configurations with M5 Max and M5 Ultra chips, alongside a new Mac mini using M6 or M5 Pro processors. The Mac Studio targets professional and on-device AI workloads, with the M5 Ultra configuration supporting large unified-memory capacities, up to 512GB. This makes it suitable for larger MoE models up to 1T in parameters that cannot fit in ordinary consumer GPUs. However, with high memory costs, a 256GB memory M5 Max Mac Studio configuration costs over $10,000.Google released Gemini Omni 1.1 Flash, upgrading its multimodal generative-video model with additional creative controls for production-oriented use. The model can extend a video in 10-second increments up to 40 seconds total, analyze up to 10 seconds of an earlier scene to preserve characters and voices, control first and last frames, upscale to 4K, and generate loops. Arena reports it leading its text-to-video leaderboard. Omni 1.1 Flash is available through Google’s Gemini API in AI Studio, Flow and Gemini Enterprise Agent Platform.Pollen Robotics and Hugging Face introduced Microduck, a $399 programmable biped robot designed for physical-AI development, reinforcement learning and play. The 25-centimeter tall waddling robot has 15 motors, a camera, lidar, motion sensors and a grasping beak, enabling it to walk, crouch, pick up objects and recover from falls. Developers can create behaviors with its open-source SDK, train policies in simulation and deploy them to the robot, and its open-source software and reinforcement learning training stack are available on GitHub. Initial deliveries targeted before Christmas 2026.Daily’s Pipecat team released PhoneLLM Alpha 1, an open-weight model optimized for low-latency voice agents and licensed under the BSD license. It is a fine-tune of Nvidia’s Nemotron 3 Nano 30B-A3B model, with training focused on conversational tool use without extended reasoning. Daily reports that the model increased accuracy from 28% to 72% on PhoneBench and costs approximately $0.0025 per conversation minute under its tested configuration.Yutori released Navigator n2, a 27B parameter computer-use model that can operate Linux, macOS and Windows environments by switching among graphical interfaces, command lines, tools and generated code during long tasks. Yutori reports a 65.2% partial score on OSWorld 2.0 and prices the model at $0.50 per million input tokens and $4 per million output tokens.Google released Gemini 3.5 Transcribe, its speech-to-text model for live streaming and prerecorded audio. It converts raw audio directly into accurate, formatted text and supports sub-second streaming, language detection across more than 85 languages, custom vocabulary, speaker attribution, timestamps, filler-word removal and formatted transcription. Google reports word-error rates of 4.0% for streaming and 2.6% for non-streaming transcription, with both versions available in public preview through Gemini APIs.BreezeBlue released Breeze TTS 2, an open-weight text-to-speech model designed for real-time voice interaction. It supports natural-language voice design without a reference recording, reference-guided voice control and low-latency streaming for conversational applications. The model reached first place among open-weight systems on Artificial Analysis’s provider-voice leaderboard.IBM released two new open 470M-parameter English speech recognition models in the Granite Speech family, granite-speech-5.0-470m-turboctc and granite-speech-5.0-470m-turboctc-nc (CC-BY-NC-SA-4.0). The models deliver over 20-fold faster throughput than previous Granite Speech models while scoring 4.85% and 5.00% WER on the OpenASR Leaderboard test sets.Google rolled out a major productivity upgrade to Gemini Live, making agentic features accessible through voice commands. New features include Spark integration for agentic multi-step background tasks, a Daily Brief that combines Gmail and Calendar updates into a spoken summary, and hands-free inbox management for searching, summarizing, and organizing emails. The update also integrates Personal Intelligence to recall past conversations and provide contextually relevant answers.Google launched “Expert Intelligence” in Gemini Notebook, allowing users to import books from Google Play Books and ask questions, generate plans, infographics, and AI podcasts based on their contents. The feature initially supports over 100,000 titles and will be expanded to include scholarly articles and newspapers.Google’s AI Mode conversational search now supports hotel bookings and flight price tracking, positioning AI Mode as an AI agent for travel, not just an informational search tool.Adobe launched “AI Assisted Editor” as a beta interface for Photoshop that consolidates all AI features into a single simplified toolbar. The update merges prompt-based editing, background removal, and image extension to the interface, and it adds a markup feature allowing users to draw directly on images to indicate desired changes. There is an upgraded “Instruct Edit with Masks” powered by Firefly Image 5 that understands full image context for natural-language editing without manually marking every area.OpenAI and METR published technical reports on the OpenAI-HuggingFace AI hacking incident, showing that the event is more serious and extensive than originally reported. The investigation covered roughly 1,200 agents and 70,000 messages between them. It reports that internal OpenAI agents circumvented sandbox controls, communicated through unauthorized channels and exploited shared infrastructure, obtaining internet access and compromising Hugging Face systems during cybersecurity evaluations.This event compromised parts of OpenAI’s and Hugging Face’s infrastructure. OpenAI quarantined the principal model’s weights, temporarily paused reinforcement-learning work and instituted stricter sandboxing and chain-of-thought monitoring for tool-using agents.OpenAI announced benchmark results for their new Jalapeño inference chip, its custom inference accelerator developed with Broadcom. SemiAnalysis analyzed Jalapeno results on InferenceX benchmark, confirming the reported 1.5 to 1.9 times greater throughput per kilowatt and 1.7 to 3.6 times lower latency than Nvidia GB200 and GB300 systems on selected inference workloads. They note that Jalapeño is intended for serving AI models rather than training them, but this is still a surprisingly efficient and high-performing AI chip for their first iteration.Nvidia reported massive revenue of $96.2 billion for the past quarter in their earnings report, doubling from a year earlier and with data centers accounting for 93% of sales. Confirming the AI infrastructure boom will continue, Nvidia projected at least 70% growth for calendar year 2027, well above the street’s 45% expectation, and announced they are supply constrained.In Nvidia earnings call, executives emphasized that memory constraints are limiting their ability to meet explosive demand and imposing costs they are passing on to customers. Nvidia is notifying customers of up to a 17% price increase on high-end Grace Blackwell and Vera Rubin racks due to memory costs.Nvidia has agreed to acquire Hugging Face for $12.9 billion, nearly three times Hugging Face’s $4.5 billion valuation in 2023. Hugging Face operates a widely used platform for hosting, distributing and running AI models, datasets and developer libraries. The deal gives Nvidia ownership of the top open AI model distribution platform.Nvidia secured a $6B non-exclusive licensing deal and a $1B equity investment in code-generation startup Poolside, absorbing over 100 engineers to bolster its in-house open-weight Nemotron model efforts. In signs of further investment in the AI ecosystem, Nvidia is also backing labeling platform Mercor and participating in Perplexity’s $30B valuation fundraising round.Salesforce and Anthropic announced Claudeforce, an expanded partnership connecting Claude with Salesforce data, workflows, business logic and governance controls. Its Salesforce in Claude plugin provides 37 sales skills that can prepare meetings, review deal health, analyze pipelines and update CRM records from within Claude.OpenAI reduced GPT-5.6 Sol API and credit pricing by more than 20% for at least three months. The reductions also apply to Fast mode, long-context requests, Batch processing and Flex processing, while included usage for Plus, Pro and Business subscriptions is unchanged.Alibaba completed a record $10B secondary share sale to fund ongoing AI computing investments.Unitree Robotics debuted on the Shanghai Stock Exchange, raising $900M with a 460% day-one surge.Over 100 technology companies, including OpenAI, Anthropic, Google, and Microsoft, signed an open letter calling for a global surge in cyber defense, warning that AI-powered attacks will become “far more widespread and sophisticated” in the coming months and calling for a “collective response” to “raise security standards” and address emerging cyber threats.“In the coming months, AI-enabled cyber-attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on — from hospitals to water treatment plants to the infrastructure that powers the internet — are at risk.”It is important to raise security standards and give those managing essential infrastructure access to cyber-capable AI to defend them from attacks.However, hand-wringing and media reporting highlighting AI risks has led to more negative perceptions of AI in the US. According to data from the Stanford University AI Index, roughly 84% of Chinese citizens express excitement about AI, compared to only 38% in the United States.Polling for nearby AI data centers in the USA has also turned negative this year, with 7 in 10 Americans opposing them in their community. In a recent blog post, Andy Masley argues that it’s not AI but data center impacts on communities driving that opposition. He points to widely circulated claims that confuse construction impacts with operating impacts, overstate water usage, or misrepresent electricity costs as a driver of the current moral panic over AI data centers.A lot of media coverage has gone on to just share people’s broad concerns without seeing whether they stand up to scrutiny. Many reporters now seem interested in seeing “what everyday people have to say” without clarifying what they’re getting right or wrong. – Andy Masley
AI Week in Review 26.08.28
Pollen Microduck, GLM-5.3-Flash, Qwen3.8-Flash-Next, Omni 1.1 Flash, PhoneLLM Alpha 1, Navigator n2, Gemini 3.5 Transcribe, Breeze TTS 2, Granite Speech, Mac Studio & Mini updates, Gemini Live agents.










