Our series about model distillation continues. We have a surprising mega interview. We dive into the Astra, Fable and Muse Spark releases to keep you up to date. We will discuss the possible “ChatGPT moments” for robotics. The AI industry has developed a peculiar new benchmark: can you finish reading a model’s system card before its replacement ships? This week, OpenAI, Anthropic, Meta, and Google turned the release calendar into a competitive sport. Somewhere, an engineer is still updating last week’s model comparison spreadsheet. Please give them space.OpenAI’s GPT-6 Astra arrives with an expansive pitch spanning computer use, software engineering, and scientific work. OpenAI reports 98% on FrontierMath Tier 4 alongside improvements in computer interaction. The practical ambition is clear: models that navigate software, execute complicated workflows, and deliver usable work with less supervision. OpenAI is also updating the Codex harness to accelerate computer use, underscoring how much performance depends on the combination of model and surrounding infrastructure. The benchmark chart is becoming a job description—and the software around the model is becoming part of the résumé.Anthropic’s Claude Fable 5.1 and Mythos 5.1 push coding, knowledge work, and scientific research forward. A revealing detail: they share the same underlying model, with different safeguards and access arrangements. Fable is generally available; Mythos is restricted to trusted access programs. Anthropic estimates that cheaper cache reads will reduce costs for typical workloads by around 25%, with larger savings possible for highly agentic work. That matters when an agent repeatedly revisits a substantial working context. Capability, deployment policy, and inference economics increasingly arrive in the same announcement.Meta’s Muse Spark 1.3 focuses on the unglamorous mechanics that make agents useful: maintaining requirements across long tasks, handling conflicting information, revising plans, and asking for help. Meta says it trained across multiple agent harnesses to improve generalization between environments. It also emphasizes keeping track of different tasks within a single conversation, including when users interrupt or redirect ongoing work. Anyone who has watched an agent confidently abandon the original task halfway through a workflow will appreciate the ambition. Remembering what you were hired to do remains an underrated capability.Google supplied the week’s best illustration of the tempo: Gemini 3.8 Flash is its third Flash release in six weeks. Google reports stronger coding and reasoning while retaining 3.7 Flash’s speed and introductory pricing. Alongside it, Flash Cyber targets vulnerability discovery and patching through restricted access. Google attributes gains in the shared foundation partly to training in cybersecurity, a demanding environment for reasoning about complex software. Even the economical workhorse now comes with a specialist security counterpart. Apparently, a six-week-old model family already needs a reunion.Taken together, these releases suggest that sustained, affordable execution is becoming the central competitive frontier. For builders, the frantic pace creates both opportunity and an adoption tax: every upgrade demands fresh evaluations, cost comparisons, and regression checks. The useful response is a disciplined learning loop grounded in real tasks, with success measured by completed workflows and fewer human rescues. Leave room in the architecture for better models, and make upgrades reversible. This is an exhilarating moment to build—provided your evaluation pipeline can refresh almost as quickly as the launch announcements.AI Lab: ByteDance SeedSummary: This paper introduces HARNESSDEV, a benchmark designed to evaluate the capability of LLMs to construct, refine, and maintain their own execution scaffolding rather than solely producing task-level outputs. The authors find that while models can build functional harnesses that match or exceed human baselines in certain domains like machine learning, evolving them stably and transferring performance across different runtime executors remains a major challenge.AI Lab: Qwen Team, Alibaba GroupSummary: The authors propose Terminal-Universe, a framework that reconstructs reusable, executable software environments directly from recorded agent trajectories using deterministic replay and agentic completion. By expanding these environments across breadth (cross-workspace tasks) and depth (multi-round user interactions), the synthesized training data substantially improves the terminal performance of fine-tuned models on benchmarks like Terminal-Bench 2.1.AI Lab: Qwen Team, Alibaba GroupSummary: This work presents E-Commerce Bench, an open-source benchmark evaluating LLM agents managing concurrent online stores over a simulated 365-day business year with deterministic market demand and counterpart negotiations. Evaluating 18 frontier and open-weight models across seven operational dimensions reveals that high profitability does not correlate with optimal performance in negotiation, fraud avoidance, or operational efficiency.AI Lab: ServiceNow AISummary: The paper establishes AgentJudgeBench to systematically evaluate how reliably LLM-as-a-judge models assess structured, dependency-driven tool-calling across varying DAG topologies and difficulty tiers. The study demonstrates that judge alignment degrades significantly with task difficulty—converging to a performance ceiling on hard queries—and that exposing ground-truth sequences can actually degrade the judgment of frontier models due to over-anchoring.AI Lab: AppleSummary: This paper introduces CoGR, a retrieval framework that trains separate language models to directly generate matching keyword representations for both queries and items via an inverted index. Using supervised fine-tuning followed by alternating reinforcement learning with GRPO against frozen opposite-side indexes, the framework achieves significant $F_1$ improvements over competitive sparse, dense, and generative retrieval baselines.AI Lab: Meta AISummary: The authors investigate logit-based knowledge distillation during language model mid-training and discover an inherent reasoning-recall tradeoff, wherein distillation improves reasoning capabilities but slows the acquisition of factual knowledge compared to standard next-token prediction. To resolve this, they introduce Switch Distillation, an entropy-routed objective that selectively applies reverse-KL distillation only to tokens where the teacher is confident, successfully improving reasoning while preserving factual recall through post-training.OpenAI released GPT-6 Astra, its most capable model to date, with saturated scores on FrontierMath and ARC-AGI-3 and a staged rollout that starts with cyber partners before reaching ChatGPT and the API.Anthropic launched Claude Fable 5.1 and Mythos 5.1, one model shipped at two safeguard levels, with better performance than Fable 5 and a 75% cut to cache-read pricing. Meta released Muse Spark 1.3, an update focused on long-horizon agentic and coding work, now live in Muse Code and the Meta Model API with open weights on the roadmap. Google rolled out Gemini 3.8 Flash, a solid step up from 3.7 Flash on coding and agentic tasks, plus a restricted Cyber variant for vulnerability discovery.NVIDIA agreed to acquire Hugging Face for roughly $12.9 billion, with Jensen Huang committing that the platform stays open, multi-cloud and multi-accelerator, and that NVIDIA compute will not be required to build or deploy on it. Crusoe raised over $3 billion at a roughly $30 billion valuation in a round co-led by Atreides Management and Valor Equity Partners, shortly after signing a reported $13 billion five-year cloud deal with Jane Street. Thinking Machines is in talks to raise $1 billion at a valuation of at least $40 billion, with existing backer Accel expected to lead, on an annual revenue run rate reported at just over $100 million. Wonderful closed a $550 million Series C led by Insight Partners at a $5 billion valuation, more than doubling its $2 billion mark from six months ago, as it repositions from customer service agents to an “AI OS” for the enterprise. AIR came out of stealth with $50 million across two seed rounds led by Sequoia and Greenoaks to build an inline firewall that vets the skills, plugins and MCP servers AI agents load at runtime. HiddenLayer raised a $100 million Series B led by Delta-v Capital to extend its AI security platform to agent runtime and coding-agent protection, after growing ARR more than 10x over the past year. TechCrunchNscale is seeking about $3.5 billion in pre-IPO financing, split between up to $1.5 billion in convertible notes and roughly $2 billion from Nvidia, ahead of a listing that could raise another $3 billion. DeepSeek plans to deploy at least 160,000 Huawei Ascend 950DT chips at a new Inner Mongolia data center for inference, while continuing to train on Nvidia hardware. Gimlet Labs raised a $300 million Series B led by Andreessen Horowitz at a $3 billion valuation to scale its multi-silicon inference cloud, which splits models across GPUs, CPUs and other accelerators.Cognition is set to raise around $1 billion at a $47 billion valuation, up from $26 billion three months ago, with annualized revenue reportedly above $900 million. No posts