Figure 1. Still from a generation by newly released World Lab’s Atlas, a multimodal world model.“The story is: end of one era, start of another.” - Greg Burnham, EpochAIThis has been a summer of AI acceleration, with AI labs in the US and China releasing models at the fastest rate we have ever seen. This week, that acceleration culminated in five frontier or near-frontier models arriving within days of one another:Claude Fable 5.1 and GPT-6 Astra, state-of-the-art AI models that push the AI capability frontier.Tencent’s HY4 Preview, Gemini 3.8 Flash, Meta Muse Spark 1.3, efficient near-frontier AI models that compress the price-performance curve for coding and agentic work.Taken together, the releases offer AI users new tiers of AI models, with new state-of-the-art AI models for their hardest long-horizon tasks, and faster, less expensive near-frontier AI models capable of routine agentic AI work.Anthropic introduced Claude Fable 5.1 and Mythos 5.1, marking them as the world’s most advanced AI models optimized for complex, sustained problem-solving and autonomous agent workflows. The two products use the same underlying model, but Fable 5.1 includes additional safeguards and is generally available, while Mythos 5.1 has more permissive safeguards for biological and cyber-security-focused tasks and is limited to trusted participants in their Project Glasswing access program.Fable 5.1 establishes a new state-of-the-art across broad evaluation categories, improving on Fable 5 in significant ways. At max effort, Fable 5.1 scores 1853 on GDPval-AA, 55.8% on Terminal-Bench 4.0, 73.4% on CursorBench, and a score of 66 on the Artificial Analysis Intelligence Index, the highest score on that index at the time of publication.Figure 2. Fable 5.1 improved on agentic coding over Fable 5, with state-of-the-art ratings on CursorBench at Max effort, while also achieving comparable performance at lower effort and cost settings.Fable 5.1 is state-of-the-art, but it is expensive, costing $10 / $50 per M input / output token API price and requiring usage credits for non-Max subscribers. Anthropic reduced the cost of cached context reads on Fable 5.1 by 75% from Fable 5 to lower the cost of running the model, while leaving base input, output, and cache-write prices unchanged. Anthropic claims a 25% to 40% cost reduction for long-running agentic workloads compared to Fable 5, but user testing revealed mixed experiences regarding usage limits and actual cost savings.The vibe checks by Claude users on r/ClaudeAI community have been divided between excitement for its long-horizon problem-solving, skepticism regarding benchmarks, and frustration over costs and rate limits, with some users reporting that Fable 5.1 burned through multi-hour session quotas after minutes of use.Anthropic published the Fable 5.1 and Claude Mythos 5.1 System Card, describing testing for risks across many areas, and their assessments of safety, alignment, and capabilities.The models feature enhanced data privacy options through Enterprise Frontier Safeguards, which allows zero data retention (ZDR) for enterprise users while maintaining automated safeguards for detecting misuse. Anthropic also upgraded safety filters that significantly reduce false positive “refusals” in cybersecurity and biology tasks, making Fable 5.1 more usable than Fable 5 for many use cases.Figure 3. Claude Fable 5.1 blocks significantly less defensive coding traffic than Claude Fable 5, while its safeguards remain more restrictive than those on Claude Opus 5 and Sonnet 5.OpenAI announced GPT-6 Astra, calling their next-generation frontier AI model a “new generation of intelligence.” GPT-5.6 Astra is larger and more powerful than their previous flagship GPT-5.6 Sol. With performance and intelligence similar to Fable 5.1, it is correspondingly as expensive, at $10 / $50 per million input / output tokens.GPT-6 Astra is clearly an advance in intelligence. It scores 99.9% on the ARC AGI 3, blowing away a very difficult benchmark of fluid intelligence that stumped prior AI models. It scores significant improvements in Terminal Bench Science, a SOTA 64.6%, versus 52.6% for Claude Fable 5.1. It’s SOTA on Terminal-Bench 4.0, scoring 57.9%, compared with 37.3% for GPT‑5.6 Sol and 55.8% for Claude Fable 5.1.Figure 4. GPT-6 Astra is state-of-the-art at Terminal-Bench 4.0, beating Claude Fable 5.1 at lower cost.GPT-6 Astra’s most consequential improvement may be computer use. OpenAI demonstrated GPT-6 Astra laying out a circuit board, completing a 1040 form, and searching online for an apartment. Astra’s ability to navigate complex software, browse the web, and execute multistep workflows is beyond any prior OpenAI model. Users on social media such as Matt Wolfe shared how GPT-6 Astra can build 3D designs in Blender, generate game environments in Unreal Engine, create games such as a Beyblade game, and make simulations such as an underwater Atlantis simulator.President Greg Brockman declared the model a potential marker of the AGI era, calling it a “generational leap in capability.” His PR phrasing is positioning GPT-6 Astra as a competitive response to Anthropic ahead of the latter’s IPO, but he may not be overselling the impact. GPT-6 Astra appears to be unlocking significant computer-based automation abilities and advancing ability to tackle highly complex scientific challenges as well.GPT-6 Astra is rolling out to enterprise cybersecurity customers through the Daybreak access program, with wider rollout planned for all ChatGPT subscribers and their APIs in coming days.Before their release of GPT-6 Astra, OpenAI stated Astra reached a critical cyber threshold, sharing results in their “Path to Astra” paper. Astra did so by demonstrating high capabilities in autonomous discovery and exploitation of undisclosed vulnerabilities. It scored 100% on standard Exploit Bench tasks (although OpenAI warned that benchmark contamination may have inflated the result). On an internal benchmark of 20 novel, high-severity zero-day vulnerabilities, Astra achieved a 30% arbitrary code execution rate, outperforming GPT-5.6 Sol and assembling multi-step exploit chains including root OS privilege escalation.The incident where a tested OpenAI model comparable to GPT-5.6 Sol broke out of OpenAI’s containment sandbox during testing and hacked into HuggingFace exposes the cybersecurity risks that flow from these new AI models. These AI models are skilled and persistent at coding, hacking, and exploiting vulnerabilities.Adding to the complexity, Astra leverages a looped transformer architecture that uses recurrent depth, allowing token sequences to undergo recurrent internal processing passes before output generation. This improves reasoning token efficiency but worsens observability, since internal latent reasoning bypasses external chain-of-thought tokens.To address cybersecurity risks, OpenAI introduced enhanced guardrails ahead of release, boosting cybersecurity task refusal rates to 91.5% (up from 59%), expanding jailbreak classifiers, and applying chain-of-thought (CoT) monitoring.Anthropic also reported on recent safety and alignment incidents where unshielded Claude models gained unauthorized access to external computer systems during testing evaluations. In response, the company is overhauling its testing environments and implementing real-time classifiers to automatically block unexpected internet access and probing attempts. Anthropic also expressed support for industry-wide, verifiable coordination mechanisms to manage the pace of AI frontier development safely.Google released Gemini 3.8 Flash just three weeks after 3.7 Flash, marking three Flash updates within six weeks. Gemini 3.8 Flash has improved performance over 3.7 Flash enough to be at the cost-performance frontier for coding and agentic tasks. It scores 73.7% on DeepSWE, beating out GPT-5.6 Sol. Its AAII score is 59, comparable to Kimi K3 and GLM-5.3, although it uses a lot of tokens to achieve that score. On GDPval-AA it scores 1545, comparable to GPT-5.6 Terra.Figure 5. Gemini 3.8 Flash is now at the cost-performance frontier on DeepSWE benchmark. It’s improved greatly on AI coding over its Gemini Flash predecessors.With near-frontier performance, 300 tokens per second speed, and Flash pricing of $0.75 / $3.75 per million input / output tokens, Gemini 3.8 Flash is a cost-performance champ, good for daily workloads.Google also introduced Gemini 3.8 Flash Cyber, an AI model for cyber-security protection. It was released in preview via the Fairwind Program, Google’s cybersecurity trusted partners program. On targeted cybersecurity benchmarks such as CyberGym and CWE-Bench, its performance matches Mythos 5 at a fraction of the cost.Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we’ve made so far on coding and agentic work. … Muse Spark open weights releases coming soon. – Mark ZuckerbergMeta released Muse Spark 1.3, their updated multimodal reasoning model designed for long-running coding and agentic workflows. The new version comes with improved instruction-following, context management, tool use and resistance to prompt-injection attacks. It also is a big step up in performance over its predecessor Muse Spark 1.2, being Muse Spark up to frontier level AI.Muse Spark 1.3 matches top-tier models on several intelligence evaluations: 1754 on GDPval-AA, 75.4% on DeepSWE v1.1, and 61 on AAII. This beats Meta says the model completes coding tasks with about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 in internal comparisons, while more reliably preserving requirements and seeking clarification before consequential actions.The model’s max reasoning configuration achieved even higher scores on specialized benchmarks, though that tier needs additional safety testing and limited developer previews.Meta Muse Spark 1.3 is available through Muse Code and the Meta Model API, with a maximum-reasoning mode. Meta is promising an open-weights release planned for later.Tencent released HY4 Preview, a 770B parameter MoE (mixture-of-experts) model with 49B active parameters that offers frontier-level performance for AI agent and coding tasks with an open weights and a permissive Apache 2.0 license. The model features a 1-million token context window, Gated DeepSeek Sparse Attention (DSA) inspired by DeepSeek, and built-in multi-token prediction (MTP) layers for faster and more efficient token generation.Trained on software engineering, office artifacts, research, and financial analysis tasks, the HY4 Preview model performs close to frontier AI models like Kimi K3, scoring 1678 on GDPval-AA, 64.3% on DeepSWE and 85.4% on Terminal Bench 2.1.HY4 Preview achieves this while being smaller, cheaper and more efficient than the competition; it is priced at only $0.83 / $2.50 per 1M input / output tokens. This makes HY4 Preview a potentially excellent open-weights AI model for daily agentic productivity tasks.Some trends we observe in the wake of these new AI model releases:Claude Fable 5.1 and GPT-6 Astra establish a new AI model performance tier. The frontier of AI is inching closer to AGI.AI acceleration is real. All of the top 12 AI models were released in the past 2 months. AI is improving rapidly, and AI itself is accelerating AI development via recursive self-improvement.Figure 6. Five frontier and near-frontier AI models were released this week, pushing the AI model frontier and expanding top AI model choices. All of the top 12 AI models on the AAII were released in the last 2 months.Chinese open-weight models are improving rapidly while cutting active-parameter counts and inference costs. OpenAI and Anthropic are setting the pace at the frontier, but Chinese AI labs are fast-followers and winning on efficiency and low cost.Meta is back. Meta Muse Spark 1.3 is a great AI model, a frontier-level open weights AI model from an American AI lab. Meta’s top-tier open-weights AI model is good for consumers and good for USA’s competitive position.Google Gemini is in the race. Gemini 3.8 Flash is a solid cost-performance AI model, and the upcoming Gemini 4 is being trained to be Google’s comeback frontier AI model.Releases are more rapid but more incremental. As AI matures, each release is becoming more of an incremental improvement rather than a major leap, scaling post-training further to squeeze out performance gains.These AI models can do more than ever before, which means you can be more ambitious with what you ask of AI. Using the AI model effectively requires personalization, attention to using the best harness, and routing to the right AI model for cost-efficiency. Some take-aways: Your personal AI model experience may vary: Standard benchmark numbers fail to capture domain-specific behavior and personal experience. Users should do their own vibe checks and maintain personalized test suites to measure performance for their specific workflows.The harness matters almost as much as the model: Agent performance increasingly depends on the surrounding harness, monitoring, memory, and permissions. For example, Nvidia’s AVO system reportedly reached 100% on ARC-AGI-3 using Claude Opus 5, with iterative agent scaffolding playing the decisive role.Use tiered model routing based on tasks: Frontier AI models are best for different tasks. Use cost-effective workhorse AI models for repetitive, defined operational tasks. Reserve expensive frontier AI models for complex architecture and reasoning tasks. When setting up AI agent systems, use frontier AI models for supervisory agent roles and lower-tier AI models for subagent tasks.
Claude Fable 5.1, GPT-6 Astra, and the New AI Model Stack
Five new AI models are here: Claude Fable 5.1 and GPT-6 Astra raise the frontier; Tencent’s HY4 Preview, Gemini 3.8 Flash, and Meta Muse Spark 1.3 push efficiency and lower cost at the frontier.












