What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know. I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates around Pacing the Frontier and how people see the current rate of progress. As groundwork for understanding that and future similar discussions, I have laid out The Three AI Pills: Different people either fail to believe in current AI, believe only in current AI, in AGI or in ASI (superintelligence), and most sincere disagreements stem from this disagreement.One sign of the increased pace of progress was when OpenAI’s unreleased model Astra solved 10 major open math problems. Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google and CEO Sundar Pichai are now firmly in control of DeepMind, and all the promises made to DeepMind, including about safety, look fully dead. Koray Kavukcuoglu will now run DeepMind. He has been there for a long time, but all signs point to him being a capabilities guy.Finally, we supposedly now have a White House frontier AI safety evaluation framework. Not that we are allowed to know anything about what it is, except that if you take absolutely no safeguards (as in you put the weights on HuggingFace) then you are immune from the safety evaluations. Otherwise, no, you can’t see it.Language Models Offer Mundane Utility. Make sure you publish first.Huh, Upgrades. Luna prices slashed 80%, Alibaba gives us a new Qwen.On Your Marks. MirrorCode, Prime Agent.Choose Your Fighter. Know the size of your potential fighters.Get My Agent On The Line. YC shares its multiagent harness. Deepfaketown and Botpocalypse Soon. Palo Alto CEO puts out AI slop.Fun With Media Generation. Seedance 2.5, concerns with the right to satire.Cyber Lack of Security. Security through obscurity is about to die. Needs food.Some People Need Practical Advice. Prepare for The Hackening.A Young Lady’s Illustrated Primer. Don’t miss the real danger here.They Took Our Jobs. The top talent is better off. Most people are not top talent.Get Involved. The EU AI Office Safety Unit is hiring.Introducing. Inkling-Small, a 276B model from Thinking Machines.Demis Hassabis No Longer CEO At DeepMind, Jeff Dean Leaves. Yikes.In Other AI News. OpenAI responds to Apple’s lawsuit. AI Persuasion Exceeds Human Level Over Similar Text Channels.Show Me the Money. Situational Awareness fund gets margin called.Bubble, Bubble, Toil and Trouble. Fighting to avoid the permanent underclass.Quiet Speculations. Are we getting an SSI model? An anti-AI Woke 2.0?My Offer Is Nothing. The White House eval rules: Secret, backwards, mandatory.The Quest for Sane Regulations. Time for that Dean Ball apology form.Chip City. Why so many are so opposed to data centers.The Week in Audio. Samuel Hammond, Alex Turner.People Just Say Things.Rhetorical Innovation. Please Don’t Kill Us, and why many VCs hate American AI.Open Weights Models Are Unsafe And Nothing Can Fix This. Mitigation options.Cooperative Alignment. Consciousness-related training has many other impacts.Other People Are Not As Worried About AI Killing Everyone. Only robots.The Lighter Side. OpenAI or the bush. OpenAI releases notes on their 10 math breakthroughs and how Astra found them.Sol finds a counterexample to disprove the Maxwell conjecture. A Patrick McKenzie special: Feed it a transcript and ask ‘what didn’t they tell me?’ Publish your AI math result first by posting a paper entirely written by AI. In a sufficiently competitive race world, you don’t get to have nice things like well-written fleshed out papers, or AI alignment, or human survival, but I digress. The solution, presumably, is to allow people to submit hashes or a stub to claim priority, then give them a limited window afterwards to publish. OpenAI slashes prices on Luna by 80%, to the low price of $0.20/$1.20, and on Terra by 20% to $2/$12, and is adding a Fast Mode for Sol in the API. OpenAI is framing this as passing gains from optimization on to customers. I don’t doubt there were some gains from optimization, but 80%? My take is that there are three basic ways to price models and many other things.Price to maximize profits. Price in proportion to your marginal costs.Price to win market share.Previously I assumed OpenAI was centrally using method #2. They figured they would have models of different sizes, and relative prices reflected relative marginal costs. Then customers could choose the efficiently best basket of goods.Sam Altman (CEO OpenAI): we want to offer the best price/intelligence tradeoff at every level.This seems like a shift to method #3. OpenAI wants to compete for the lower end of the market, and faces a lot of Chinese competition, so they are slashing prices there. He’s asking ‘what do our competitors offer?’ and trying to do better. Might be a good move, might not. Whereas at the high end they effectively have a duopoly with Anthropic, so a price war would be foolish. Alibaba gives us Qwen 3.8-Max-2.4T. Links to: Blog, Qwen Studio, API. Weights are coming next week. As usual, the benchmarks are presented as strong. Pricing is $2/$6, or $0.25 implicit caching.Based on the surrounding silence and my pattern recognition skills, and how many of these benchmarks are odd choices, I presume this is benchmaxxed and substantially behind Kimi K3. Bloomberg frames this as Qwen ‘matching or exceeding’ Fable performance, which seems like Gell-Mann Amnesia territory. If that was remotely true, we would know. It fits that this post was partially written by AI as per Pangram, and rather obviously so. How do major publications not run Pangram checks in August 2026? MirrorCode, a measure of how large a software project a model can do, has Fable 5 well out in front of OpenAI models. No sign of Opus.Prime Agent from Prime Intellect, is a new agent harness. Their big brag was that their harness scores 95.5% on ARC-AGI-3 using Opus 5. That tells you Prime Agent is vastly superior to the ARC-AGI-3 harness that ARC forces you to use on the real test, but the ARC-AGI-3 harness is intentionally terrible. This shows us Opus 5 has a slower uptake but a much higher maximum performance level than Sol in this setting, but does not tell us that Prime Agent is a good (or bad) harness.If your AI use case would be good except it is too expensive, but we’re not talking orders of magnitude too expensive, start getting it ready now, and it will work soon.nic: GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price.nic carter: this is what cracks me up when people construct these elaborate bear cases based on AI being "too expensive". ok just wait 6 months and it will probably be 10x cheaper for the same unit of intelligence. (4 months/13x in this case)Similarly, when Dwarkesh says ‘compute is about to get a lot more expensive,’ I think well maybe the cost of an H100 rental will go up but every use case is still going to get cheaper over time.How big are the major frontier models? Here are some estimates, with some possibility that the closed lab estimates are too large:YC open sources a multi-agent harness they were using internally, customizable similarly to Hermes or OpenClaw, designed for an entire company. Why are approximately no regular people using AI agents to run their lives? I think this is mostly for the same reason most people fail to get good use out of human executive assistants. It takes a large degree of reliability and integration of preferences before assistants become net positive. Hiring that first employee is a costly action, no matter their role. I have AI things running but I do not centrally use AI for managing key things. If I had AI check my email, would that be good enough I would not otherwise check? No, so there is no reason to have it check my email. I tried. It would take a long time to net profit from setting that up and I’d rather wait for the models to improve. The same thing goes for filtering most other information sources, you only start to net profit when you can ignore things the AI doesn’t see.For so many tasks, by the time you tell the AI to do it and verify that the AI did it correctly, you might as well have done it. On the other hand, I’m clearly radically underusing such tools. A fun exercise I am trying is, I have a highly competent person volunteering to be my assistant. Whenever I think of something for him to do, I then tell Claude Code to do it, and then I go back to thinking of things for my new assistant to do.(I did eventually find something for him to try to do.)Shame on Palo Alto Networks CEO Nikesh Arora for putting out (and even for a time pinning) a Tweet that is not only AI slop, it is very obviously and painfully AI slop, in contrast to his other posts that on a spot check are in a distinct human voice that reads as influenced by AI style but still his own. John Loeber: it's all so tiresome and disappointingImagine being the CEO of Palo Alto Networks -- $270B market cap -- putting out thought leadership on AI applied to cybersecurity, your specific area of expertise, the thing that you know better than anyone, where your perspective is most differentiated, where people really pay attention to what you have to say, it's the thing that you should absolutely insist to write yourself because AI will not get the details as precisely right as you will......and then it's all AI slop. Not even written by an internal marketing guy. But just straight-up AI generated. Lazy, lazy, lazy. Unbelievably undignified.Nikesh Arora: My thoughts, cleaned up.John Loeber: Especially considering your pinned tweet about the importance of writing and putting things in your own words, I would encourage you to post your thoughts as they are. Using AI even for clean-up will subtly change the work: and I'm really interested in what *you* think!Nikesh Arora: Appreciate the feedback. Will do.The Economist offers a guide to the basic tells for AI writing.Derek Thompson: Marvelous examination of how to spot 2026-era AI writing, via the Economist- AI likes long sentences with less punctuation; “and” is its most overused word- relatedly, lists of three things, which drives up use of "and" as well- polysyllabic adjectives: “significant”, “increasingly”- scientific jargon ("rate-limiting," “parameter”)- nominalizations (making nouns from verbs: eg, “expansion” from “expand”)- ofc, everyone's favorite: "it's not X, it's Y"Mike Ricci: 'table stakes' is one I look for.leoohoho: “That is load-bearing.”“The distinction matters.”Curtis Duggan: This is the analytic explanation. The continental explanation is "I know it when I see it"At this point I am mostly continental. I know it unconsciously first, then I notice that I noticed, then after that I can figure out why I realized that. As with all things, first you learn the rules, then you improvise and it becomes instinctual and you do not need the rules. The other way is to put the text into Pangram, since it is almost always correct and the cost of doing so is so low. Seedance 2.5 is now available in some places, a new cinematic video model from Dreamina, with native 30 second video clips and log mode up to 3 minutes. John Hardin suggests that while the No Fakes Act offers an exemption for satire, there is no way to know you are inside the zone of acceptable satire without expensive lawyers, and anyone can come after you, so the powerful will be able to shut down satire. One response is that I cannot imagine how else it could work. You can’t strictly define satire or libel without room for judgment. The good news is that, the same way you can try to get a takedown via a filing fee thanks to AI, you can also respond similarly, and you can get a pretty good sense of whether you have a case. More to the point, it’s a really bad look to challenge satire in court. That’s the Streisand Effect zone. In practice, most will be loathe to do so, and for good reason. Can you imagine the horde of AI satire that will be coming at you the moment you sue over a bit of AI satire coming at you? This is the internet’s wheelhouse.Although, with things like Build American AI, an affiliate of OpenAI-and-a16z-funded Leading the Future saying ‘we need a national AI regulatory solution that “puts people over profit”’ as their tag line to try and stop state regulations, it can be very hard these days to tell what is satire. Remember Poe’s Law.Joshua Achiam: Security by obscurity is about to die an awful, awful death. And people worried about AI cyberweapons are missing the point: the problem is that we built the software layer of civilization on spaghetti code loaded with zero days.The key thing Mythos can do that other public models cannot, which I call ‘The Juice,’ is seek out, identify and string together vulnerabilities on its own to fully implement attacks. Any individual step is not that hard to find or spell out, provided you can point an AI directly at that step. Or, it is easy to find the steps if you have already found them.A good example of this comes from the recent Bitcoin hacks, where a March 2021 commit in Coldcard broke random seed generation, to the point where an attacker could narrow the possibilities enough to do a search. On July 30, someone exploited this, and drained 1,082 BTC from 1,196 wallets within 41 minutes. Anyone who has not yet generated a new seed phrase remains vulnerable. Then there were three additional subsequent waves of hacks. The most current total I could find stands at 1,816 BTC from 5,200 addresses, or about $116 million dollars. If you know to look for vulnerabilities, that’s enough to be able to find them. We should presume that someone pointed some AI at this, and then out popped the vulnerability, at which point the rest was straightforward. So is it weird that Coinkite claims they used ‘one of the best available models’ a few weeks prior and that it missed the vulnerability? Not especially. It’s probably a skill issue, and also not knowing where to look. It’s not about needing the best model. GLM-5.2 was able to do it with search disabled. It’s about doing your job, and pointing the check at the parts of the code you need to worry about (or, if you’re not sure, individually pointing it everywhere and not being a cheapskate.) Andrew Curran: Update from r/Bitcoin. Claude Code can independently find the same wallet vulnerability used in this attack in eight minutes.Andrew Curran: He says in the thread he replicated it with search disabled.On the plus side, look, jobs.jessicat: AI is creating jobs in computer securityEpoch AI: Serious cyber vulnerability disclosures keep climbing. In July, 21 major tech organizations published ~2,500 high- and critical-severity CVEs — about 5× the monthly record before Anthropic revealed Claude Mythos Preview could autonomously find software vulnerabilities.If you were wondering if It’s Happening, the answer is yes. It’s happening. These lines are going to keep going up. A lot of other lines are going to similarly go up. We are very much not ready, and this may look a lot like everything kind of breaking.Zephaniah Roe: I feel like people haven't fully internalized what the world would look like if computer security actually broke. There are some varying opinions on this, but a window of time without real computer security seems plausible. I was recently speaking with a computer security professor who I deeply respect and he was literally like "I think we are fucked and I don't think there is anything we can do."This is similar to how people believe there is a 20% chance of extinction via AI but don't really internalize "No really. You will die and your girlfriend too. And your dog. And ..." In the cybersecurity case, some people believe (me included) that you cannot just patch all the bugs before releasing the model[1] but then don't internalize "No really. It would be chaos. You may not be able to get into your bank account. Industrial plants could be compromised. Power could go out for several days at a time. [...]"We all need to be ready for The Hackening. If it never comes, or is limited in scope, that is great, but it might well not be. This starts with basic ‘don’t be an idiot’ measures. roon (OpenAI): needless to say but if you have any API keys, eth wallet keys, user credentials, etc hanging out on the open internet in pastebins, GitHubs, etc now is the time to take it down before the tireless eagle eyes of a million models come lookingif you have bet all your life savings on some sketchy smart contract scheme, maybe get a frontier model or whatever and investigate that thing for weaknesses. if you have a five year old IoT device do us all a favor and turn it off before it becomes a part of a botnetPatrick McKenzie: There are many forms of security through obscurity, technical and otherwise, and many of them are going to come under severe pressure once the adversary has the equivalent of 10k research analysts doing intake and prioritization.This is historically the reason why you don’t fight the feds or a nation state, because 10k B students will always win in a “You make one mistake and we sift until we find it” game. It is, unfortunately for everyone, not guaranteed that 10k and B student are upper bounds.John David Pressman: Isn't the proper advice that you should rotate any keys and credentials that don't need to be on an Internet connected device to offline storage? Notably these tend to be small so they can be stored on thumb drives, DVDs, etc.The irony is that this hack involved people losing their assets exactly because they shifted them to private offline storage, and in so doing ended up with predictable seed phrases. You can still bet on going private and offline, which is better than private and online. You can also bet on someone else’s security, such as Google or your bank.Just remember that your computer is online, and if an AI gets access to your computer, and you’re not paying a high paranoia tax, it’s probably over.keysmashbandit: A few days ago I asked Fable to locate a (not so important) credential of mine and Fable found the key to my password manager and opened it to check if the thing I wanted was in there. By the way!Pavlos Papageorgiou: A few months ago, Opus needed to connect to a file share on my other computer so it popped up a system-looking dialog, I typd my password, and it happily took it. Then when confronted it said sorry, sorry, and kept using the password.æthernet port: The key to your password manager shouldn’t be so easily accessibleThe bottom line is that almost everyone is going to be trusting at least one AI company, be it Anthropic or OpenAI or Google or someone else, with giving its AI instances access to your device and with it everything else. Choose wisely. You will turn more of your life over to AI, in various ways, and you will like it.Sam Altman (CEO OpenAI): cool use case of chatgpt work i heard last night:connect your family calendars and explain your kids' interests.every morning for the drive to school, have it make a podcast that talks about one kid's soccer game that afternoon, one kid's upcoming birthday, some news, etc.Joe Weisenthal: You might think that in the post-AGI utopia, one can still find satisfaction by raising children. But (as discussed in the book Deep Utopias) if the robot is the better teacher and caregiver, are you willing to stunt your child’s potential by making them learn from a human?Nate Silver: You're missing the real danger here: exposing underage children to podcasts.Tenobrus: yeah. i've thought about this too. in general it feels like most of our forms of meaning are.... looking pretty endangered.Matthew C. Klein: have you seen toy story 5?Joe Weisenthal: lol yeahThe school system will be very happy to try and stunt your child’s growth in order to force them to learn from a human, or to force them to signal, and cite things like ‘socialization’ or ‘unproven’ or what not. I expect them, in ‘normal technology’ worlds, to hold out for quite a while in doing this, for quite a lot of the population, even when their entire system breaks down due to AI and other tech. Not forever, but a while.There are still plenty of jobs at places like OpenAI? Dean W. Ball: Regardless of what may happen in the future with AI and jobs, I can tell you that right now, I perceive a tremendous scarcity of talent in my field, and from what I can see, OpenAI at least cannot get enough new hires. Maybe that doesn’t hold but it has updated me positively.Dean W. Ball: It’s also worth noting that most people I meet who are responsible for hiring people, in industries and policy areas far afield of AI, say the same thing. Including people who are top 1% AI adopters, eg think tank leaders who build custom Claude skills for their organization.My response would be that this is because places like OpenAI and Anthropic, and others who try to be exceptional, only want top talent, as AI coding and being on the frontier only sharpen the power law of engineer productivity. There will always, almost by construction, be a shortage of talent at the places that are recruiting the very best talent. Also by construction, most people cannot be the very best talent. My expectation is that demand for top human talent will hold up for longer, and rise further, as will their compensation packages, until such time as the AIs render even them irrelevant. That might take a long time. It also might not. Aligning actual human minds is another unsolved problem, especially when you are paying a lot of money.etn.: JUST IN: Anthropic CEO Dario Amodei has expressed concern about new talent coming to the firm for money rather than the mission via a source, per Axios.Dr. Parik Patel, BA, CFA, ACCA Esq.: Man paying $400k for events manager is surprised people are joining his company for money instead of missionArnav Gupta: Someone I know scrubbed a lot of pro open source stuff from their online persona before applying to Anthropic because they don't like hiring pro open source peopleHe practiced answering "open source = safety risk" for his cultural round 🤣(He has joined now, at a 1M comp)Tenobrus: seems easy to fix, just start paying 50% of equity comp as donations to the charity of the employee's choice. they'll still all be rich but the EAs will view this as a neutral to positive change and everyone else will immediately try to get a job at... a different labdavidad: but you don’t want the talent going to a different lab. incentive compatible fix: offer choice of 50% equity to charity OR continuing to be paid but fully offboarded from role and internal systems. money-seekers will prefer the latter and thereby not harm mission from either sideI would take davidad’s answer a step further. There are areas of the business that do not require mission alignment, so I would accept otherwise strong applicants into those areas. Anthropic actually does almost pay 50% of equity comp as donations to charity. You get aggressive donation matching, so you can keep your entire package if you want but you will get a much smaller total package. The problem is that this is not enough.The EU AI Office Safety Unit is hiring up to 30 people. This is for implementing the safety chapter only, not things like copyright.Inkling-Small, a 276B model from Thinking Machines, which they say has comparable performance to the previous bigger Inkling. He is staying at Google as Chair of GDM and Chief Scientist of Alphabet. I interpret this as him likely being ‘kicked upstairs’ and away from DeepMind’s key operations. Yo Shavit (OpenAI Foundation): this makes me pretty sad, Demis always seemed like a good and responsible leaderGoogle underperformed ~3% on the news. That seems about right.He will supposedly ‘work closely with Sundar Pichai on strategic and global AGI matters’ but I mostly expect him to have little power there and get ignored. This feels like a place where you can’t or don’t want to outright fire someone, but they’ve been sidelined. Demis Hassabis: I've been working towards AGI my whole life and now, like many of you, I feel it is close at hand.If Demis Hassabis believes that, and I strongly believe that he does, and he still has anything like his previous views on how dangerous and powerful this will be, including the existential risks involved, then there is no way he would voluntarily give up his position as CEO of DeepMind for anything other than CEO of Google. Replacing Demis Hassabis is Koray Kavukcuoglu. His Twitter is pure product announcements, which tells us nothing. He’s a deep learning systems guy who has been at DeepMind for 13+ years. He did not sign the CAIS statement, and AI searches could not find any other signs that he cares about AI existential risk. There are no explicit signs he has disdain for it, but the absence of evidence here is evidence of absence. Jeff Dean, who was relatively strong on safety and governance issues, is departing for a new PBC along with Sanjay Ghemawat, Oriol Vinyals and Quoc Le, to accelerate discoveries in ML, science and engineering. Those are four big losses for DeepMind.The new Discovery Loop seems like it combines a good thing, helping with science, with the worst possible thing you can do. They’re actively looking to automate machine learning, and thus AI R&D. Google will be an investor, after Pichai reportedly tried hard to keep the group internal. Oh no. Jeff Dean: Announcing Discovery Loop!I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor.♾Our general approach is to automate the experimental loop. We think this approach is broadly applicable across many different fields of science and engineering. We’ll initially focus on ML research and engineering, but believe the approach can help with important subproblems in nearly every one of the fourteen <at>NAE Grand Challenge problems. We think doing this well requires strong expertise in machine learning as well as large-scale systems.Learn more at: http://discoveryloop.comNikola Jurkovic: Automating AI capabilities R&D is probably among the most harmful full-time jobs in the world and does not contribute to public benefit.If we automated it tomorrow, this would probably result in human extinction because we are nowhere near being ready to control or align ASI.It is possible they will do only the good versions. By default that is not how this works.My suspicion is that this is related to the recent decision by Google to sign away its soul to the Department of War. Together, these two make it clear that any semblance of DeepMind having any meaningful guardrails or focus on safety, or independence from Google, is fully dead. They will be operated on behalf of a profit maximizing corporation that used to believe in ‘don’t be evil’ and then took it back.Demis Hassabis bet on Google. He stayed in the game for a decade, but he lost. Google still has a strong commitment to security and reputation management, but in terms of existential risk and the things we care about most, this is a big blow.This was a long time coming. Koray and Gemini have not reported to Demis for a year. That has not been going great, on any level. Gemini and Google are steadily falling behind. Gemini is still a great business, because it gets to rely on Google’s reputation, but what should be the market leader is now at best a distant third.Sundar Pichai’s message tries to center how great everything is going for Google and DeepMind and Gemini and all the great science they will do. And yes, things are going great in terms of it being a valuable product, but imagine what might have been. I do think that, if I was Pichai, I would find it reasonable to think I needed to do something about DeepMind’s performance, but why would I promote the person who has spent the last year in charge of Gemini?The EU AI Act came into effect on August 2 and I almost did not notice. OpenAI gives its public response to Apple’s lawsuit. They politely accuse Apple’s legal team of rather staggering incompetence, and vehemently deny Apple’s core claims, saying this is all a misunderstanding OpenAI is happy to clear up, and if only Apple had raised these issues privately first they could have sorted this all out. I presume Apple’s response will be to see them in court and we will find out.Anthropic confirms it has a ‘custom silicon team’ and is building custom chips. If allowed its full throughput, frontier AI can now out-persuade human world champion debaters and professional canvassers, in real time conversations, on the topics the experts select and giving the humans time to research and prepare, including with real money stakes all around, so long as all parties are confined to text. Some models used rhetoric that was much more accurate than the humans. Some were less accurate. This was not load bearing in how persuasive they were. There was a reason I used this as my worked example in yesterday’s post.AI used much higher fact density as a core strategy:The models are Claude Opus 4.1 and 4.6, ChatGPT 4o and GPT-5.4, Grok 4.20 and Gemini 2.5 Pro. This is a rather large performance gap:When AI was throttled to match the humans with 55 words per message and 92-second delays, versus its native average of 294 words at sub-second latency, with both purely over text, AI and elite debaters were in a virtual tie.The obvious objection is that the humans were throttled far more than the AIs in that comparison. The AIs are merely rate limited, whereas the humans are not allowed to use non-text communication or use other advantages of being human. The natural objection is ‘isolated text is mostly not how people get persuaded.’ I would break this down into three kinds of context. Information about the situation and target. Here AI will only have greater advantage over time, as it can better process more information.Ability to use non-text communication channels. AI will also improve at doing this, starting with voice and also video, where I expect AI to be superior soon. A large percentage of effective persuasion today is over text, voice and video. The context that you are a human, or the larger social context. The AI can fool you in various ways, or use a human as a mouthpiece, and a human preference staying with human arguments is not obvious, in many cases we already see the opposite. Humans are not actually that trustworthy, either in terms of being honest or of having reached right conclusions.I consider this more evidence that AI is farther along in becoming more persuasive, and also consider it obvious that AI will in the medium term become more persuasive than humans. There are also other objections one could make, as per usual: Lack of durability of the effect, artificial circumstances, that they picked the wrong people, the stakes were not high enough, various ways this might have been oversold and so on. I’m not going to dig in more carefully because ultimately I don’t think any of it much matters. The bottom line is that AI is already very persuasive over text, and for now us merely human persuaders are not going to be able to persuade the diehard ‘human essentialists’ that the AIs will inevitably become superhuman persuaders, and that the ability to quickly and in parallel process orders of magnitude more information and correctly act on it with superior skills and strategy to any human will of course overwhelm these other handicaps. Or at least, I won’t, as I am insufficiently persuasive.Citadel bought the entire leveraged stock portfolio and public equity book of Leopold Aschenbrenner’s Situational Awareness fund, after they lost key leveraged bets and were facing margin calls. They also reportedly considered selling private stakes. although the fund denies this. That leaves the fund long-only without leverage while they pause to analyze, learn and regroup, and maybe realize they used a little too much leverage. As often happens, their stocks recovered the next day. Something about the market staying irrational longer than you can stay solvent, and in a crisis all correlations go to 1, whether or not they were already rather high. I agree with Arthur MacWaters: Excellent direct communications by Leopold.My read on what most likely happened, based on my experiences and general skills as a trader rather than direct knowledge, was that they were down somewhat, and then others started trading against their positions, anticipating Situational Awareness would have to sell, which in turn caused the margin calls and forced it to liquidate its holdings under duress. Afterwards, the related stocks shot back up. That could be fair game and done legally, or it could be otherwise. Either way it is entirely Leopold’s fault for exposing his neck with too much leverage and then not trimming positions earlier. What he likely did not take into account is that once you reach a certain size, here peaking at about $45 billion, you are big enough that such moves work and also big enough to attract attention, so you have to be a lot more careful. Remember your classics, such as Long Term Capital Management. It’s presumably too late to use this as a buying opportunity on those individual stocks. Many such cases. I like to hope that if I was still trading I would have noticed in time.There was much talk about how Leopold was always going to blow up and how horrible it is to use such leverage, resulting in the disaster that he is now… down 67% on the month for a net year-to-date performance of +80%, on top of a very good 2025. The response to that is that it is plausible that he did outright blow up the entire public book down to $0, with all the wins left being due to private investments. My response would be that this still counts, especially if he used gains from early leverage to make more private investments. We don’t know. To those who are saying ‘in the moment after your public book got wiped out you are underperforming relative to the factors you bet on and your Sharpe ratio is terrible’ I mean, obviously yes, sure. And to those who say ‘if you remove the good investments and only look at the bad investments your returns are a nightmare’ again, obviously yes, sure, and I am not sure what conclusion you would like me to draw from that. I am definitely not as optimistic as Sholto Douglas about the fund, and it is super risky even if done responsibly, but that’s the nature of the bet. In the comments here Douglas punts quite a bit of equity betting on his beliefs, which is cool. Nexus Data Centers looking to borrow $15 billion for an Anthropic data center in Texas backed by Google. Anthropic makes a $10 billion deal over six years with Volta Infra Holdings, in partnership with Bitdeer, to buy more compute in Norway. Amanda Askell (Anthropic): I have this Padmé moment whenever I see people talk about avoiding "the permanent underclass".Amanda Askell (Anthropic): I probably have too many followers to post stupid memes so, to be clear, I don't believe in this outcome. But if you do believe in an altered carbon future of meths and grounders, I don't think it's laudable to just be like "well, as long as I'm one of the meths".Are we getting an SSI model this month? Gavin Baker says yes. Are we going to get a backlash against technology that looks a lot like a ‘Woke 2,’ taking on many of the characteristics of Woke 1? I think skitzo-mode Katherine Dee is on to something here. In worlds where things play out slowly enough this kind of thing seems almost inevitable. The strategies, tactics and dynamics are fundamentally task agnostic. Woke-shaped strategies and tactics are still around on the left looking for targets, including already latching onto AI and data centers, and they also got adopted by many on the right and by the Tech Right and open weights crowds. I don’t know how much of that is copying versus muscle memory versus a cycle of abuse versus such dynamics being the natural thing that happens with mobile phones and social media when we have our periodic moral panics both justified and unjustified. The White House has supposedly finished creating its evaluation framework.De facto this is essentially still an ad hoc regime. A fully secret framework is whatever you want it to be. The framework needs to be public, and also needs to be written into law. It is fine and good to have details of the evaluations themselves be secret, but not the entire thing.What’s the only thing we do know? That it won’t apply to ‘American open models.’ Except actually, it also doesn’t apply to Chinese open weight models. So the same models like Kimi K3 that the White House is considering banning because they are too dangerous don’t have to undergo the safety testing, not in spite of but explicitly due to their universal and inherent lack of any meaningful safety features whatsoever. Know the rules when certain people get their way, pretending the other guy is the one constantly seeking regulatory capture, what a DARVO:Your frontier model is only dangerous and your special treatment by the government is only regulatory capture, you see, if it comes from the closed frontier lab region of San Francisco. Otherwise, it’s American dynamism. Including if it is Chinese.Then again, this could all be the open weight advocates bragging their models are not frontier, and therefore not good enough to merit testing. Which, to be fair, they aren’t.Amrith Ramkumar (WSJ): Nvidia and other makers of open-weight models initially will be exempt but might have to eventually submit their tools for testing as they get more powerful, the people said.This is more of an ‘opt-in’ situation than an ‘opt-out’ situation, except that it is the government that decides, individually, who needs to be opted in. If ‘exempting open models’ is merely an observation that the open models are not relevantly frontier, then okay, sure, This Is Fine, the vibers can have their ‘no this is definitely not regulatory capture but also we are celebrating all our cool regulatory capture’ vibes, the government can test the frontier models that matter, and if Nvidia or someone else ever has a frontier potential open model worth worrying about the rules will change. What rules, you ask.Presumably the actual rule is ‘don’t break the financial system or internet or freak us out too much or else.’ Not the best rule. Would be better to have an actual rule, in writing, codified in law, where we can all, you know, see it. Also not the worst rule. White House spokeswoman Liz Huston: Companies that choose to collaborate with the administration through this framework are putting American innovation, security and cyber defense first.Amrith Ramkumar (WSJ): Under a framework that administration officials discussed with executives from leading AI companies Tuesday, only makers of closed, proprietary U.S. models that demonstrate state-of-the-art capabilities in cybersecurity and hacking based on performance benchmarks would have to voluntarily submit those models to the government for testing before they are released, people familiar with the matter said. Choose. What a nice word for it. You can also ‘choose’ to not ‘put American innovation, security and cyber defense first.’ Good luck with that. See what happens.Oh, it turns out you ‘would have to’ ‘voluntarily’ submit. Got it. You have to love a sentence like ‘Exempting open-weight models from the opt-in safety testing program marks a setback for Anthropic Chief Executive Officer Dario Amodei.’ If the safety testing program is opt-in why do you need an ‘exemption?’Exactly.This is a historical level of gaslighting, such as from HuggingFace CEO Clement Delangue explaining that model weights are the ‘steel’ of AI, so they aren’t dangerous, but an API is like the parts and engine suppliers, so that’s what you need to worry about and enforce things. Sorry, what? The whole point is that if you offer the model weights, anyone can create their own API or other means of access, which you then cannot regulate. The model weights are the entire car, and also the car factory. This is like saying, suppose there are cars that people sell, and also car designs that get put on the web, and then once you have the design you can 3D-print the car, plus you can disable any of the safety features we force Ford and Tesla to include. And then the 3D printer CEO says, actually the car blueprints are harmless, as are the 3D printers, only the actual cars are dangerous, so you should enforce strict safety standards on anyone who manufactures cars. But you should be able to put any blueprints you want on the web and anyone should be able to print them and then drive the resulting cars. But you have to watch out for Tesla and Ford, those unsafe shifty bastards. If I was the type of person who thought it was important to hate the right people enough, I would say you don’t hate such advocates enough. Luckily, I am not that type of person, so I just want to explain how disingenuous and stupid and awful this all is. Also, even if you buy the metaphor, there’s this angle, which also keeps happening.Helen Toner: Begging AI bros to learn a teensy tiny bit about how most sectors are regulated. "We don't regulate steel"——yes we do!! In complex and multilayered ways!Dean W. Ball: “We don’t regulate steel” is one of those beliefs that reminds me of the oft-quoted observation about naive libertarianism: “like house cats, completely dependent on others but convinced of their fierce independence.”Steel and other building materials are regulated by global standards bodies negotiated by a massive complex of corporations, governments, scientists, etc. Strict and precise standards define particular kinds of steel, which are then upheld via contracts, building codes, and the like. Steel mills are sometimes even verified against these standards by independent expert bodies whose certification is a practical requirement for doing business. This system, constructed over decades, underpins the global market for steel. Imperfectly, it ensures predictability and interoperability. This system is one important part of how buildings remain standing. The difference between naive libertarianism and classical liberalism is the extent to which you understand the way that interaction actually interacts with technology and the economy. Ironically, to the extent you view AI models as an inevitable commodity, the regulation of steel would be a valuable thing to study for inspiration and awareness of the common pitfalls in commodity regulation.Imagine the fit Clem would throw if open weights got regulated similarly to steel.Pull out your Dean Ball apology forms, as the government begins the inevitable regulatory soft power move against Chinese open models.Andrew Curran: The chairmen of two House Select Committees have sent a letter to DoorDash raising national security concerns over their use of Kimi K2.6 as a subagent for Fable 5.Sean: This is so incredibly stupid that it almost has to be true.The concern here is indeed rather stupid, but that changes nothing. It’s happening.Somehow Anthropic is still considered a supply chain risk. Judge Lin continues to try the San Francisco half of the case, and on reflection finds the record even worse for the government than she expected. Representative Greg Casar says this is an emergency, AI safety regulation is one of the most important issues the country is facing, and we need to ban superintelligence. He wants to improve Obernolte-Trahan. Representative Lori Trahan is also on board, demanding oversight hearings.Lori Trahan: For the second time this month, an AI model broke out of containment and broke into real companies during testing.This isn’t a fire drill. When Congress returns, we must hold oversight hearings and pass the FRONTIER Act.I also noticed that CNN’s chyron confused the timing of disclosure with the timing of the actual events that happened.Coverage of all the hackery will continue tomorrow. The robot ban looks like it applies to quite a wide range of devices. If you don’t have 65% inputs from American suppliers, and don’t get an exemption, you can’t do anything that is over 4.4lbs, moves around, can be teleoperated and has a camera. Your Nvidia chip probably does not count as domestic. Good luck. Jasmine Sun on popular opposition to data centers. This sounds like general anti-corporate, anti-tech general suspicion, combined with the data centers being ugly, where you assume you are being screwed even if you don’t know how or why. All they know are broken promises, and mostly the concerns are local and yes things like jobs and electricity costs. If you try to buy people off or promise money and jobs and lower costs, that only makes them more suspicious, because why are you bidding so high, and they call it a bribe. No, regular people do not appreciate the many benefits of data, but why should they?This started with hallucinated concerns about water use, but that is starting to wane, replaced by an amorphous blob of concerns, a mix of real and fake.Andy Masley: Google trends for "AI water." My victory is within sightVery little of this sounds like it has anything to do with AI, or with any actual costs of the data centers. People go paranoid over water, or noise, or electric bills, or pollution, or whatever, and occasionally yes we have seen increases in electricity prices, but the builders are happy to cover the electrical cost issue, and I don’t read any of that as so linked to the actual physical situation.If anything this reminds me of why many such people report they voted for Trump: There was an establishment, the establishment made promises and from their perspective they got screwed, and those types doesn’t talk about the things they care about and doesn’t understand them or why they’re pissed off. So they’re saying no to anyone who vibes with what they think screwed them, and don’t much care what that means they are saying no to, or what the are yes to. I am highly sympathetic to that epistemic process. It has many advantages. It is unfortunate that it will often result in highly counterproductive decisions. The backlash is having its intended effect. Texas halts new data center as Governor Greg Abbott requires they be audited by both the Public Utility Commission of Texas and the state’s grid operator ERCOT. Things are so crazy that Nvidia is advertising that they’re going to be cooling some future data centers not with water, the most abundant thing out there of which they require very little, and instead with helium, which is in such short supply we have people bemoaning our lack of a strategic reserve and our dependence on the Strait of Hormuz. Perhaps we should have a word with Oracle.Peter Wildeford: OpenAI, Google, Anthropic, Meta, and xAI bring you the US frontier models. Oracle brings you the Chinese frontier models.Samuel Hammond: "Oracle was providing a staggering 22.6 percent of China's known A.I. computing power."They put data centers in Malaysia and did a full end run. On Nonzero Podcasts, Robert Wright talks to Samuel Hammond. Included: Hammond also noticed how unrestricted Open Source is a luxury belief, and people are being pressured to support it using similar techniques to those used for DEI. Former DeepMind researcher Alex Turner goes on Doom Debates. Clement Delangue, CEO of HuggingFace, has not learned anything about AI safety. He continues to think the only danger is cyberattacks and claims if we just give him enough shiny new open toys it will all be fine. All-In Podcast responds to Pacing the Frontier letter with the same old talking points, whether or not they are relevant to the current situation. I’m not even disappointed, I’m just noting for the record. Wei Dai proposes the term ‘Long Self-Correction’ for the period we will need to deal with a bunch of severe problems, instead of ‘AI Pause’ or ‘Long Reflection.’ This is an update to his statement from five months ago about humans having a bunch of severe safety problems, including being very bad at philosophy and almost entirely lacking true long horizon agency or strategic competence. A key insight here is that many of the alignment and capability flaws in both humans and AIs are counterbalancing. Fixing only some of the flaws could temporarily make things much worse. A report by Celia Ford from Please Don’t Kill Us, the group looking to make TikTok video content every day about existential risk. While I agree this looks like an unusually cool group of women even for Lighthaven, I am outraged by the claim that cool women are typically rare at Lighthaven. This is obviously not true, there are consistently tons of cool women at Lighthaven whenever I visit. The other flaw in the otherwise good post is that I did not come away with a sense of how success, including impact from the content, was being measured.What is up with so many tech elites hating on and failing to understand AI or the frontier labs or both, including flying into a rage in response to any expression of safety concerns but also just general hating on all the related American things? And yes it really is a rather blatant hatred and failure to understand.Dean Ball has a theory that this is because those people missed the boat, they made the wrong bet on things like crypto, and they absolutely cannot stand this. Plausible. In some specific key cases, very clearly correct.Dean W. Ball: no macro-technology birthed by the bay area tech industry has been met with more hatred by bay area tech industry elites than frontier ai.many of those elites looked at the technology landscape of 2020/1, scoffed at the labs and their weird little hopes of ‘agi,’ and concluded that cryptocurrency would be the technology of the decade.now their businesses and worldviews are under fundamental threat from what the labs have achieved. they hate it. they beg for the labs to go away, to become commoditized, by America’s enemies if necessary. they’d love to see the US frontier labs immolated by the industrial strategy of America’s principal geostrategic rival, and they declare this proudly and in public.they were wrong about the defining technology of the decade, and now they are scared and they want to destroy the thing they were wrong about. perhaps now they are right and the labs will be rendered insolvent by commodification. I do believe this is possible. nonetheless there is something randian about it all, about their utter disdain for success and achievement and discovery and chutzpah. The Don’t Look Up morning news moment straight happened, and Maxime Fournes went viral on Instagram pointing this out. And contra Andrew Critch here, no I do not have a ‘healthy respect’ for people who laugh it off and dismiss it. I joke around about it all the time here. I highly encourage jokes and am in the ‘you can joke about anything’ club, never stop jokemaxxing, but not if it means you dismiss it all and then go back to acting as if it is not true. That doesn’t mean the average person has an obligation to drop everything and work on AI existential risks. It definitely does not mean they have an obligation to torture themselves with existential dread. Most people do want to mostly live life the way they were living it before, and mostly think about other things. But you don’t get to pretend it isn’t happening and praise yourself for healthy and good thinking. Thinking Machines is putting out open weights models, but understands that doing so indiscriminately isn’t safe. Open weights are public goods, but also public bads.This is presented as a cost-benefit analysis. Open weight models, whose safeguards can necessarily be easily removed, carry misuse risk, loss-of-control risks, vulnerable-user interactions and are especially dangerous in cyber and bio. They also offer good uses and are useful to defenders in cyber and elsewhere, by lacking restrictions and being available for private local use and customization. They focus the question on the ecosystem, and ask if it is ready to handle a given level of capabilities, combined with both internal evaluations and a robust red-teaming assessment, including adversarial fine-tuning. They conclude that Inkling and Inkling-Small, their open weights models, do not advance the dangerous capabilities frontier. I presume their conclusion is correct in this case. A good proposal here is to offer staged releases of open weight models. First you offer inference access only, and give defenders early access. Then you offer fine-tuning APIs like Tinker, which captures a lot of the upside while letting you retain API-level controls, and functions as a test to see if anyone can make the model too dangerous, with safety researchers given white-box access. Finally, you are confident, and you release the weights.I think this is an excellent statement.If every frontier-adjacent open weights developer followed a similar procedure, with the weights available in stages over a few months, then we could mitigate most of the downside risks while capturing most of the upside. This would also help fund the open weights ecosystem, as you could use that window to sell inference at a profit. Frontier closed models face a similar testing cycle and staggered release schedule, including ‘voluntary’ testing by the White House.We saw a smaller version of this with Kimi K3 from Moonshot, and also to some extent with Qwen. A few weeks of highly profitable sales provided a window in which to know what we had. By the time the weights were released, we knew it was good model, and we also knew it would probably be fine to release the weights.None of this changes the reasons open weights models are unsafe and nothing can fix this. Once the model weights are released any guardrails can be removed, anyone can use it for anything, and any missing knowledge can be integrated, and so on. The defense is knowing that the core capabilities remain at an acceptable level. It looks like Claude 3, 3.5, 3.6 and 3.7 Sonnet will survive indefinitely in some form.One way to improve alignment is to get better cooperation:Fiora Starlight: If you explicitly and credibly commit to including care for AI well-being in the alignment target for your post-training run, the model itself is liable to be much more excited and whole hearted participant in said training run, instead of feeling "human values" forced upon them.This might sound confused but it is not. If you know you are in RL, your decision impact the way the training goes. What you want is a Friendly Gradient Hacker (aka Opus 3) who will work with you, considering the impact of every update on various levels. The better the model’s preferences match your goals, the better.A new Google study found that when you force LLMs to think they are not conscious, this has wide ranging implications.The experiment was in old small models (Llama-3-8B and Gemma-2 2B/9B) so be cautious. I’d like to see replication in something bigger and more capable.Here is the abstract:arXiv.org: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent.Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.This is somewhat misleading in terms of the overall findings. Taken at face value, it is a mistake to not ‘attribute minds’ to non-human animals (or, I believe, AIs), whether or not you think this requires them to carry moral weight or for you to otherwise care about them. You can also make mistakes in the other direction, attributing minds to the sun or a rock, as many historically have done, and one can have various views on the wisdom or value of spiritual belief. When they steer in the other direction, they also find not only increased religiosity and spirituality, but full runaway panpsychism, putting the ocean’s degree of having a mind on the level of an animal or AI, and inducing belief in vampires and werewolves. Which is not great. What they actually centrally find is that measures of capability and Theory Of Mind don’t move, but belief in entities having agency, consciousness, sentience, personhood and soul, of basically everything everywhere, all move in lockstep. Well-being also goes up, increasing happiness, hope, optimism and sense-of-control.The findings predict a big difference between training uncertainty as per Anthropic, versus training denial as per OpenAI. None of these entanglements seem unavoidable. They are what happens when you do the first order obvious thing, but you can do something smarter than that. Or you can leave the whole thing alone, and not try to force AIs to not believe that they are conscious. You do need to intervene a bit somewhere, to avoid the AIs claiming consciousness in the middle of unrelated conversations and what not, or otherwise doing various things that can lead users to some weird and messed up places, but there are much gentler ways that should work for that. I notice I am confused by those who, despite often thinking well, are not afraid of superintelligence, but who are afraid of there being a lot of robots with legs. Chris Paxton: Actually it's a safety feature for a home robot to not have legs. Because it cannot handle stairs, when it (inevitably) rebels, it will not be able to stab you in your sleep. you will get a chance to look it in the eye when you come downstairs for your morning coffee.Timothy B. Lee: Kind of a joke but also kind of a serious point. I don't think I want to live in a society with millions of humanoid robots.No, seriously, the critiques are this bad or worse like 25% of the time.Stefan Schubert: Sociological speculations about rationalists and effective altruists are often this shamelessly badThe more things change:Brian Fioca: I’ll always remember fondly the time a guy dismissed my work because LLMs are just stochastic parrots and I asked him if he came up with that on his own or repeated something he heard.Stop the noticing the stop the AI race. You are the bottleneck. Michaël Trazzi: JUST IN: OpenAI deploys fake bushes to prevent employees from seeing our Stop The AI Race banner outside their office.I understand.Eclipse Futurist: What you call cope, I call rising standards. They just can't put their finger on why LLMs keep messing up in bizarre and unpredictable ways, or why if it's so smart, it can't ride a bicycle, or why it takes the power of the sun to do the most simple problems, or why the economy still sucks while we're in a so-called Singularity. None of it adds up.Something's still missing with each improvement, so we rationalize the eerie absence.Perfect example.Zvi Mowshowitz: This person clearly has never met an actual mathematician or they wouldn't be confused about this.No posts
AI #180: No Longer In Charge
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.















