There comes a point in every guitarist's life when the digital modeller gets unplugged and the hand-wired amplifier takes the stage. The modeller imitates a hundred amps. The hand-wired circuit does one thing, and it wastes zero watts pretending.Google has reached that moment in silicon.About The AuthorAt heart, I am a storyteller drawn to the watershed moments that bend the technology landscape. I braid narrative with data, humanise statistics, and trace the arc from first spark to world-changing impact. My reportage, features and reviews are witty, sardonic, visual and vivid, using anecdote to illuminate rather than eviscerate.
As a technology journalist with over sixteen years of experience, I have travelled the world and the seven seas, covered every major tech conference worth its lanyard, chronicled the defining breakthroughs of the last decade and a half, and played a pivotal role in launching some of India’s most important technology publishing platforms across web, print and TV.
In my current role as Editor of Gadgets Now Studios, I bring that experience, instinct and editorial firepower to the table, with the mandate of scaling the brand to towering heights.
When I am off the clock, I am usually lost in music, from underground electronic and progressive rock to stone-cold blues. I am also an incurable F1 nut, a hangover from my previous life as an auto journalist, and always game for a jam session with friends, where I do my best to make my guitar gently weep.The company is developing the Frozen v2, a server-class AI chip built around its own Gemini models, etching pieces of their architecture directly into the circuitry and trading a measure of the TPU's versatility for a single-minded return: six to ten times more AI tokens served per unit of power, according to a report in The Information. Deployment could begin as early as 2028. Investors delivered their verdict within hours of the story breaking, lifting Alphabet shares about 3 per cent on Monday morning, two days before the company reports second-quarter earnings. The market called it correctly. Google is admitting, in silicon, that the artificial intelligence race has become a power-physics problem, and that admission will reshape the economics of every AI data centre on the planet, including the one Google is building in Visakhapatnam.Key TakeawaysGoogle's Frozen v2 is a server-class AI chip that hardwires parts of the Gemini architecture into silicon, targeting six to ten times more tokens per watt than today's TPUs, with deployment as early as 2028.The chip answers a compute capacity crunch inside Google: Cloud declined external customer deals, the backlog crossed $460 billion, and 2026 capex guidance stands raised to $180–190 billion.Every AI giant now builds custom silicon: OpenAI's Jalapeño promises a 50 per cent inference cost cut by end-2026, Anthropic is in early Samsung talks, and Microsoft, Meta, Amazon, Alibaba and Huawei field their own.Nvidia keeps the training crown while inference migrates to custom ASICs, a category analysts see reaching about 40 per cent of the accelerator market; Broadcom co-designs for both Google and OpenAI.India carries the largest stakes: Jio's 18-month Gemini AI Pro offer, worth Rs 35,100, and the $15 billion Visakhapatnam hub make tokens-per-watt economics decisive in rupees.What Exactly Is Google's Frozen V2 Chip?Frozen v2 is a server processor that hardwires portions of Gemini's architecture into the chip itself, so the model's weights stop travelling. A conventional AI processor, Google's TPU included, keeps those weights in memory and shuttles data back and forth for every query it answers. Each shuttle costs energy. Frozen v2 hardwires portions of Gemini's architecture into the chip itself, collapsing the distance data must travel and, with it, the electricity bill attached to every token generated. The design sits alongside the TPU family rather than replacing it, and the extent of the hardwiring is still being settled inside Google, The Information reported.Google's own response to the report sounds like a confirmation wearing a laboratory coat. "Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers," the company said when asked about the chip, adding that "this rigorous exploration is central to our full stack approach" and that hardware and software are co-designed from the ground up. A company batting away a rumour says the rumour is wrong. Google chose instead to explain its philosophy of building exactly this kind of chip. Draw your own circuit diagram.The TPU line itself is moving fast. At Cloud Next in April, Google introduced its eighth-generation chips in two specialised flavours: TPU 8t for training, with three times the processing power of the Ironwood generation, and TPU 8i for inference, delivering 80 per cent better performance per dollar than its predecessor. Ironwood, the seventh generation, already doubled performance per watt against Trillium, the chip that powered Gemini 2.0. So why does Google need another family on top of that cadence?Because inference at Google's scale is a utility-bill problem, and utility bills punish generalists. Google's first-party models processed more than 16 billion tokens per minute via direct API use in the first quarter, up from 10 billion the quarter before. Every one of those tokens carries an electricity cost. A general-purpose chip spends a slice of every watt on flexibility the workload leaves untouched; a hardwired chip spends close to everything on the answer. Think of it as the signal-chain purist's rig: every pedal a guitar signal passes through colours the tone and draws current, so the player chasing the cleanest possible sound runs the shortest chain the song allows. Frozen v2 proposes the shortest chain AI has seen. Gemini becomes the circuitry.More articles by AuthorTrending StoriesWhy Has Tokens Per Watt Become The Deciding Metric?Tokens per watt decides the race because electricity, rather than silicon, has become the binding constraint of the AI buildout. Gartner forecasts that power shortages will restrict 40 per cent of AI data centres by 2027, a supply wall that has already turned grid connections into the industry's most contested asset. Hyperscalers are signing decade-long power purchase agreements, reviving nuclear plants and negotiating directly with utilities, because the chips they ordered years ago are only as useful as the megawatts behind them. The binding constraint of the AI buildout has migrated from silicon to substations.Once power becomes the ceiling, tokens per watt stops being an engineering metric and becomes the income statement. A six-fold efficiency gain means a single gigawatt of capacity serves six times the inference volume. Every data centre Google has already paid for gets revalued upward overnight, because its power-limited output multiplies. Depreciation per token collapses. Margins on every API call widen. Monday's 3 per cent pop in Alphabet shares follows from that logic: the market was repricing stranded capacity back into productive capacity.The timing is equally commercial. Enterprise buyers have started auditing the electricity embedded in every AI contract, and investors have spent a year asking whether AI spending earns its keep. Efficiency has become the sales pitch. A chip that serves the same model at a fraction of the wattage lets Google Cloud quote prices its rivals either match at a loss or decline to match at all.The Capacity Crunch That Forced Google's HandFrozen v2 arrives from a position of shortage, and the shortage is severe enough to shape product decisions across the company. Sundar Pichai has spoken on a recent podcast about being acutely constrained on compute, a constraint he said he reviews almost every week. The Information's reporting adds the operational colour: the scarcity has fuelled internal tensions at Google, and its Cloud unit has walked away from external customer deals it lacked the capacity to serve. Engineers inside the company are rationing access to AI coding tools, even as 75 per cent of all new code at Google is AI-generated and approved by engineers, up from half last autumn.The demand side of that squeeze is historic. Google Cloud's revenue grew 63 per cent year on year to $20 billion in the first quarter, and its backlog nearly doubled quarter on quarter to more than $460 billion. "Our AI investments and full stack approach are lighting up every part of the business," Pichai told investors on the April earnings call. The full stack is indeed lit, and it is lit precisely at the layer where Google can supply compute, which is why the constraint bites so hard: every gigawatt of capacity still on the drawing board is revenue the backlog has already promised.The money being thrown at that wall is staggering. Alphabet spent $35.7 billion on capital expenditure in the first quarter alone, more than double the year-ago figure and roughly $400 million a day. The company then raised its full-year guidance to between $180 billion and $190 billion, up from $175 billion to $185 billion, folding in the Intersect acquisition. At roughly Rs 90 to the dollar, the upper bound approaches Rs 17.2 lakh crore in a single year. Chief financial officer Anat Ashkenazi went further on the call: "We expect our 2027 CapEx to significantly increase compared to 2026." Spending at that altitude invites one persistent question from shareholders, and Frozen v2 is part of the answer. Efficiency is the chapter of the payback story that turns a power-constrained data centre into a licence to print tokens.Wednesday's earnings will test the appetite for that story. Analysts expect earnings per share of about $2.90 on revenue near $101 billion, with Google Cloud projected to grow around 67 per cent to $22.8 billion. Any language from Pichai or Ashkenazi touching model-specific silicon will move the stock more than the print itself.Every AI Giant Now Builds Its Own AmpGoogle's move lands in the middle of a custom-silicon arms race that has compressed chip programmes from half-decade slogs into product cycles measured in months. OpenAI introduced its first custom processor, an inference chip called Jalapeño co-designed with Broadcom, in June, promising a 50 per cent cut in inference costs after a nine-month design cycle, with deployment targeted for the end of this year. Anthropic, which trains and serves its Claude models on more than a million Google Ironwood TPUs, is simultaneously in early talks with Samsung about a custom chip of its own. Microsoft has Maia 200 built on TSMC's 3-nanometre process running its newest models. Meta, Amazon, Alibaba and Huawei all field their own silicon, and ByteDance has held talks with Qualcomm about doing the same.CompanySilicon ProgrammeClaimed GainTimelineGoogleFrozen v2, model-hardwired, alongside TPU 8t/8i6–10× tokens per watt (reported)As early as 2028OpenAIJalapeño inference ASIC with Broadcom~50% inference cost cutDeployment end-2026AnthropicEarly custom-chip talks with SamsungPrivateNegotiation stageMicrosoftMaia 200, TSMC 3nmPowers newest GPT-class modelsIn deploymentMeta / Amazon / Alibaba / HuaweiMTIA 300–500, Trainium, Zhenwu M890, Ascend 950DTVaries by programme2025–2027 rolloutsThe Anthropic thread is the one worth pulling. Alphabet owns about 14 per cent of Anthropic, sells it more than a million TPUs, and now watches it shop for a foundry in Seoul. Google is arming a competitor, profiting from the arming, and racing that same competitor towards the efficiency frontier, all in the same quarter. The Jalapeño precedent shows how quickly a hardwired bet can travel from whiteboard to wafer: nine months from design start to deployment target, a cycle time that would have embarrassed the industry five years ago and now counts as table stakes. Google choosing 2028 points to something deeper than a derivative design. The company is re-architecting the relationship between model and metal, rather than tuning one for the other, and re-architecture consumes the years that tuning skips.Broadcom has become the silent co-author of the custom era, pairing its SerDes and packaging expertise with whoever walks in carrying a model and a grudge against Nvidia's gross margin. Marvell plays a similar role further down the menu. The barrier to bespoke silicon has collapsed from a billion-dollar decade to a design engagement, and every foundation-model company with enough traffic to amortise a tape-out now treats a chip team the way it once treated a research lab: as core infrastructure. The scramble is a tribute to how completely inference economics have eaten the industry's attention.Nvidia Keeps The Training Crown And Watches The MeterThe second-order casualties and beneficiaries sort into two camps. Nvidia still commands close to 80 per cent of AI accelerator revenue, a share built on training workloads, the CUDA software moat and networking that stitched a generation of clusters together. Training stays loyal to that stack. Inference, the volume business of the AI era, migrates towards whatever custom silicon each hyperscaler etches for its own models. Every Jalapeño, Maia and Frozen that ships is inference volume that bypasses an Nvidia invoice. Analyst estimates already see custom ASICs climbing towards a 40 per cent share of the accelerator market over the coming years, and Google's 3 per cent rally suggests investors are starting to price the divergence: the general-purpose era's champion selling into a market that is going bespoke.Broadcom's position is more elegant and more exposed. It co-designs Google's TPUs and OpenAI's Jalapeño, so it collects rent from the custom wave on both sides of the fiercest rivalry in AI. Yet efficiency carries a paradox. A chip that serves six times the tokens per watt means hyperscalers need a fraction of the sockets to serve the same demand, and chip vendors sell sockets. The historical counterweight is Jevons paradox: every efficiency gain in computing has expanded consumption faster than it shrank unit demand. Cheaper tokens mean more tokens, more agents, more video, more everything. If Jevons holds, Nvidia, Broadcom and the ASIC wave all win together, and the only loser is the status quo cost of intelligence itself.The Gemini Delay Hanging Over The SiliconOne awkward fact hangs over the whole design. The weights Google wants to etch into metal belong to a model family still being rebuilt. Gemini 3.5 Pro, introduced at Google I/O in May and promised for June, has missed its window as engineers wrestle its coding performance up to internal standards; Alphabet shares slid about 4 per cent when the postponement surfaced. "We're shipping quickly across a wide range of models while keeping them highly cost-effective for customers," the company said in response, adding that 3.5 Pro is in partner testing. The slip leaves Google fronting a leaderboard fight with Gemini 3.1 Pro while Anthropic shipped Fable 5 in June, OpenAI shipped GPT-5.6 Sol earlier this month and Moonshot AI released its 2.8-trillion-parameter Kimi K3 this week.Set Frozen v2 against that backdrop and its strategic shape sharpens. Frontier models are converging in capability with every release cycle, and convergence turns the model layer into a commodity fight. The durable moat migrates downward, into the cost of serving intelligence rather than the possession of it. Google has spent a decade insisting its full stack is the advantage; Frozen v2 is that doctrine taken to its logical extreme, a chip whose reason to exist is making the economics of Gemini superior to any rival's even on days when Gemini itself is second to market. It is a hedge dressed as hardware, and the 3 per cent rally says shareholders prefer the hedge to the wait.The engineering wager inside the hedge is the sharpest edge of the whole story. Models mutate every few months; silicon takes years to etch. Hardwire the wrong primitives and Frozen v2 ships as a museum of Gemini 3's assumptions, beautiful and obsolete. Google is therefore betting that the deep grammar of its architecture, the attention patterns and data flows that survive every version bump, has settled enough to cast in metal. The codename carries the thesis in plain sight: freeze the model, beat the meter. Rival labs iterating on general-purpose chips keep their flexibility and pay for it in watts. Google keeps its edge and pays for it in commitment, which is a fine trade right up until the day the architecture shifts underneath the silicon.What Does Frozen V2 Mean For India?India feels tokens-per-watt economics more personally than any other market, because India is where Google has placed its two largest AI bets. The first is distribution: Jio users are being offered 18 months of Gemini AI Pro at zero cost, a package worth Rs 35,100, which amounts to the single largest distribution subsidy in the history of consumer AI. The second is infrastructure: a $15 billion AI hub in Visakhapatnam, around a gigawatt of capacity built with AdaniConneX and Airtel, which broke ground in April, alongside Reliance Intelligence's role as Google Cloud's strategic TPU partner for Indian enterprises. Subsidised users generate inference forever. A subsidised army of hundreds of millions is a permanent electricity bill, and the only mechanism that converts that bill from burn into margin is exactly the kind of hardwired efficiency Frozen v2 promises.The domestic numbers sharpen the point. For Indian enterprises, the efficiency race lands as a procurement question with a direct line to margins. A startup serving vernacular-language agents at scale cares about one number above every benchmark: the rupee cost of a million tokens delivered. If Google's hardwired economics flow through to Cloud pricing by the end of the decade, Visakhapatnam becomes the cheapest place on earth to serve intelligence at population scale, and every Indian company building on Gemini inherits that cost curve as a birthright. If the economics stay in America, India remains a market that subsidises adoption today and pays rack rate tomorrow. IndiaAI Mission has provisioned more than 38,000 GPUs for the country's startups and researchers, with 20,000 more announced, yet India's data centre builders wrestle power costs, grid queues and land constraints that American hyperscalers rarely face at the same intensity. Efficiency is how a gigawatt-starved market serves a billion-user appetite. Cricket maps Google's play most cleanly: the Jio offer is the Powerplay assault, runs on the board early with wickets risked, and the Visakhapatnam hub is the long innings behind it. Frozen v2 is about the economy rate. In a market where the customer pays in rupees and the grid bills in gigawatts, the side conceding the fewest watts per token wins the match, and every rival, from OpenAI's Indian pricing pushes to Anthropic's enterprise courtship, is about to be judged on that figure.Alphabet reports on Wednesday, and the scoreboard gets its next update. Watch the capex line for direction, the Cloud backlog for appetite, and any syllable Pichai spends on model-specific silicon for confirmation. The chip itself arrives in 2028, but the race it confesses to is already mid-innings. The crowd just started counting the run rate and the economy rate in the same breath. FAQsWhat is Google's Frozen v2 chip?Frozen v2 is an in-development Google server processor reported by The Information in July 2026. It hardwires portions of the Gemini model architecture directly into silicon, cutting the energy spent moving data for every query. Google has confirmed a broad hardware-software co-design research programme while stopping short of confirming the chip itself.When will Frozen v2 launch?Deployment could begin as early as 2028, according to the report. The scope of the hardwiring is still being finalised inside Google, so the timeline carries two years of design risk before a single wafer ships.How is Frozen v2 different from Google's TPU?TPUs are versatile accelerators spanning training, in the TPU 8t, and inference, in the TPU 8i. Frozen v2 sacrifices a slice of that versatility by etching Gemini's architecture into the metal, targeting six to ten times more tokens per watt. The new family will sit alongside TPUs rather than replace them.Does Frozen v2 threaten Nvidia?Training workloads remain loyal to Nvidia's CUDA stack, which holds close to 80 per cent of AI accelerator revenue. The exposure sits in inference: every custom chip Google, OpenAI or Microsoft ships routes serving volume around an Nvidia invoice. If cheaper tokens expand demand faster than efficiency shrinks socket counts, the dynamic economists call Jevons paradox, both camps grow together.What does Frozen v2 mean for India?India hosts Google's two biggest AI bets: the Jio Gemini AI Pro offer worth Rs 35,100 per user and the $15 billion Visakhapatnam data centre hub of around a gigawatt. Hardwired efficiency decides whether subsidised inference for hundreds of millions of users becomes margin or burn, and whether Indian startups eventually buy the cheapest tokens on earth.end of article










