(Image credit: AMD)
The demand for AI compute is already insatiable, and it seems only poised to grow in the wake of the introduction of frontier-class open models like Kimi K3 that anybody can potentially fine-tune and serve. Against this backdrop, Microsoft and AMD are teaming up to get Redmond more AI FLOPS for both internal and external use. The two companies announced this morning that Microsoft will commit to adding AMD's Helios rack-scale AI accelerator in volume to run frontier-model workloads in its own data centers, as well as for Azure AI infrastructure customers and services.Go deeper with TH Premium: AI and data centersThe partnership makes next-gen AMD AI compute available to Azure customers like AI labs for AI training and inference serving workloads, and it’ll also underpin managed compute for enterprise customers looking to deploy AI workloads through Microsoft Foundry.The two companies didn't indicate the exact size of Microsoft's Helios deployment in either watts or dollars, but the commitment would seem to be another major win for AMD as it seeks to grab data center GPU share from Nvidia. AMD has struck massive partnerships with OpenAI and Meta in the past year with gigawatts of compute installations and hundreds of billions of dollars potentially hanging in the balance.For a quick refresher, the Helios rack-scale accelerator will take the fight to Nvidia’s Vera Rubin NVL72 system when it arrives later this year. Helios joins together 72 next-generation Instinct MI455X GPUs with an aggregate of 31.1TB of HBM4 memory capacity across the system. Those GPUs offer as much as 1.4 exaFLOPS of FP8 compute and 2.9 exaFLOPS of FP4 for AI models using those OCP AI data types.AMD is targeting 260 TB/s of scale-up bandwidth within the rack, on par with Nvidia’s Vera Rubin NVL72 rack-scale system, and 43 TB/s of scale-out bandwidth using UALink over Ethernet, or about twice that of Vera Rubin, although the performance of UALink over Ethernet in practice remains to be seen.Microsoft and AMD also announced that Azure will add two new VM series built on AMD's upcoming sixth-gen Epyc Venice CPUs: the HDv2 series for "agentic AI and data pipelines," and the HXv2 for semiconductor design workflows. Microsoft will also leverage its existing deployment of AMD Pensando DPUs to integrate that hardware into its Azure Boost offerings to accelerate networking and storage processing operations.Tom’s Hardware will be on the ground at AMD’s Advancing AI event this week, where we expect to learn more about AMD’s AI ambitions for the second half of this year and beyond. Stay tuned for our coverage from that event.Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.










