From metal to model: Private cloud gets an assembly line for production AI

Enterprises moving artificial intelligence from pilot projects into production are discovering that the hard part is no longer the model. It’s the infrastructure beneath it. Cost, tokenomics, data privacy and the manual labor of stitching together graphics processing units, servers, networking and software stacks have become gating factors for deployment at scale.

Those pressures are pushing more production AI workloads back into the data center, where organizations can keep models close to their data. Broadcom Inc. is betting that turnkey automation can simplify that shift, according to Prashanth Shenoy (pictured, right), chief marketing officer and vice president of the VMware Cloud Foundation Division at Broadcom.

“A lot of our customers are looking at private cloud in an on-premises environment to deploy their production AI workloads at scale. But as they’ve been trying to do this, it’s been a very complex process from what we call the metal to model,” he said. “Setting up GPUs, servers, networking, Kubernetes, containers, AI software stack, testing, validating which models to use. It’s an extremely manual and complex process.”