Off-Prem

Deal to fund 'large' deployment of Nvidia's last-gen HGX B300 systems launching in Q1 2027

With power, datacenter capacity, and other supply chain constraints, AI infrastructure is in short supply, and service providers will rent all the compute they can get, even if it means working with competing cloud providers.Together AI is the latest example. This week the service provider announced a $240 million deal to run its open weights inference platform on a “large cluster” of Nvidia GPUs housed in IBM Cloud.“Together AI selected IBM with Nvidia because of their innovative product roadmaps and their ability to deliver GPU capacity at the pace required for rapid AI scaling and lowest token cost,” IBM's announcement says.

Translation: Together AI tapped IBM because Big Blue had the capacity it needed when it needed it.

Together AI sits toward the top of the AI inference ecosystem. Its business model largely revolves around renting compute from cloud or neocloud providers, many of which themselves rent floorspace and capacity from bit barn operators and operate competing services. More recently, the company has begun deploying GPU compute in datacenters in Maryland, Memphis, and Sweden, but like most AI service providers, its main value proposition remains making it easier for users to consume the hardware for applications like inference, fine-tuning, and training. For inference, this largely boils down to an OpenAI-compatible API endpoint.The service provider isn’t particularly picky about what hardware that API endpoint runs on top of either, so long as price performance is favorable. For example, Together AI will be deploying services atop SambaNova’s heterogeneous compute platform built in collaboration with Intel and leveraging Nvidia GPUs for prefill processing. Those systems went live in Vector Core Compute’s new AI bit barn earlier this year.