A multi-year deal with the startup Together AI will put Nvidia Blackwell systems on IBM Cloud, and the wager is that enterprises now care more about the cost of running AI than the prestige of the model doing the running.
IBM has decided that the money in artificial intelligence is no longer only in building the cleverest model, but in running it cheaply, and it has put $240m behind that conviction.
The company has signed a multi-year agreement with Together AI, a San Francisco startup, to stand up a large-scale AI inference cluster on IBM Cloud, and the pitch is squarely aimed at enterprises trying to trim their AI bills.
The cluster will run on Nvidia HGX B300 systems, built on the Blackwell architecture that Nvidia markets as tuned specifically for inference, knitted together with the company’s Spectrum-X Ethernet networking, which is the same recipe that a wave of specialists have been chasing as they bet that AI’s real profits lie in cheap inference rather than in ever-larger training runs.
Together AI is an interesting partner to pick. Valued at $8.3bn as of July, its platform lets companies train and run workloads on open-source models, including DeepSeek, MiniMax and Kimi, and it sells itself as a cheaper, more flexible alternative to the closed, proprietary systems that dominate the headlines.










