IBM is putting $240 million behind a bet that enterprise AI inference, the process of actually running trained models at scale, is where the real money is heading. The company signed a deal with Together AI to deploy a large-scale inference cluster on IBM Cloud, with full deployment targeted by 2027.
The partnership pairs IBM’s enterprise cloud infrastructure with Together AI’s GPU-accelerated platform, which specializes in high-performance model training and inference. Together AI claims its infrastructure can deliver up to 2x faster inference times compared to standard setups, along with meaningful cost reductions.
Why inference matters more than training now
IBM’s move here is essentially an infrastructure play. Rather than trying to build the models themselves, they’re positioning IBM Cloud as the place where enterprises run them.
Together AI, founded in 2020, has carved out a niche as an infrastructure layer for open-source AI models. The company provides the tooling and optimized compute that lets organizations deploy models without building their own GPU clusters from scratch. For IBM, bringing that capability in-house means enterprise clients get access to high-performance inference without leaving the IBM ecosystem.










