IBM has built the first large-scale inference cluster for Together AI on IBM Cloud, leveraging NVIDIA HGX B300 systems, to support enterprises in rapidly and efficiently deploying their AI workloads to production.
IBM announced a collaboration with Together AI to provide IBM and NVIDIA AI infrastructure. Under a multi-year, $240 million agreement between IBM and Together AI, IBM will deploy a large-scale cluster of NVIDIA HGX B300 systems on IBM Cloud, scheduled for availability in the first quarter of 2027. Together AI will leverage this cluster to provide inference services for open-source models. This deployment marks the first large-scale inference-dedicated cluster built on IBM Cloud using HGX B300 systems and NVIDIA Spectrum-X™ Ethernet networking. According to NVIDIA, this configuration is designed to deliver 30 times the processing power of the AI factory compared to the previous generation.
This collaboration aims to help Together AI deliver higher performance and superior token efficiency as companies efficiently scale their AI adoption. Together AI operates on the belief that open-source models are essential to the future of AI, and that developers should be able to develop using an open and modular stack. The company recently raised $800 million in Series C funding at a valuation of $8.3 billion to expand its AI Native Cloud. Its platform provides capabilities across inference, training, fine-tuning, and agent-based workflows. The company also states that its inference services are growing rapidly and are currently processing 400 trillion tokens per month.
Also Read: Firebird Opens Armenia AI Factory for Global Growth
Together AI selected IBM and NVIDIA because of their innovative product roadmaps, their ability to provide GPU capacity at the pace needed to rapidly scale AI, and their ability to deliver low token costs. Leveraging IBM’s expertise in delivering enterprise cloud capabilities, this collaboration aims to support Together AI’s continued growth in the enterprise market and make open-source AI more accessible to developers and businesses worldwide.
Vipul Ved Prakash, CEO of Together AI, stated: “Enterprises are looking for superior performance from cutting-edge models without the high costs of closed models. This is only possible if the underlying infrastructure is fast, reliable, and scalable at scale. Our collaboration with IBM and NVIDIA provides us with that infrastructure. This cluster will enable us to deliver production-level inference to more enterprises more quickly, and represents a significant step forward in our efforts to make open-source AI a clear choice for businesses.”
Source: PRTimes


