Red Hat has cooked up Red Hat AI Factory with Nvidia, a co-engineered platform that mashes Red Hat AI Enterprise and Nvidia AI Enterprise into one stack for building and scaling AI apps.
It is pitched at outfits trying to get past pilot projects and into production, where the bills land and the infrastructure arguments get loud.
The platform stitches together Red Hat’s enterprise Linux and AI tooling with Nvidia’s accelerated computing software, running on systems from Cisco, Dell Technologies, Lenovo and Supermicro.
Enterprise AI spending is forecast to exceed $1 trillion by 2029, with agentic AI pushing constant inference and turning GPUs into the new office politics.
That rising demand is squeezing infrastructure, particularly GPU capacity and model-serving efficiency, which is where the pair reckon they can become indispensable.
The stack is meant to span on-premises, cloud and edge environments while supporting high-performance inference, model tuning, customisation and agent deployment with centralised management.
It ships with pre-configured models, including IBM Granite, Nvidia Nemotron and Nvidia Cosmos open models, delivered as Nvidia NIM microservices.
If your data is too sensitive to share, you can refine models using internal sources with Nvidia NeMo, promising to shave tuning time and costs.
On inference, it pulls in vLLM, Nvidia TensorRT-LLM, and Nvidia Dynamo, which should keep the GPU meters spinning more efficiently.
It adds built-in observability to track performance and service levels, so IT teams can match model workloads to GPU resources.
GPU orchestration uses pooled infrastructure with on-demand allocation, which sounds tidy until everyone wants the same kit at the same time.
Automatic checkpointing protects long-running jobs and reduces the risk of data loss when something crashes mid-run.
Security is anchored in Red Hat Enterprise Linux, with compliance controls and isolation baked in, while Nvidia DOCA microservices add runtime protections to support zero trust across hybrid deployments.
Red Hat, CTO and SVP global engineering Chris Wright said: “The shift from AI experimentation to industrial-scale, enterprise-wide production requires a fundamental change in how we manage the AI computing stack. We’re accelerating the path to deploying AI and moving quickly to production with Red Hat AI Factory and Nvidia. With a stable, high-performance foundation driven by our proven hybrid cloud offerings, we’re enabling our customers to own their AI strategy and scale with the same rigour they apply to their core IT platforms.”
Nvidia, VP enterprise AI platforms Justin Boitano said: “Enterprises are building AI factories that turn data into intelligence at scale during inference, requiring production-grade infrastructure and software that span the hybrid cloud. Red Hat AI Factory with Nvidia provides the software foundation that helps organisations keep pace with rapid infrastructure innovation while reliably building and deploying the next generation of agentic AI applications.”
Hardware partners, including Cisco, Dell Technologies, Lenovo and Supermicro, have confirmed support across their AI-focused systems, with TD SYNNEX and WWT lining up as distribution and integration partners.







