Cloudera and VAST Data partner on hybrid AI factory platform

The two vendors say their joint platform addresses GPU starvation by piping enterprise data continuously into AI training and inference workloads.

A brightly lit, long aisle in a modern data center is flanked by rows of dark server racks featuring blue liquid cooling tubes and overhead piping.

Cloudera and VAST Data have announced a strategic partnership to deliver what they describe as a unified AI factory: a production environment in which enterprise data is continuously ingested, refined, governed and fed to AI models for training and inference. The joint solution is available immediately through both companies' sales teams and partner ecosystems, with reference architectures and industry-specific deployment patterns to follow throughout 2026.

At its core, the collaboration integrates Cloudera's containerised lakehouse services with VAST's AI Operating System, which bundles high-performance storage, a vector database and a global namespace into a single stack. The combined architecture is built around the NVIDIA AI Data Platform reference design, with NVIDIA NIM microservices accelerating Cloudera's AI Inference Service and NVIDIA cuDF providing GPU-accelerated processing for Apache Spark workloads running on top of VAST storage.

The GPU starvation problem

The central technical claim is that traditional data architectures were not designed for continuous AI pipelines, causing expensive GPU clusters to sit idle while waiting for data. Abhas Ricky, Chief Business Officer and GM of Applied AI at Cloudera, put the point directly: "Enterprises are investing billions in GPUs, yet many struggle to achieve full utilisation due to data bottlenecks. Our partnership with VAST eliminates GPU starvation and enables customers to build true AI factories where data flows seamlessly from ingestion to insight."

Jeff Denworth, co-founder of VAST Data, framed the opportunity as one of unlocking latent value: "Most enterprises already have the data they need for AI. The challenge is unlocking the value in data to create a continuous pipeline of AI inference, fine-tuning, and data analysis to build the next generation of intelligent applications."

The two companies say their combined customer base manages roughly 60 exabytes of data, which they describe as the addressable pipeline for the joint offering. Neither company disclosed committed contract value, customer names, or specific throughput benchmarks for the combined stack.

Market context and competitive positioning

The AI infrastructure market is undergoing rapid consolidation around the concept of the "AI factory," a term popularised by NVIDIA to describe GPU-dense facilities optimised for continuous model training and inference. Cloudera and VAST are competing in an increasingly crowded field that includes hyperscaler-native stacks (AWS, Azure and Google Cloud each offer managed data and inference pipelines), as well as specialist vendors such as Databricks, which combines a lakehouse platform with MLflow and model-serving capabilities, and NetApp, which has positioned its all-flash arrays around similar GPU-feeding narratives.

VAST's differentiation rests on its Disaggregated Shared Everything architecture, which the company positions as eliminating the traditional trade-off between performance, scale and resilience. Cloudera's value proposition is hybrid and multi-cloud portability: the ability to run the same containerised data services on-premises, in private cloud and across public cloud providers. For large enterprises in regulated industries such as financial services, healthcare and government, this combination addresses a genuine pain point: the need to keep sensitive training data on-premises while still benefiting from GPU-accelerated AI workloads.

The partnership's emphasis on sovereign and private AI is notable. With the EU AI Act's general-purpose AI obligations beginning to phase in and UK and US government customers demanding data-residency controls, a "silicon-to-application" stack that never routes data through a public cloud is a commercially meaningful differentiator. Compliance with frameworks such as FedRAMP and ISO 27001 will be a prerequisite for large regulated-industry deals, and neither company has yet detailed its certification posture for the joint solution.

Investors will be watching for named enterprise customers and independently validated benchmark data as the clearest signals that the partnership moves beyond reference architecture into production revenue.