AWS and NVIDIA plan 2 million more GPUs for agentic and physical AI

The two companies are expanding their strategic collaboration to deploy 2 million additional NVIDIA GPUs across AWS infrastructure in 2027 and 2028.

AWS and NVIDIA plan 2 million more GPUs for agentic and physical AI

Amazon Web Services and NVIDIA have announced a significant expansion of their long-standing partnership, with plans to deploy two million additional NVIDIA GPUs across AWS's global infrastructure during 2027 and 2028. The announcement, made on 26 August 2026, builds on an earlier commitment at NVIDIA GTC 2026 to add more than one million GPUs from 2026 onwards, a target the companies say has already been overtaken by customer demand.

The expanded collaboration spans compute, networking, memory interconnects, open models, data processing and robotics, reflecting the breadth of the infrastructure requirements that large-scale AI workloads now place on cloud providers.

What the deal covers

The two million additional chips will include NVIDIA Blackwell Ultra, Rubin and Rubin Ultra GPUs. AWS is also the first major cloud provider to offer compute instances powered by the RTX PRO 4500 Blackwell Server Edition, which the companies say delivers 4.6 times the AI inference performance and 2.1 times the graphics throughput of the previous-generation G6 instance.

Beyond raw GPU counts, the collaboration introduces several deeper technical integrations. NVIDIA's Vera CPU is being brought to AWS to support agentic workloads that require high-performance CPU compute alongside accelerated infrastructure. NVIDIA NVLink Fusion is being extended to work with NVIDIA's custom high-bandwidth memory (NVHBM), allowing Amazon's Annapurna Labs to integrate Trainium and GPU silicon within a shared rack-scale architecture.

On the data side, GPU-accelerated processing on Amazon EMR using the NVIDIA cuDF library is reported to deliver up to 3.7 times faster processing and 30 per cent better price performance compared with CPU-based configurations. GPU-accelerated vector indexing on Amazon OpenSearch is said to be up to nine times faster at one quarter of the cost, a figure that will attract attention from enterprises running retrieval-augmented generation pipelines at scale.

The collaboration also targets US federal customers. AWS and NVIDIA plan to build AI factories for the US government, including a cluster of 100,000 GPUs running on secure AWS infrastructure cleared for Impact Level 6 and above workloads covering national-security use cases.

Jensen Huang, founder and chief executive of NVIDIA, said the partnership was "expanding across the full stack, GPUs, CPUs, networking, open models and software, to make agentic and physical AI real at an unprecedented pace and scale."

Market context and competitive read-across

The announcement underlines the degree to which the hyperscaler-GPU vendor relationship has shifted from a straightforward procurement arrangement to a tightly co-engineered stack. Microsoft Azure's partnership with NVIDIA and its own investment in custom Maia accelerators, alongside Google Cloud's TPU programme and its own NVIDIA fleet, show that every major cloud provider is pursuing a hybrid strategy: proprietary silicon for cost efficiency, NVIDIA for frontier model performance.

AWS's position is slightly distinctive: it is simultaneously deepening its NVIDIA dependency while continuing to develop its own Trainium line, with the NVLink Fusion integration an attempt to make the two architectures interoperable rather than competing. That heterogeneous approach lowers customer lock-in risk and gives AWS flexibility as the mix of training versus inference workloads evolves.

The federal AI factory commitment also signals a commercially significant shift. Government AI spending is growing rapidly, and IL6-cleared GPU infrastructure is a scarce resource. Winning that footprint early creates a durable competitive moat that is difficult for rivals to replicate quickly given the security-accreditation timelines involved.

For enterprise buyers, the near-term milestones to watch are the availability timeline for Vera CPU instances on AWS, the general rollout of NVLink Fusion with NVHBM, and whether the benchmark figures for EMR and OpenSearch acceleration hold up under independent testing.