Back Home

AI 基礎設施/資料中心

AWS Receives First Vera CPU/Vera Rubin Systems, Plans to Add 2 Million NVIDIA GPUs Over Two Years

AWS has received its first Vera CPU servers and Vera Rubin GPUs, while the two companies plan to deploy an additional 2 million NVIDIA GPUs from 2027 to 2028. The collaboration also incorporates NVLink Fusion, NVHBM, Nitro, and EFA into heterogeneous rack designs, though most capacity and performance figures remain future commitments.

Wikideas1 · CC0 · Image source
zh-Hant

NVIDIA confirmed on August 27 that AWS had received its first Vera CPU servers and Vera Rubin GPUs. This represents the delivery of physical systems, but it does not mean customers can already launch corresponding EC2 instances. The more concrete commercial plan is for AWS to add 2 million Blackwell Ultra, Rubin, and Rubin Ultra GPUs across its global infrastructure and AI factories between 2027 and 2028. This capacity is in addition to a previously announced expansion plan involving more than 1 million GPUs.

The technical significance goes beyond GPU count. AWS plans to introduce standalone Vera CPUs as well as Vera CPUs paired with Rubin, allowing tool execution, sandboxes, data pipelines, and orchestration for agentic workloads to run without competing for accelerator resources. Meanwhile, Amazon subsidiary Annapurna Labs will combine NVLink Fusion with NVIDIA NVHBM, enabling next-generation Trainium chips and NVIDIA GPUs to operate within a shared rack-scale scale-up architecture. Scale-out connectivity across racks will continue to use the AWS Nitro System, Elastic Fabric Adapter, and Spectrum networking being jointly optimized by the two companies. This suggests that AWS-designed silicon and NVIDIA platforms may share more memory, interconnect, and management layers instead of remaining in entirely separate clusters.

The companies also announced performance claims for adjacent data pipelines: EC2 G7 instances powered by RTX PRO 4500 deliver 4.6 times the inference performance of G6; EMR with cuDF provides up to a 3.7-fold speedup; and GPU-accelerated vector indexing in OpenSearch is up to nine times faster while reducing costs to one-quarter. These are official figures based on specific configurations, and complete details covering datasets, power consumption, network topology, and cost-equivalent comparisons have not yet been provided. Engineering teams should next watch for the general availability dates, regions, and quotas for Vera/Rubin instances, the proportion of inter-GPU connectivity, reservable capacity, and actual pricing. Until those details are published, the 2 million GPUs should be viewed as a multiyear deployment plan rather than compute capacity that can be scheduled immediately.

Sources

  1. Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now
  2. AWS and NVIDIA to Deliver 2 Million Additional GPUs and Next-Generation Infrastructure for Agentic and Physical AI