Back Home

資料中心基礎設施

Azure to Adopt AMD Helios, Add Compute Instances for Large-Scale Inference, Data Pipelines, and Chip Design

Microsoft plans to deploy AMD Helios at scale on Azure and launch HDv2 and HXv2 instances powered by sixth-generation EPYC processors, along with MI455X inference instances. The infrastructure update co-designs GPUs, CPUs, DPUs, networking, and software, though its real-world price-performance remains to be demonstrated through public testing.

zh-Hant

Microsoft and AMD are expanding their Azure infrastructure partnership, centered on bringing AMD Helios rack-scale systems to the cloud for frontier models, Azure AI services, and customers’ large-scale inference workloads. Helios is not a standalone accelerator product but a complete platform integrating Instinct MI455X GPUs, sixth-generation EPYC CPUs code-named Venice, Pensando networking components, and ROCm software. AMD says shipments to customers, including Microsoft, will begin in the second half of 2026. Azure’s customer-facing inference offering will primarily use ND MI455X v7 instances, targeting inference, search, and agentic workloads.

The same set of updates also addresses bottlenecks beyond accelerators. HDv2 is designed for data preparation, search, reinforcement learning, and agent orchestration, with specifications including nearly 500 physical EPYC cores, 4 TB of memory, 32 TB of local NVMe storage, and 400 Gb Azure Boost networking. HXv2 targets RTL simulation, electronic design automation, and distributed scientific computing, offering 176 EPYC cores, clock speeds above 5 GHz, 50% more addressable cache per core, nearly 2 TB or 4 TB of memory, and 800 Gb InfiniBand. Azure will also expand its deployment of Pensando DPUs and integrate them with Azure Boost, offloading connectivity and some network processing from host compute resources.

The technical significance is that the unit of competition in cloud AI is shifting from the GPU to the entire rack and data path. Inference-service throughput is constrained simultaneously by memory for model weights, intra-node interconnects, cross-node networking, CPU data processing, and scheduling software. If Helios can become a standardized, repeatable deployment configuration on Azure, engineering teams will gain an alternative path to a closed, single-vendor stack.

However, the two companies have not disclosed pricing, regional availability, a general availability date, power consumption, or independently verified inference performance per dollar or per watt for ND MI455X v7. “Deployment at scale” also comes without a publicly stated rack count. Key areas to watch include ROCm’s model and kernel support, the stability of inter-GPU communication, actual supply availability, and tail latency for mainstream models under long-context, sparse MoE, and concurrent agentic workloads.

Sources

  1. Microsoft expands Azure AI and HPC infrastructure with AMD
  2. Microsoft to Deploy Next-Gen AMD Instinct and AMD EPYC Processors