Amazon engineers published a technical guide on May 11, 2026 that maps the building blocks for running the full foundation-model lifecycle on AWS. The post, hosted on the Hugging Face blog, argues that pre-training, post-training and inference have converged on the same requirements: tightly coupled accelerators, a high-bandwidth network and scalable shared storage.
From One Scaling Law to Three
Scaling is no longer a single curve. The authors point to pre-training, post-training through supervised fine-tuning and reinforcement learning, and test-time compute as three regimes that share infrastructure needs but differ in how workloads are scheduled and how resources move.
Compute: From H100 to Blackwell Ultra
- p5.48xlarge: eight H100 GPUs, 640 GB HBM3
- p5e.48xlarge: eight H200 GPUs, 1,128 GB HBM3e
- p6-b200.48xlarge: eight B200 GPUs, 1,440 GB HBM3e
- p6-b300.48xlarge: eight B300 GPUs, 2,100 GB HBM3e and 800 GB/s EFA
For communication-heavy workloads such as mixture-of-experts training, the size of the NVLink domain becomes a first-order constraint. P6e-GB200 UltraServers answer this by composing up to 72 Blackwell GPUs and 13.4 TB of HBM3e inside a single NVLink fabric.
Networking and Storage
EFA provides OS-bypass RDMA across nodes using the Scalable Reliable Datagram protocol. EFA v3 on P5en reduces packet latency by roughly 35% against EFA v2, while EFA v4 on P6 adds an 18% improvement in collective performance. Storage is tiered: local NVMe instance store, FSx for Lustre for shared throughput, and Amazon S3 for durable checkpoints.
Orchestration: Slurm and Kubernetes
Slurm schedules at the job level and supports backfill, multi-factor priority and topology-aware placement. Kubernetes needs extra layers for tightly coupled training. Kueue handles gang admission and hierarchical quotas, while Volcano and NVIDIA KAI Scheduler add topology-aware pod placement. SageMaker HyperPod offers both Slurm and EKS modes, including checkpointless and elastic training.
The Software Stack
The runtime is five layers deep: kernel drivers, CUDA and kernel libraries, the NCCL communication substrate, PyTorch, then distributed training and inference frameworks. Training options span Hugging Face Transformers, NVIDIA Megatron Core and veRL for RLHF. Serving leans on vLLM and SGLang, both of which integrate with NVIDIA Dynamo for disaggregated prefill and decode.
Observability and GPU Health
Prometheus and Grafana remain the default telemetry stack. DCGM-Exporter surfaces GPU utilization, memory, power and health metrics, and the guide flags XID 63, 64, 94 and 95 events as signals that warrant immediate node replacement.
What This Means
The post is less a product announcement than a map. It shows teams where bottlenecks hide, from a saturated EFA link to a misconfigured driver, and it confirms that the open-source software stack now covers the entire foundation-model lifecycle.