Skip to content

Repository files navigation

AWS Logo

AWS Deep Learning Containers

One stop shop for running AI/ML on AWS

Docs · Available Images · Tutorials

Auto Release - PyTorch 2.14 Auto Release - TensorFlow Training 2.21 Auto Release - TensorFlow Inference 2.20 Auto Release - vLLM Auto Release - vLLM-Omni Auto Release - SGLang Auto Release - Ray Auto Release - Base cu130 Auto Release - Base cu132


About

AWS Deep Learning Containers (DLCs) are pre-built Docker images for running AI/ML workloads on AWS. Each image is tested and patched for security vulnerabilities. For more details, visit our documentation.


🔥 What's New

🚀 Release Highlights

  • [2026/10/01] vLLM Server ARM64 CPU v1.0 (AL2023) · EC2: server-cpu-v1.0 · SageMaker: server-sagemaker-cpu-v1.0 · Initial release: vLLM 0.30.0 CPU backend on AWS Graviton (ARM64), no GPU required; bf16 kernels for Graviton 3 and later, RAM-based KV-cache defaults, and the OpenAI-compatible API on EC2 and SageMaker AI.
  • [2026/09/28] PyTorch v2.14.0 — EC2: 2.14-cu133-amzn2023 · SageMaker: 2.14-cu133-amzn2023-sagemaker · PyTorch 2.14.0 with torchvision 0.29.0; CUDA 13.3.1, EFA 1.50.0, TE 2.18.0, DeepSpeed 0.19.6.
  • [2026/09/26] vLLM Server v2.6 (AL2023) — EC2: server-cuda-v2.6 · SageMaker: server-sagemaker-cuda-v2.6 · DeepEP v2 over EFA, built from the amazon-contributing/DeepEP fork; NCCL 2.31.2; EFA 1.50.0. vLLM stays at 0.30.0 (ec4a3a5).
  • [2026/09/26] vLLM-Omni v1.8 (AL2023) — EC2: omni-cuda-v1.8 · SageMaker: omni-sagemaker-cuda-v1.8 · vLLM-Omni 0.29.0rc1 (up from 0.28.0); DeepEP v2 over EFA; FlashInfer 0.6.18, NCCL 2.31.2, EFA 1.50.0.
  • [2026/09/24] SGLang Server v1.4 (AL2023) — EC2: server-cuda-v1.4 · SageMaker: server-sagemaker-cuda-v1.4 · SGLang 0.5.19 (up from 0.5.17); DeepEP v2 over EFA; PyTorch 2.13.0, NCCL 2.31.2, EFA 1.50.0.
  • [2026/09/23] vLLM v0.30.0 (Ubuntu) — EC2: 0.30.0-gpu-py312-ec2 · SageMaker: 0.30.0-gpu-py312 · DeepSeek-V4.1-Flash, GLM-5.3-Flash, K2-Horizon; Fast Start GPU weight cache; Gumbel-max watermarking; transformers capped below 5.17.
  • [2026/09/23] vLLM Server v2.5 (AL2023) — EC2: server-cuda-v2.5 · SageMaker: server-sagemaker-cuda-v2.5 · vLLM 0.30.0 (up from 0.27.1), built from ec4a3a5 for Transformers 5.17; FlashInfer 0.6.18.post1.
  • [2026/09/21] vLLM-Omni v1.7 (AL2023) — EC2: omni-cuda-v1.7 · SageMaker: omni-sagemaker-cuda-v1.7 · vLLM-Omni 0.28.0 (up from 0.26.0); PEFT LoRA adapters validated on SageMaker — per-request lora selection on /v1/images/generations; FlashInfer 0.6.16.post3, NCCL pinned 2.30.7, CUDA-13 mooncake wheel.
  • [2026/09/19] SGLang v0.5.20 (Ubuntu) — EC2: 0.5.20-gpu-py312-ec2 · SageMaker: 0.5.20-gpu-py312 · GLM-5.3-Flash, Hy4-Preview, Qwen3.8-Flash-Next, K2 Horizon; RL sampling masks; unified radix tree; DeepSeek-V4 on Blackwell.
  • [2026/09/11] vLLM v0.29.0 (Ubuntu) — EC2: 0.29.0-gpu-py312-ec2 · SageMaker: 0.29.0-gpu-py312 · New models: Hy4-preview (Tencent 770B/49B-active MoE with Gated DeepSeek Sparse Attention and native MTP), Qwen3.8-Flash-Next (BF16/FP8/NVFP4, MTP), GraniteSWA, GraniteMoeSWA, NemotronH_Omni_Reasoning_V3 (MTP), and Kimi K3 NVFP4 checkpoints.
  • [2026/09/07] SGLang v0.5.19 (Ubuntu) — EC2: 0.5.19-gpu-py312-ec2 · SageMaker: 0.5.19-gpu-py312 · Qwen3.8, Ling-3.0, Spark2.5, Granite 4.2; beam search; DeepEP v2 MoE all-to-all.
  • [2026/09/01] Ray LLM v1.0 (2.58.0, AL2023) — EC2/EKS: serve-llm-cuda-v1.0 · Initial release: OpenAI-compatible LLM serving with Ray Serve and vLLM 0.26.0 on PyTorch 2.11.0 / CUDA 13.0.2 / Python 3.13; ray[llm]'s build_openai_app runs vLLM behind Ray Serve — single GPU on EC2, and multi-node serving on EKS via KubeRay.
  • [2026/09/01] Ray Train v1.1 (2.58.0, AL2023) — EC2/EKS: train-ml-cuda-v1.1 · EFA 1.49.0 (up from 1.47.0).

📢 Support Updates

  • [2026/04/28] We cannot guarantee security patching on Ubuntu-based vLLM and SGLang images due to the lack of Ubuntu Pro licensing. Customers may continue using these images at their own discretion and risk. We recommend migrating to our Amazon Linux-based images.
  • [2026/02/10] Extended support for PyTorch 2.6 Inference containers until June 30, 2026
    • PyTorch 2.6 Inference images will continue to receive security patches and updates through end of June 2026
    • For complete framework support timelines, see our Support Policy

📝 Blog Posts

🎓 Workshop


License

This project is licensed under the Apache-2.0 License.

About

One stop shop for running AI/ML on AWS.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1.2k stars

Watchers

49 watching

Forks

Used by

Contributors

Languages