Compatibility

Hardware, platform, and feature support for Dynamo backends
View as Markdown

Compatibility by version

Dynamo v1.3.0GA release

Released Jul 20, 2026 · Release notes · UCX 1.20.x

SGLang0.5.14
NIXL1.3.0
CUDA 13.0Driver 580.xx+
TensorRT-LLM1.3.0rc19
NIXL1.0.1
CUDA 13.1Driver 580.xx+
vLLM0.23.0
NIXL1.1.0
CUDA 13.0Driver 580.xx+
GPU
BlackwellHopperAda LovelaceAmpere
OS
Ubuntu 24.04Ubuntu 22.04CentOS Stream 9 · experimental
Arch
x86_64ARM64 (Ubuntu 24.04 only)

CUDA 12 container images discontinued; EFA variants go multi-arch as -efa; GA wheels published as 1.3.0.post1 (containers stay :1.3.0); UCX 1.20.x.

Backend versions listed are the versions tested and supported for the selected release. TensorRT-LLM does not support Python 3.11.

For extended driver compatibility beyond the listed minimums, including forward compatibility and cuda-compat packages, see the CUDA Compatibility documentation.

See Release Artifacts for the full artifact inventory — container images, wheels, Helm charts, and crates — and Model Early Access Builds for per-model early access container builds.

Find a Compatible Release

Pick your backend and the CUDA generation your host driver supports to see which Dynamo releases you can run — and what to pull for the current one.

What runs where

Releases that match your backend and driver

Backend
CUDA driver situation
v1.3.0CUDA 13.0driver 580.xx+Current
v1.2.1CUDA 13.0driver 580.xx+
v1.2.0CUDA 13.0driver 580.xx+
v1.1.1CUDA 13.0driver 580.xx+
v1.1.0CUDA 13.0driver 580.xx+
v1.0.2CUDA 13.0driver 580.xx+
v1.0.1CUDA 13.0driver 580.xx+
v1.0.0CUDA 13.0driver 580.xx+
v0.8.1CUDA 13.0Experimental imagedriver 580.xx+
v0.8.0CUDA 13.0Experimental imagedriver 580.xx+
v1.2.1CUDA 12.9driver 575.xx+
v1.2.0CUDA 12.9driver 575.xx+
v1.1.1CUDA 12.9driver 575.xx+
v1.1.0CUDA 12.9driver 575.xx+
v1.0.2CUDA 12.9driver 575.xx+
v1.0.1CUDA 12.9driver 575.xx+
v1.0.0CUDA 12.9driver 575.xx+
v0.9.1CUDA 12.9driver 575.xx+
v0.9.0CUDA 12.9driver 575.xx+
v0.8.1CUDA 12.9driver 575.xx+
v0.8.0CUDA 12.9driver 575.xx+
v0.7.1CUDA 12.8driver 570.xx+
v0.7.0CUDA 12.9driver 575.xx+
v1.3.0CUDA 13.1driver 580.xx+Current
v1.2.1CUDA 13.1driver 580.xx+
v1.2.0CUDA 13.1driver 580.xx+
v1.1.1CUDA 13.1driver 580.xx+
v1.1.0CUDA 13.1driver 580.xx+
v1.0.2CUDA 13.1driver 580.xx+
v1.0.1CUDA 13.1driver 580.xx+
v1.0.0CUDA 13.1driver 580.xx+
v0.9.1CUDA 13.0driver 580.xx+
v0.9.0CUDA 13.0driver 580.xx+
v0.8.1CUDA 13.0driver 580.xx+
v0.8.0CUDA 13.0driver 580.xx+
v0.7.1CUDA 13.0driver 580.xx+
v0.7.0CUDA 13.0driver 580.xx+
No release ships TensorRT-LLM for a CUDA 12 driver. TensorRT-LLM ships CUDA 13 images only — switch the driver filter.
v1.3.0CUDA 13.0driver 580.xx+Current
v1.2.1CUDA 13.0driver 580.xx+
v1.2.0CUDA 13.0driver 580.xx+
v1.1.1CUDA 13.0driver 580.xx+
v1.1.0CUDA 13.0driver 580.xx+
v1.0.2CUDA 13.0driver 580.xx+
v1.0.1CUDA 13.0driver 580.xx+
v1.0.0CUDA 13.0driver 580.xx+
v0.8.1CUDA 13.0Experimental imagedriver 580.xx+
v0.8.0CUDA 13.0Experimental imagedriver 580.xx+
v1.2.1CUDA 12.9driver 575.xx+
v1.2.0CUDA 12.9driver 575.xx+
v1.1.1CUDA 12.9driver 575.xx+
v1.1.0CUDA 12.9driver 575.xx+
v1.0.2CUDA 12.9driver 575.xx+
v1.0.1CUDA 12.9driver 575.xx+
v1.0.0CUDA 12.9driver 575.xx+
v0.9.1CUDA 12.9driver 575.xx+
v0.9.0CUDA 12.9driver 575.xx+
v0.8.1CUDA 12.9driver 575.xx+
v0.8.0CUDA 12.9driver 575.xx+
v0.7.1CUDA 12.9driver 575.xx+
v0.7.0CUDA 12.8driver 570.xx+

Driver floors from the CUDA & driver history; pull commands shown for the current release only.

Platform Notes

Dynamo ships multi-arch (x86_64 + ARM64) container images. Wheels are built in a manylinux_2_28-compatible environment and validated on CentOS Stream 9 and Ubuntu 22.04/24.04; other Linux distributions are expected to work but are not officially verified.

Cloud Service Providers

Amazon Linux 2023 (AWS) · x86_64 · Supported

AL2023 TensorRT-LLM limitation: there is a known issue with the TensorRT-LLM framework when running the AL2023 container locally with docker run --network host ... due to a bug in mpi4py. Replace the --network host flag with precise networking configuration by mapping only the necessary ports (4222 for NATS, 2379/2380 for etcd, 8000 for the frontend).

Feature Support

Feature support by backend
SupportedCaveatExperimentalNot supported
SGLang9 / 15
TRT-LLM9 / 15
vLLM14 / 15
Disaggregated Serving
KV-Aware Routing
SLA-Based Planner
KV Block Manager
Multimodal (Image)
Multimodal (Video)
Multimodal (Audio)
Request Migration
Request Cancellation
LoRA
Tool Calling
Speculative Decoding
GPU Memory Service
Shadow Engine Failover
Dynamo Snapshot
  1. Disaggregated Serving · vLLM: Prefill/decode separation with NIXL KV transfer
  2. KV Block Manager · SGLang: Work in progress across all combinations
  3. Multimodal (Image) · SGLang: Not compatible with KV-aware routing. Disagg patterns: EPD, E/PD, E/P/D (not traditional EP/D)
  4. Multimodal (Image) · TRT-LLM: Image URLs + pre-computed embeddings. Disagg: EP/D + E/P/D. KV-aware routing via dedicated MM Router Worker (requires KV event publishing)
  5. Multimodal (Image) · vLLM: With KV-aware routing, image-aware routing on documented paths
  6. Multimodal (Video) · vLLM: Video input with frame sampling
  7. Multimodal (Audio) · vLLM: Qwen2-Audio, experimental
  8. Request Migration · TRT-LLM: Work in progress with multimodal
  9. Request Cancellation · SGLang: Remote-prefill-phase cancellation not supported in disaggregated mode
  10. Request Cancellation · TRT-LLM: Engine temporarily not notified of cancellations — resources for cancelled requests are not freed (known issue)
  11. LoRA · vLLM: Dynamic load/unload; KV-aware routing supports adapter affinity
  12. Speculative Decoding · SGLang: Code hooks exist; no examples or docs yet
  13. Speculative Decoding · vLLM: Eagle3
  14. GPU Memory Service · SGLang: Weights and KV; upstream integration remains in progress
  15. GPU Memory Service · TRT-LLM: Weights only; multinode and upstream integration remain in progress
  16. GPU Memory Service · vLLM: Weights and KV; upstream integration remains in progress
  17. Shadow Engine Failover · SGLang: No KV-cache reuse or hardware fault tolerance
  18. Shadow Engine Failover · TRT-LLM: No KV-cache reuse or hardware fault tolerance
  19. Shadow Engine Failover · vLLM: Software-process failover only; no KV-cache reuse or hardware fault tolerance
  20. Dynamo Snapshot · SGLang: Single-GPU supported; multi-GPU and multinode remain in progress
  21. Dynamo Snapshot · TRT-LLM: Single-GPU aggregated text-worker path only
  22. Dynamo Snapshot · vLLM: Single-GPU supported; multi-GPU is highly experimental and multinode remains in progress

Superscripts reference the numbered notes above; full per-backend detail follows.

Per-Backend Detail

vLLM offers the broadest feature coverage in Dynamo, with full support for disaggregated serving, KV-aware routing, KV block management, LoRA adapters, and multimodal inference including video and audio.

Source: docs/backends/vllm/README.md

FeatureSupported?Notes
Disaggregated ServingPrefill/decode separation with NIXL KV transfer
KV-Aware Routing
SLA-Based Planner
KV Block Manager
MultimodalImage + video; audio experimental (Qwen2-Audio). With KV-aware routing, image-aware routing on documented paths (Source)
Request Migration
Request Cancellation
LoRADynamic load/unload; KV-aware routing supports adapter affinity
Tool Calling
Speculative DecodingEagle3 (Source)
GPU Memory ServiceWeights and KV; upstream integration remains in progress
Shadow Engine Failover!Software-process failover only; no KV-cache reuse or hardware fault tolerance
Dynamo Snapshot!Single-GPU supported; multi-GPU is highly experimental and multinode remains in progress

Feature Interactions

Pairwise feature-by-feature compatibility within each backend. Each cell reports whether the row feature works together with the column feature. A marks the diagonal or a combination that does not apply; blank cells are the mirror of the populated lower triangle.

Legend: Supported · WIP Work in Progress / Experimental / Limited

vLLM Feature Interactions

FeatureDisaggregated ServingKV-Aware RoutingSLA-Based PlannerKV Block ManagerMultimodalRequest MigrationRequest CancellationLoRATool CallingSpeculative Decoding
Disaggregated Serving
KV-Aware Routing
SLA-Based Planner
KV Block Manager
Multimodal1
Request Migration
Request Cancellation
LoRA2
Tool Calling
Speculative Decoding

Notes:

  1. Multimodal + KV-Aware Routing: Image-aware KV routing is supported in the documented vLLM paths. The default Rust frontend path supports model families handled by llm-multimodal; the Python chat-processor path delegates to vLLM’s multimodal processor. (Source)
  2. KV-Aware LoRA Routing: vLLM supports routing requests based on LoRA adapter affinity.
  3. Audio Support: vLLM supports audio models like Qwen2-Audio (experimental). (Source)
  4. Video Support: vLLM supports video input with frame sampling. (Source)
  5. Speculative Decoding: Eagle3 support documented. (Source)

SGLang Feature Interactions

FeatureDisaggregated ServingKV-Aware RoutingSLA-Based PlannerKV Block ManagerMultimodalRequest MigrationRequest CancellationLoRATool CallingSpeculative Decoding
Disaggregated Serving
KV-Aware Routing
SLA-Based Planner
KV Block Manager🚧🚧🚧
Multimodal21🚧
Request Migration🚧
Request Cancellation🚧3🚧🚧
LoRA🚧
Tool Calling🚧
Speculative Decoding🚧🚧🚧🚧🚧

Notes:

  1. Multimodal + KV-Aware Routing: Not supported. (Source)
  2. Multimodal Patterns: Supports simple Aggregated EPD, E/PD, and E/P/D patterns. Traditional Disagg EP/D is not supported. (Source)
  3. Request Cancellation: Cancellation during the remote prefill phase is not supported in disaggregated mode. (Source)
  4. Speculative Decoding: Code hooks exist (spec_decode_stats in publisher), but no examples or documentation yet.

TensorRT-LLM Feature Interactions

FeatureDisaggregated ServingKV-Aware RoutingSLA-Based PlannerKV Block ManagerMultimodalRequest MigrationRequest CancellationLoRATool CallingSpeculative Decoding
Disaggregated Serving
KV-Aware Routing
SLA-Based Planner
KV Block Manager
Multimodal12
Request Migration🚧
Request Cancellation333333
LoRA
Tool Calling
Speculative Decoding

Notes:

  1. Multimodal Disaggregation: Supports EP/D (Traditional) and E/P/D (Full Disaggregation) image flows, including image URLs and pre-computed embeddings. (Source)
  2. Multimodal + KV-Aware Routing: The native Rust frontend routes supported models using image-aware KV overlap. TRT-LLM workers must publish KV events with block reuse enabled. (Source)
  3. Request Cancellation: Due to known issues, the TensorRT-LLM engine is temporarily not notified of request cancellations, meaning allocated resources for cancelled requests are not freed.