Motif-3 NVFP4
Motif-3 NVFP4
Serve Motif-3 NVFP4 with Dynamo and the Motif vLLM runtime on NVIDIA B200 GPUs.
This experimental recipe serves Motif-Technologies/Motif-3-NVFP4 on two NVIDIA B200 GPUs with vLLM tensor parallelism, expert parallelism, Dynamo MTP2 speculative decoding, and a 262K-token context. Only an aggregated target is provided; there is no disaggregated target for Motif-3 yet.
Deployment target
Prerequisites
- A Kubernetes cluster with the Dynamo platform installed and 2xB200 GPUs available.
- Create a namespace and an
hf-token-secretcontaining access to the model. The token is used only by the model-download Job.
Deploy
Edit storageClassName in the model-cache manifest, then create the cache and download the checkpoint:
Apply the aggregated chat manifest:
The base manifest has no cluster-specific scheduling. To add node selectors or tolerations for your cluster, use the Kustomization described in the recipe README.
Smoke Test
First, forward the frontend port for your target:
Send a request:
Benchmark
The AIPerf manifest replays the 8k_1k_70kv_chat_new_noschedule_short_15perc.jsonl trace at concurrency 11 against the Dynamo MTP2, TP2 deployment. Stage that trace at /shared-model-cache/traces/8k_1k_70kv_chat_new_noschedule_short_15perc.jsonl, then apply the Job:
The deployment uses actual MTP verification by default. For benchmark-only synthetic acceptance-length trials, change the SPECULATIVE_CONFIG ConfigMap key in the deployment manifest from speculative-config to speculative-config-synthetic. These acceptance lengths were calculated by running SPEED-Bench on the coding domain:
Expected Performance
This is a benchmark-only synthetic proxy, not the expected performance of the shipped manifests. It was measured with speculative-config-synthetic (MTP2, acceptance length 2.13 from SPEED-Bench coding) on the chat trace. The deployment manifest uses real MTP verification, and the AIPerf Job does not select the synthetic configuration, so applying them as shipped will not reproduce these figures. A real-MTP chat-trace result is not yet available.
Under the synthetic proxy, the run meets user output throughput P50 ≥ 50 tokens/second/user and TTFT P50 < 5 seconds. Thirty-five trace requests exceeded the configured 262,144-token context limit.
Notes
- The recipe uses the public
nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.5.0-motif-3-dev.1image; no image-pull secret is required. - The model image carries Motif-specific vLLM compatibility code. Do not replace its vLLM package with an unrelated nightly wheel.