This document is relevant for: Trn2, Trn3
Tutorial: Deploy Wan2.2-A14B with vLLM Omni Neuron#
This tutorial walks through deploying the Wan2.2-A14B diffusion video models with the vLLM Omni Neuron plugin. It covers environment setup, model download, stage configuration, and offline video generation. Wan2.2 comes in two variants, both served by this plugin:
Wan2.2-T2V-A14B — text-to-video. See its model card.
Wan2.2-I2V-A14B — image-to-video. See its model card.
Both are Mixture-of-Experts models (~27B total, 14B active per step) that generate
5-second clips (81 frames at 16fps) at 480P and 720P. Generation runs offline through
the vLLM Omni Neuron entrypoint — there is no online chat endpoint. BF16 is the default precision
throughout; FP8 (fp8_row_mx) is an optional T2V path covered in the optional step.
Variant |
Entry script |
Default resolution |
Cores (recommended) |
|---|---|---|---|
T2V |
|
832x480 |
64 (TP4 × CP8 × CFG2) |
I2V |
|
832x480 |
64 (TP4 × CP8 × CFG2) |
Prerequisites:
One SSH-accessible
trn2.48xlargeortrn3instance. Step 1 prepares it.
Step 1: Set up your environment#
Complete the setup guide, choosing manual or DLC installation. It configures the dependencies, device access, and persistent caches, then verifies the environment.
Run these commands in an SSH session on the instance host, using its DLC container
vllm-omni-neuron. For manual installation,
run the Python commands in the activated environment, omit
sudo docker exec vllm-omni-neuron, and replace /workspace with
$VLLM_OMNI_HOME.
Step 2: Download the model (optional)#
sudo docker exec vllm-omni-neuron hf download \
Wan-AI/Wan2.2-T2V-A14B-Diffusers \
--local-dir /workspace/huggingface/Wan2.2-T2V-A14B-Diffusers
Note
This step is optional. You can pass the Hugging Face model ID (e.g.
Wan-AI/Wan2.2-T2V-A14B-Diffusers) directly via --model-path, and the weights are
downloaded automatically on first run into the mounted cache. For I2V, download
Wan-AI/Wan2.2-I2V-A14B-Diffusers instead.
Step 3: Review the stage configuration#
The parallelism layout and diffusion settings live in a stage-config YAML, not on the
command line. The entry scripts default to
examples/wan22/wan22_stage.yaml (T2V) and
examples/wan22/wan22_i2v_stage.yaml (I2V);
override with --stage-config. The default 64-core T2V config is:
# engine_args in the stage config
dtype: bfloat16
boundary_ratio: 0.875 # I2V uses 0.9 — the MoE high→low-noise expert switch point
flow_shift: 5.0
vae_use_tiling: true # tile the VAE decode to fit 720P within HBM
model_config:
tp_sequence_parallel: true # Megatron sequence parallelism (720P headroom)
parallel_config:
tensor_parallel_size: 4 # TP — shards the model
ring_degree: 8 # CP degree — sets sequence_parallel_size = 1 * 8 = 8
cfg_parallel_size: 2 # runs the two guidance passes in parallel
vae_patch_parallel_size: 16 # default; the 720P config uses 32 ranks
ring_degree(notsequence_parallel_size) sets the context-parallel degree; vLLM Omni derivessequence_parallel_size = ulysses_degree(1) * ring_degree. CP self-attention runs ring attention. See the context parallelism design doc.boundary_ratioshould be0.9for I2V and0.875for T2V — the I2V stage config already sets this. A wrong value switches MoE experts at the wrong step.
Step 4: Run inference#
Text-to-video (T2V)#
Quick smoke test (5 frames, low resolution) to validate the setup:
sudo docker exec vllm-omni-neuron python /workspace/plugin/examples/wan22/run.py --dev
Full 480P generation (81 frames, 832x480, 40 steps):
sudo docker exec vllm-omni-neuron python /workspace/plugin/examples/wan22/run.py \
--num-frames 81 --height 480 --width 832 --guidance-scale 5.0 \
--prompt "A fluffy orange cat walking gracefully across a sunny garden path, high quality, detailed" \
--output 480P.mp4
720P generation (1280x720) uses the 720P stage config, with VAE tiling and 32-way VAE patch parallelism:
sudo docker exec vllm-omni-neuron python /workspace/plugin/examples/wan22/run.py \
--stage-config /workspace/plugin/examples/wan22/wan22_stage_tp4cp8cfg2_720p.yaml \
--num-frames 81 --height 720 --width 1280 --guidance-scale 5.0 \
--prompt "A fluffy orange cat walking gracefully across a sunny garden path, high quality, detailed" \
--output 720P.mp4
Image-to-video (I2V)#
I2V animates a conditioning first frame passed with --image:
sudo docker exec vllm-omni-neuron python /workspace/plugin/examples/wan22/run_i2v.py \
--image /workspace/plugin/examples/wan22/i2v_input.JPG \
--num-frames 81 --height 480 --width 832 --guidance-scale 5.0 \
--prompt "Summer beach vacation style, a white cat wearing sunglasses sits on a surfboard." \
--output i2v_480P.mp4
sudo docker exec vllm-omni-neuron python /workspace/plugin/examples/wan22/run_i2v.py \
--stage-config /workspace/plugin/test/neuron/configs/wan22_i2v_stage_tp8cp4cfg1_720p.yaml \
--image /workspace/plugin/examples/wan22/i2v_input.JPG \
--num-frames 81 --height 720 --width 1280 \
--output i2v_720P.mp4
Note
CFG parallelism (cfg_parallel_size=2) is not yet supported for I2V at 720P.
vLLM Omni Neuron compiles the model on the first run.
With the DLC flow, copy the output with
docker cp vllm-omni-neuron:/workspace/480P.mp4 .; with manual installation, it is in your working directory.
Optional: FP8 for T2V#
T2V can run its QKV, output, and FFN projections in FP8 (fp8_row_mx) from a
ComfyUI-repackaged FP8 checkpoint. Skip this for the default BF16 path.
sudo docker exec vllm-omni-neuron python /workspace/plugin/examples/wan22/run.py \
--quantization fp8_row_mx \
--comfyui-fp8-model-path Comfy-Org/Wan_2.2_ComfyUI_Repackaged \
--num-frames 81 --height 480 --width 832 \
--output 480P_fp8.mp4
--modules-to-not-convert selects which FP8 modules to CPU-dequant; pass
attn1 attn2 ffn for the full CPU-dequant reference.
Troubleshooting#
Symptom |
Likely cause |
Fix |
|---|---|---|
|
Driver/runtime incompatibility or unavailable devices |
Follow the host prerequisites and device checks in the setup guide. |
Silent crash / |
|
Start the container with |
OOM during VAE decode at 720P |
Single-shot decode exceeds HBM |
Set |
Very long first-run latency |
Compiler generating NEFFs on first inference |
Wait for the first compilation to complete. |
Conclusion#
You have deployed Wan2.2-A14B on Trainium — text-to-video and image-to-video, at 480P and 720P, in BF16 (and optionally FP8 for T2V). Generation runs offline through the vLLM Omni Neuron entrypoint.
Next steps#
Tune throughput and memory — choosing TP/CP/CFG degrees with roofline analysis and profiling: Optimizing offline video generation.
This document is relevant for: Trn2, Trn3