This document is relevant for: Trn2, Trn3

Wan2.2-I2V-A14B Model Card#

Introduction#

Wan2.2-I2V-A14B is an image-to-video diffusion model developed by Wan-AI. Given a conditioning image (the first frame) and a text prompt, it animates the image into a short video that stays faithful to the input frame. Like its text-to-video sibling, it uses a Mixture-of-Experts (MoE) architecture with two experts — a high-noise expert for overall layout during early denoising stages and a low-noise expert for detail refinement during later stages. The model has ~27B total parameters but only 14B active parameters per inference step, keeping computation and memory roughly equivalent to a single 14B dense model. It generates 5-second videos (81 frames at 16fps) at 480P and 720P resolutions with cinematic-level aesthetics and complex motion.

Wan2.2-I2V-A14B is now supported for inference serving with vLLM Omni using the Neuron SDK on AWS Trainium2 (trn2) and Trainium3 (trn3) hardware. For how the image conditioning path works on Neuron, see Image conditioning on Neuron.

Compatible model checkpoints:

Model

HuggingFace

Hardware

Quantization

Wan2.2-I2V-A14B

Wan-AI/Wan2.2-I2V-A14B-Diffusers

Trn2, Trn3

BF16

I2V runs in BF16. FP8 is not currently supported for image-to-video.

Features#

Per-model feature availability for Wan2.2-I2V-A14B. See the README for configuration details.

Category

Feature

Status

Generation

Image-to-Video

✅

832x480 resolution (480p)

✅

1280x720 resolution (720p)

✅

Up to 81 frames

✅

Quantization

BF16

✅

Parallelism

Tensor Parallelism (TP)

✅

Context Parallelism (CP)

✅

Megatron Sequence Parallelism (SP)

✅

CFG Parallelism

✅

VAE Patch Parallelism

✅

Performance

Classifier-Free Guidance

✅

Spatial Tiling (VAE)

✅

Temporal Chunking (VAE)

✅

Continuous request batching

Limited

Compilation

torch.compile

✅

Status legend:

  • ✅ Supported: integrated and tested for Wan2.2-I2V-A14B

  • Limited: accepted, but concurrent requests run serially rather than as a batched forward pass

Image conditioning on Neuron#

The conditioning image is passed through multi_modal_data ({"image": <PIL.Image>}); an optional last_image enables FLF2V. On Neuron the DiT runs SPMD on every rank, but the VAE is materialized only on the VAE rank(s), so the pipeline VAE-encodes the image on the VAE rank and broadcasts the fixed-shape condition latent to all ranks before the denoise loop. The seeded noise latents are RNG-reproducible and identical on every rank, so they need no broadcast.

Accuracy Evaluation#

Benchmark: VBench-I2V is the image-to-video track of VBench, a comprehensive benchmark suite for video generation models. It scores generation quality (subject/background consistency, motion smoothness, dynamic degree, aesthetic quality, imaging quality) alongside how faithfully the video preserves the conditioning frame and camera motion (the I2V dimensions).

See the VBench paper for the original benchmark and the VBench++ paper for its image-to-video extension.

Wan2.2-I2V-A14B-Diffusers (BF16), 480p

Platform

Total Score

Quality Score

I2V Score

Trn2

88.14%

79.49%

96.78%

  • Quality Score dimensions: subject_consistency, background_consistency, motion_smoothness, dynamic_degree, aesthetic_quality, imaging_quality.

  • I2V Score dimensions: i2v_subject, i2v_background, camera_motion.

  • Total Score is the simple average of the Quality and I2V scores.

  • Scores use VBench’s normalization and weights (see vbench2_beta_i2v).

For externally published results, see the image-to-video results on the VBench Leaderboard. Results from different evaluation configurations are not directly comparable.

Reproduce: Serve the model following the quickstart, then run VBench-I2V evaluation:

git clone https://github.com/Vchitect/VBench.git
cd VBench && pip install . && cd ..
python VBench/evaluate_i2v.py \
    --videos_path wan_i2v_output_dir \
    --custom_image_folder wan_i2v_input_images \
    --dimension i2v_subject i2v_background camera_motion \
        subject_consistency background_consistency motion_smoothness \
        dynamic_degree aesthetic_quality imaging_quality \
    --ratio 16-9 \
    --mode=custom_input

Known limitations#

  • Batch size is limited to one. The current Wan2.2 pipeline generates one video per request; concurrent requests run serially rather than as a batched forward pass.

  • CFG parallelism (cfg_parallel_size=2) is not validated at 720P. Use the recommended TP8CP4CFGP1 config for the I2V 720P case.

Tutorials#

This document is relevant for: Trn2, Trn3