This document is relevant for: Trn2, Trn3

vLLM Omni Neuron (Beta)#

vLLM Omni Neuron is a Neuron hardware plugin for vLLM Omni that runs diffuser-type multimodal generation models — such as video and image generation — on AWS Trainium. Diffuser models are not supported by upstream vLLM, so they are not served by vLLM Neuron, the Neuron LLM plugin. They live in vLLM Omni, an extended library under the vLLM project for multimodal generation. vLLM Omni Neuron adds the Neuron backend that library needs.

The plugin provides Neuron-optimized reference implementations — including NKI kernels, custom compilation, and hardware-aware scheduling — for models whose upstream implementations live in vLLM Omni. It uses native PyTorch (torch.compile) through the libtorch_neuronx_lite package, the same native PyTorch infrastructure used by vLLM Neuron.

Note

vLLM Omni Neuron is in Beta and under active development. It is supported on Trainium 2 (trn2) and Trainium 3 (trn3) instances only.

The source code for the vLLM Omni Neuron plugin is hosted in the vLLM Omni Neuron GitHub repository.


Get started#

Get started

Choose manual installation or a Neuron DLC, then run the online and offline serving quickstarts.

Guides

Features guide and offline video generation optimization on Trainium.

Model recipes

Production-ready deployment recipes for supported models on Trainium.

Tutorials

End-to-end walkthrough for deploying Wan2.2-A14B on Trainium.

Model development

Model onboarding, accuracy evaluation and debugging, and Neuron kernel implementation references.

Concepts & architecture

Engine and model integration, and context parallelism for diffusion inference.

This document is relevant for: Trn2, Trn3