This document is relevant for: Trn2, Trn3
vLLM Omni Neuron (Beta)#
vLLM Omni Neuron is a Neuron hardware plugin for vLLM Omni that runs diffuser-type multimodal generation models — such as video and image generation — on AWS Trainium. Diffuser models are not supported by upstream vLLM, so they are not served by vLLM Neuron, the Neuron LLM plugin. They live in vLLM Omni, an extended library under the vLLM project for multimodal generation. vLLM Omni Neuron adds the Neuron backend that library needs.
The plugin provides Neuron-optimized reference implementations — including NKI
kernels, custom compilation, and hardware-aware scheduling — for models whose
upstream implementations live in vLLM Omni. It uses native PyTorch
(torch.compile) through the libtorch_neuronx_lite package, the same native
PyTorch infrastructure used by vLLM Neuron.
Note
vLLM Omni Neuron is in Beta and under active development. It is supported on Trainium 2 (trn2) and Trainium 3 (trn3) instances only.
The source code for the vLLM Omni Neuron plugin is hosted in the vLLM Omni Neuron GitHub repository.
Get started#
This document is relevant for: Trn2, Trn3