This document is relevant for: Trn2, Trn3

Concepts & architecture#

How vLLM Omni Neuron works under the hood — engine and model integration, and context parallelism for diffusion inference on Trainium.

vLLM Omni Neuron overview

Engine, worker, and model integration — how the plugin attaches to vLLM Omni.

Context parallelism

Sequence-parallel context handling for diffusion inference on Neuron.

This document is relevant for: Trn2, Trn3