This document is relevant for: Trn2, Trn3
Concepts & architecture#
How vLLM Omni Neuron works under the hood — engine and model integration, and context parallelism for diffusion inference on Trainium.
This document is relevant for: Trn2, Trn3
This document is relevant for: Trn2, Trn3
How vLLM Omni Neuron works under the hood — engine and model integration, and context parallelism for diffusion inference on Trainium.
This document is relevant for: Trn2, Trn3