This document is relevant for: Trn2, Trn3

Get started#

Set up vLLM Neuron and run your first inference — from install through your first API request.

Setup guide

Install and configure vLLM Neuron on Trainium or Inferentia.

Online serving quickstart

Launch an OpenAI-compatible API server and send your first chat request.

Offline serving quickstart

Run batch inference with the vllm.LLM Python API.

Migration from NxD Inference

Migrate existing NxDI deployments to vLLM Neuron.

This document is relevant for: Trn2, Trn3