This document is relevant for: Trn2, Trn3
How to set up vLLM Neuron#
Overview#
This topic discusses how to install and configure the vLLM Neuron plugin on AWS Trainium or Inferentia so you can serve large language models with vLLM on Neuron. When you have completed it, you will have a working environment ready for the online or offline serving quickstarts.
Note
This guide is for the vLLM Neuron plugin introduced in Neuron 2.31. If you are using vLLM with the NxD Inference library, see the vLLM + NxD Inference documentation. For migration guidance from NxD Inference to vLLM Neuron, see migrate from NxD Inference to vLLM Neuron.
Prerequisites#
Instance: A supported Trainium EC2 instance (for example,
trn2.48xlarge).Neuron SDK version: Neuron 2.31 or later.
Python: Python 3.11 or later.
Hugging Face access (optional): An accepted model license and a Hugging Face token if you plan to pull gated models.
Instructions#
Option A: Install from source#
Clone the repository and install the plugin from source:
git clone https://github.com/vllm-project/vllm-neuron.git
cd vllm-neuron
pip install --extra-index-url=https://pip.repos.neuron.amazonaws.com -e .
This installs the vLLM Neuron plugin along with vLLM and all required Neuron SDK packages.
Option B: Use the Neuron DLAMI#
Launch an instance with the Neuron Multi-Framework Deep Learning AMI. It ships with the Neuron SDK, PyTorch NeuronX, and a pre-built virtual environment that includes vLLM Neuron.
Launch an instance with the DLAMI by following the Neuron DLAMI setup instructions. After you connect, activate the vLLM virtual environment:
source /opt/aws_neuronx_venv_pytorch_inference_vllm_0_21_0_1_0_0/bin/activate
Option C: Use a container#
Use the vLLM Inference NeuronX DLC which bundles the SDK and all dependencies.
Verify the installation#
Confirm that vLLM discovers the Neuron platform plugin:
python -c "import vllm; from vllm.platforms import current_platform; print(current_platform.device_name)"
The output should be:
neuron
Confirm the Neuron runtime is visible:
neuron-ls
You should see output listing the available NeuronCores on your instance.
Configure environment variables#
The following environment variables control common behaviors:
Variable |
Purpose |
Default |
|---|---|---|
|
Root directory for the Neuron compile cache. Set to local NVMe for best performance. |
|
|
Plugin log verbosity ( |
|
|
Hugging Face access token for gated model downloads. |
(none) |
export VLLM_CACHE_ROOT=/local/cache/vllm
export HF_TOKEN=hf_your_token_here
Confirm your work#
Run a minimal vllm serve command to confirm the full stack is wired up. This
example uses gpt-oss-20b with a small context and a single compiled bucket to
keep compilation time short:
vllm serve openai/gpt-oss-20b \
--tensor-parallel-size 8 \
--max-num-seqs 1 \
--max-model-len 1024 \
--max-num-batched-tokens 512 \
--hf-overrides '{"quantization_config": {}}' \
--additional-config '{"neuron_config": {"num_batched_tokens_buckets": [512], "num_seqs_buckets": [1]}}'
Wait for the server to print Uvicorn running on .... Then, in a second terminal
on the same instance:
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-20b",
"messages": [{"role": "user", "content": "Hello, Neuron!"}],
"max_tokens": 16
}'
A JSON response containing generated text confirms that the plugin, Neuron
runtime, and model compilation are all working. Stop the server with Ctrl+C.
Next steps#
Quickstart: Online serving — Launch an OpenAI-compatible API server.
Quickstart: Offline serving — Run batch inference with the Python API.
Features guide — Configure features like prefix caching, speculative decoding, data parallelism, and quantization.
Configuration options — Full parameter reference.
Common issues#
Plugin fails to load the Neuron platform#
Possible solution: Confirm you activated the correct virtual environment and that
pip show vllm-neuronreports version 0.21.0.1.0.0 or later. If you installed vLLM separately before the plugin, reinstall the plugin to ensure version alignment.
neuron-ls shows no devices#
Possible solution: Verify you are on a Neuron-supported instance type (
trn2ortrn3). On a fresh instance, the Neuron runtime may need a reboot after driver installation. Checkdmesg | grep neuronfor driver messages.
Compilation times out or is very slow on first run#
Possible solution: First-time model compilation on Neuron can take several minutes depending on model size. Subsequent runs use the compile cache. Ensure
VLLM_CACHE_ROOTpoints to fast local storage (NVMe) rather than a network filesystem.
Python version mismatch#
Possible solution: The vLLM Neuron plugin requires Python 3.11 or later. Run
python --versionto check. The DLAMI virtual environment ships with a compatible Python version.