This document is relevant for: Trn2, Trn3

Deploy & serve#

Configure, tune, and operate vLLM Neuron for production workloads — features, profiling, and configuration reference.

Features guide

Configure and tune all serving features — bucketing, quantization, DI, speculation, and more.

Configuration reference

All Neuron-specific options in additional_config and environment variables.

Profiling workloads

Capture Neuron Runtime profiles via built-in profiler endpoints.

This document is relevant for: Trn2, Trn3