This document is relevant for: Trn2, Trn3

vLLM integration#

Design documentation for how vLLM Neuron integrates with the vLLM core framework. For user-facing configuration, see the features guide and configuration options.

Topic

Description

KV cache integration

KV cache integration points with vLLM

Async scheduling and execution

Async scheduling and execution design

Metrics

Production metrics design

Neuron profiling

Profiling integration

Neuron scheduler

Holdback queue and admission control

Prefix caching

Prefill segmentation and KV reuse

Disaggregated inference

DI architecture, NIXL transport, hybrid TP

This document is relevant for: Trn2, Trn3