This document is relevant for: Trn2, Trn3

Tutorials#

End-to-end guided walkthroughs for specific deployment scenarios and performance optimization.

Deploy Llama 3.3 70B (FP8)

Deploy Llama 3.3 70B in static FP8 on a single Trn2/Trn3 instance.

EAGLE3 speculative decoding (Llama 3.1)

Run Llama 3.1 8B with an EAGLE3 draft model for higher throughput.

EAGLE3 speculative decoding (GPT-OSS)

Run GPT-OSS-120B with an EAGLE3 draft model for higher throughput.

Disaggregated inference: 1P1D and xPyD

Configure disaggregated inference topologies.

Deploy gpt-oss

Deploy gpt-oss 20B and 120B, single-instance or disaggregated.

Prefix caching benchmark

Measure TTFT improvement from prefix caching with GPT-OSS.

Deploy Qwen3-VL-32B

Serve the multimodal Qwen3-VL-32B model (BF16 or MXFP8).

Disaggregated encoder: 1E1PD and xEyPD

Configure encoder-disaggregated (EPD) multimodal topologies.

Deploy Qwen3-Embedding-8B

Serve embeddings via /v1/embeddings with a pooling model.

This document is relevant for: Trn2, Trn3