This document is relevant for: Trn2, Trn3
Tutorials
End-to-end guided walkthroughs for specific deployment scenarios and performance optimization.
Deploy Llama 3.3 70B (FP8)
Deploy Llama 3.3 70B in static FP8 on a single Trn2/Trn3 instance.
EAGLE3 speculative decoding (Llama 3.1)
Run Llama 3.1 8B with an EAGLE3 draft model for higher throughput.
EAGLE3 speculative decoding (GPT-OSS)
Run GPT-OSS-120B with an EAGLE3 draft model for higher throughput.
Disaggregated inference: 1P1D and xPyD
Configure disaggregated inference topologies.
Deploy gpt-oss
Deploy gpt-oss 20B and 120B, single-instance or disaggregated.
Prefix caching benchmark
Measure TTFT improvement from prefix caching with GPT-OSS.
Deploy Qwen3-VL-32B
Serve the multimodal Qwen3-VL-32B model (BF16 or MXFP8).
Disaggregated encoder: 1E1PD and xEyPD
Configure encoder-disaggregated (EPD) multimodal topologies.
Deploy Qwen3-Embedding-8B
Serve embeddings via /v1/embeddings with a pooling model.
This document is relevant for: Trn2, Trn3