This document is relevant for: Trn2, Trn3

Tutorials#

End-to-end guided walkthroughs for specific deployment scenarios and performance optimization.

Disaggregated inference: 1P1D and xPyD

Configure disaggregated inference topologies.

Deploy gpt-oss

Deploy gpt-oss 20B and 120B, single-instance or disaggregated.

Prefix caching benchmark

Measure TTFT improvement from prefix caching with GPT-OSS.

Deploy Qwen3-VL-32B

Serve the multimodal Qwen3-VL-32B model.

This document is relevant for: Trn2, Trn3