This document is relevant for: Trn2, Trn3

Feature and model compatibility matrix#

Use this page to check feature support before configuring a deployment. For configuration details, see the features guide.

Feature/model matrix#

Feature

GPT-OSS

Qwen3-VL

Continuous batching

Segmented prefill

Prefix caching (APC)

Speculative decoding (EAGLE3)

FP8 weight quantization (static)

MXFP8 weight quantization

MXFP4 weight quantization

✅ ¹

KV cache FP8

Multimodal (image input)

Disaggregated inference (1P1D / xPyD)

Structured outputs / tool calling

On-device sampling

Tensor parallelism

Data parallelism

Expert parallelism

N/A

Vision encoder parallelism

N/A

¹ Trn3 only.

Legend#

  • ✅ Supported and tested

  • ❌ Not supported

  • N/A Not applicable to this architecture