This document is relevant for: Trn2, Trn3

Model Recipes#

Production-ready deployment recipes for specific models on AWS Trainium and Inferentia. Each recipe includes instance sizing, configuration, and performance guidance.

Deploy Llama 3

Model recipe for the Llama 3 family (1B, 8B, 70B) on Trn2/Trn3.

Deploy GPT-OSS

Model recipe for GPT-OSS 20B and 120B (MoE) on Trn2/Trn3.

Deploy Qwen3-VL 32B

Model recipe for Qwen3-VL 32B (multimodal) on Trn2/Trn3.

Deploy Qwen3-Embedding 8B

Model recipe for Qwen3-Embedding 8B (pooling / embeddings) on Trn2/Trn3.

This document is relevant for: Trn2, Trn3