This document is relevant for: Trn2, Trn3
Model Recipes
Production-ready deployment recipes for specific models on AWS Trainium and Inferentia. Each recipe includes instance sizing, configuration, and performance guidance.
Deploy Llama 3
Model recipe for the Llama 3 family (1B, 8B, 70B) on Trn2/Trn3.
Deploy GPT-OSS
Model recipe for GPT-OSS 20B and 120B (MoE) on Trn2/Trn3.
Deploy Qwen3-VL 32B
Model recipe for Qwen3-VL 32B (multimodal) on Trn2/Trn3.
Deploy Qwen3-Embedding 8B
Model recipe for Qwen3-Embedding 8B (pooling / embeddings) on Trn2/Trn3.
This document is relevant for: Trn2, Trn3