This document is relevant for: Trn2, Trn3

Model development#

Onboard new models, evaluate and debug accuracy, and study the Neuron-optimized kernels used in diffusion model development on AWS Trainium.

Onboard a new model

Implement the Neuron pipeline and model components, register them, compile, and validate.

Accuracy evaluation and debugging

Evaluate output quality, detect regressions, and debug numerical accuracy.

Kernel implementations

Per-kernel implementation references for the Neuron-optimized kernels used by vLLM Omni Neuron.

Optimizing offline video generation

Reason about the three-stage cost model and choose a parallelism scheme to maximize quality per compute.

This document is relevant for: Trn2, Trn3