Skip to content

Course catalog

1 course matching “attention”

Reset search: attention
Abstract cover artwork for Transformers and Attention, Implemented Line by Line, a Deep Learning course Deep Learning
Deep Learning advanced

Transformers and Attention, Implemented Line by Line

Build a transformer from scratch: scaled dot-product attention, multi-head projections, positional encodings, KV caching and the modern variants (RoPE, grouped-query attention, RMSNorm) that make inference cheap. You will train a small model, profile it, and understand precisely where the FLOPs go.

HT Hiroshi Tanaka 4.9

$109.00

17 hours