Deep Learning
Transformers and Attention, Implemented Line by Line
Build a transformer from scratch: scaled dot-product attention, multi-head projections, positional encodings, KV caching and the modern variants (RoPE, grouped-query attention, RMSNorm) that make inference cheap. You will train a small model, profile it, and understand precisely where the FLOPs go.
$109.00
17 hours