https://youtu.be/XoGvCBRnwLs?si=0SthOvsy0EHTDoXB
ربع ساعت اول بحث غیرفنیست(کجا بودم چکار کردم و چی شد و چرا...و معنی یادگیری، چرا باید هنوز یاد گرفت و مسیر و تجارب خودش و....)
کد:
https://github.com/hkproj/torchfeather
Non exhaustive list of papers cited:
Attention Is All You Need - https://arxiv.org/abs/...
GShard - https://arxiv.org/abs/...
Scaling Laws for Fine-Grained Mixture of Experts - https://arxiv.org/abs/...
GPipe: https://arxiv.org/abs/...
DeepSeek V2: https://arxiv.org/abs/...
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts - https://arxiv.org/abs/...
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM - https://arxiv.org/abs/...
Fast Transformer Decoding: One Write-Head is All You Need - https://arxiv.org/abs/...
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints - https://arxiv.org/abs/...
Zero Bubble Pipeline Parallelism - https://arxiv.org/abs/...
DeepSeek V3 - https://arxiv.org/abs/...
Training Compute-Optimal Large Language Models - https://arxiv.org/abs/...
Scaling Laws for Neural Language Models - https://arxiv.org/abs/...
Chapters
00:00:00 - Introduction
00:16:50 - Model Architecture, Parameters and Training FLOPs
00:51:57 - RoPE from First Principles
01:11:25 - Implementing RoPE and YaRN
01:50:16 - Building the Transformer and Weight Initialization
02:19:12 - Attention, the KV Cache and Arithmetic Intensity
03:02:56 - Coding Multi-head Latent Attention (MLA)
03:15:35 - Block Matrix Multiplication and MLA Internals
03:34:58 - Deriving MLA and Decoupled RoPE
04:10:54 - MLA Weight Absorption
04:32:59 - Autograd and the Mathematics of Distributed Training
05:04:19 - Distributed Computation Graphs and DDP
05:10:47 - Building the Training Loop
06:06:17 - Pipeline Parallelism from First Principles
06:26:52 - Pipeline Schedules: GPipe, 1F1B and Zero Bubble
07:10:29 - Datasets, Tokenization and Data Parallelism
07:42:10 - Coding Pipeline Parallelism
08:37:34 - Device Meshes and Combining PP with DP
09:48:52 - Distributed Communication Collectives
10:49:41 - Implementing Device Meshes, DDP and FSDP
12:13:08 - Tensor Parallelism from First Principles
13:44:57 - Coding Tensor Parallelism
14:47:43 - Context Parallelism and Ring Attention
15:41:22 - Metrics, Optimizers, Schedulers and Checkpointing
16:00:33 - Combining Parallelism in the Training Loop
16:49:58 - Mixture of Experts from First Principles
18:02:52 - Tensor Parallelism for MoE
18:36:24 - Expert Parallelism
19:01:56 - All-to-All Token Dispatch and Combine
19:27:38 - Expert Tensor Parallelism