عُمَر جَمیل با ویدیو ۱۹ ساعتِ بعد از ۱ سال برگشت… — Bit x Wisdom — TG.ME

عُمَر جَمیل با ویدیو ۱۹ ساعتِ بعد از ۱ سال برگشت
https://youtu.be/XoGvCBRnwLs?si=0SthOvsy0EHTDoXB

ربع ساعت اول بحث غیرفنی‌ست(کجا بودم چکار کردم و چی شد و چرا...و معنی یادگیری، چرا باید هنوز یاد گرفت و مسیر و تجارب خودش و....)


کد:
https://github.com/hkproj/torchfeather



Non exhaustive list of papers cited:

Attention Is All You Need - https://arxiv.org/abs/...​
GShard - https://arxiv.org/abs/...​
Scaling Laws for Fine-Grained Mixture of Experts - https://arxiv.org/abs/...​
GPipe: https://arxiv.org/abs/...​
DeepSeek V2: https://arxiv.org/abs/...​
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts - https://arxiv.org/abs/...​
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM - https://arxiv.org/abs/...​
Fast Transformer Decoding: One Write-Head is All You Need - https://arxiv.org/abs/...​
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints - https://arxiv.org/abs/...​
Zero Bubble Pipeline Parallelism - https://arxiv.org/abs/...​
DeepSeek V3 - https://arxiv.org/abs/...​
Training Compute-Optimal Large Language Models - https://arxiv.org/abs/...​
Scaling Laws for Neural Language Models - https://arxiv.org/abs/...​



Chapters

00:00:00​ - Introduction
00:16:50​ - Model Architecture, Parameters and Training FLOPs
00:51:57​ - RoPE from First Principles
01:11:25​ - Implementing RoPE and YaRN
01:50:16​ - Building the Transformer and Weight Initialization
02:19:12​ - Attention, the KV Cache and Arithmetic Intensity
03:02:56​ - Coding Multi-head Latent Attention (MLA)
03:15:35​ - Block Matrix Multiplication and MLA Internals
03:34:58​ - Deriving MLA and Decoupled RoPE
04:10:54​ - MLA Weight Absorption
04:32:59​ - Autograd and the Mathematics of Distributed Training
05:04:19​ - Distributed Computation Graphs and DDP
05:10:47​ - Building the Training Loop
06:06:17​ - Pipeline Parallelism from First Principles
06:26:52​ - Pipeline Schedules: GPipe, 1F1B and Zero Bubble
07:10:29​ - Datasets, Tokenization and Data Parallelism
07:42:10​ - Coding Pipeline Parallelism
08:37:34​ - Device Meshes and Combining PP with DP
09:48:52​ - Distributed Communication Collectives
10:49:41​ - Implementing Device Meshes, DDP and FSDP
12:13:08​ - Tensor Parallelism from First Principles
13:44:57​ - Coding Tensor Parallelism
14:47:43​ - Context Parallelism and Ring Attention
15:41:22​ - Metrics, Optimizers, Schedulers and Checkpointing
16:00:33​ - Combining Parallelism in the Training Loop
16:49:58​ - Mixture of Experts from First Principles
18:02:52​ - Tensor Parallelism for MoE
18:36:24​ - Expert Parallelism
19:01:56​ - All-to-All Token Dispatch and Combine
19:27:38​ - Expert Tensor Parallelism
August 19, 2026 33