عُمَر جَمیل با ویدیو ۱۹ ساعتِ بعد از ۱ سال برگشت https://youtu.be/XoGvCBRnwLs?si=0SthOvsy0EHTDoXB ربع ساعت اول بحث غیرفنیست(کجا بودم چکار کردم و چی شد و چرا...و معنی یادگیری، چرا باید هنوز یاد گرفت و مسیر و تجارب خودش و....) کد: https://github.com/hkproj/torchfeather Non exhaustive list of papers cited: Attention Is All You Need - https://arxiv.org/abs/... GShard - https://arxiv.org/abs/... Scaling Laws for Fine-Grained Mixture of Experts - https://arxiv.org/abs/... GPipe: https://arxiv.org/abs/... DeepSeek V2: https://arxiv.org/abs/... Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts - https://arxiv.org/abs/... Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM - https://arxiv.org/abs/... Fast Transformer Decoding: One Write-Head is All You Need - https://arxiv.org/abs/... GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints - https://arxiv.org/abs/... Zero Bubble Pipeline Parallelism - https://arxiv.org/abs/... DeepSeek V3 - https://arxiv.org/abs/... Training Compute-Optimal Large Language Models - https://arxiv.org/abs/... Scaling Laws for Neural Language Models - https://arxiv.org/abs/... Chapters 00:00:00 - Introduction 00:16:50 - Model Architecture, Parameters and Training FLOPs 00:51:57 - RoPE from First Principles 01:11:25 - Implementing RoPE and YaRN 01:50:16 - Building the Transformer and Weight Initialization 02:19:12 - Attention, the KV Cache and Arithmetic Intensity 03:02:56 - Coding Multi-head Latent Attention (MLA) 03:15:35 - Block Matrix Multiplication and MLA Internals 03:34:58 - Deriving MLA and Decoupled RoPE 04:10:54 - MLA Weight Absorption 04:32:59 - Autograd and the Mathematics of Distributed Training 05:04:19 - Distributed Computation Graphs and DDP 05:10:47 - Building the Training Loop 06:06:17 - Pipeline Parallelism from First Principles 06:26:52 - Pipeline Schedules: GPipe, 1F1B and Zero Bubble 07:10:29 - Datasets, Tokenization and Data Parallelism 07:42:10 - Coding Pipeline Parallelism 08:37:34 - Device Meshes and Combining PP with DP 09:48:52 - Distributed Communication Collectives 10:49:41 - Implementing Device Meshes, DDP and FSDP 12:13:08 - Tensor Parallelism from First Principles 13:44:57 - Coding Tensor Parallelism 14:47:43 - Context Parallelism and Ring Attention 15:41:22 - Metrics, Optimizers, Schedulers and Checkpointing 16:00:33 - Combining Parallelism in the Training Loop 16:49:58 - Mixture of Experts from First Principles 18:02:52 - Tensor Parallelism for MoE 18:36:24 - Expert Parallelism 19:01:56 - All-to-All Token Dispatch and Combine 19:27:38 - Expert Tensor Parallelism
عُمَر جَمیل با ویدیو ۱۹ ساعتِ بعد از ۱ سال برگشت… — Bit x Wisdom — TG.ME
August 19, 2026 33