عُمَر جَمیل با ویدیو ۱۹ ساعتِ بعد از ۱ سال برگشت… — Bit x Wisdom — TG.ME

عُمَر جَمیل با ویدیو ۱۹ ساعتِ بعد از ۱ سال برگشت https://youtu.be/XoGvCBRnwLs?si=0SthOvsy0EHTDoXB ربع ساعت اول بحث غیرفنی‌ست(کجا بودم چکار کردم و چی شد و چرا...و معنی یادگیری، چرا باید هنوز یاد گرفت و مسیر و تجارب خودش و....) کد: https://github.com/hkproj/torchfeather Non exhaustive list of papers cited: Attention Is All You Need - https://arxiv.org/abs/...​ GShard - https://arxiv.org/abs/...​ Scaling Laws for Fine-Grained Mixture of Experts - https://arxiv.org/abs/...​ GPipe: https://arxiv.org/abs/...​ DeepSeek V2: https://arxiv.org/abs/...​ Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts - https://arxiv.org/abs/...​ Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM - https://arxiv.org/abs/...​ Fast Transformer Decoding: One Write-Head is All You Need - https://arxiv.org/abs/...​ GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints - https://arxiv.org/abs/...​ Zero Bubble Pipeline Parallelism - https://arxiv.org/abs/...​ DeepSeek V3 - https://arxiv.org/abs/...​ Training Compute-Optimal Large Language Models - https://arxiv.org/abs/...​ Scaling Laws for Neural Language Models - https://arxiv.org/abs/...​ Chapters 00:00:00​ - Introduction 00:16:50​ - Model Architecture, Parameters and Training FLOPs 00:51:57​ - RoPE from First Principles 01:11:25​ - Implementing RoPE and YaRN 01:50:16​ - Building the Transformer and Weight Initialization 02:19:12​ - Attention, the KV Cache and Arithmetic Intensity 03:02:56​ - Coding Multi-head Latent Attention (MLA) 03:15:35​ - Block Matrix Multiplication and MLA Internals 03:34:58​ - Deriving MLA and Decoupled RoPE 04:10:54​ - MLA Weight Absorption 04:32:59​ - Autograd and the Mathematics of Distributed Training 05:04:19​ - Distributed Computation Graphs and DDP 05:10:47​ - Building the Training Loop 06:06:17​ - Pipeline Parallelism from First Principles 06:26:52​ - Pipeline Schedules: GPipe, 1F1B and Zero Bubble 07:10:29​ - Datasets, Tokenization and Data Parallelism 07:42:10​ - Coding Pipeline Parallelism 08:37:34​ - Device Meshes and Combining PP with DP 09:48:52​ - Distributed Communication Collectives 10:49:41​ - Implementing Device Meshes, DDP and FSDP 12:13:08​ - Tensor Parallelism from First Principles 13:44:57​ - Coding Tensor Parallelism 14:47:43​ - Context Parallelism and Ring Attention 15:41:22​ - Metrics, Optimizers, Schedulers and Checkpointing 16:00:33​ - Combining Parallelism in the Training Loop 16:49:58​ - Mixture of Experts from First Principles 18:02:52​ - Tensor Parallelism for MoE 18:36:24​ - Expert Parallelism 19:01:56​ - All-to-All Token Dispatch and Combine 19:27:38​ - Expert Tensor Parallelism

August 19, 2026 33