Tags
- AI Infra 2
- Alignment 1
- AllToAll 1
- Angular Update 1
- Checkpoint 1
- Communication 1
- Communication Overlap 2
- CUDA 3
- Dataset 1
- DDP 1
- DeepEP 1
- Distributed 1
- Distributed Training 2
- Effective Learning Rate 1
- Expert Parallelism 1
- Export 1
- Femtotron 1
- FP8 1
- FSDP 1
- GEMM 2
- GQA 1
- Hyperparameter Transfer 1
- Initialization 1
- Kernel Fusion 1
- KV Cache 2
- LLM 28
- LLM System 21
- LLM Theory 6
- MCore 3
- Meeting Notes 1
- Megatron 8
- MoE 4
- MuP 3
- NCCL 3
- Normalization 1
- NVLink 1
- Optimization 3
- PagedAttention 1
- PD Disaggregation 1
- Pipeline Parallel 3
- Pretraining 3
- Probability 1
- PyTorch 1
- Ray 1
- RDMA 1
- ReduceScatter 1
- Resharding 1
- RL 1
- RLHF 1
- Schedule 3
- Serving 3
- SGD 1
- SGLang 2
- Shared Expert 1
- SMD 2
- Training 12
- Training Framework 9
- Transformer Engine 2
- Ulysses 1
- VLLM 1
- Weight 1
- Weight Decay 2
- 人生感悟 1
- 内存管理 1
- 分布式训练 1
- 哲学 1
- 哲学研究 1
- 基础知识速查 1
- 性能 1
- 时间管理 1
- 生成式推荐 1
- 生成式搜索 1
- 维特根斯坦 1
- 职业发展 1
- 行业经验 1
- 读书笔记 1