Skip to content
Yihua Zhang
Articles
Posts
EN
中
◐
Articles
All
13
Post-training
6
Systems
5
Music
2
Archive
20
Series
3
Anatomy of DeepSeek
3
The GRPO lineage
13 articles
2026-02-13
Vogel im Käfig: How Sawano Builds Hope and Despair from the Smallest Interval
Music
·
zh
Sawano
2025-10-02
A Role Shift for AI Infra: From Foundational Support to a Core Engine of Innovation
AI Market Insights
Systems
·
en
2025-08-08
From GRPO to DAPO and GSPO: What, Why and How
Post-training
·
en / zh
·
18 equations
The GRPO lineage 2
GRPO
2025-07-02
Re-understanding KL Approximation from an RL-for-LLM Lens: Notes on “Approximating KL Divergence
Post-training
·
en / zh
·
18 equations
The GRPO lineage 3
GRPO
2025-04-10
Decorators in Machine Learning Projects
Systems
·
en / zh
Python
2025-03-08
Bauklötze: A Musical Dissection
积木崩塌时的命运回响,泽野弘之用音符砌筑的巨人悲歌
Music
·
zh
Sawano
2025-02-27
DualPipe Explained: A Comprehensive Guide to DualPipe That Anyone Can Understand—Even Without a Distributed Background
Systems
·
en / zh
Anatomy of DeepSeek 3
Distributed
2025-02-11
Navigating the RLHF Landscape: From Policy Gradients to PPO, GAE, and DPO for LLM Alignment
Post-training
·
en / zh
·
88 equations
RLHF
2025-02-07
DeepSeek-R1 Dissection: Understanding PPO & GRPO Without Any Prior Reinforcement Learning Knowledge
Post-training
·
en / zh
·
16 equations
The GRPO lineage 1
GRPO
RLHF
2025-02-02
Why Cache 32 Heads When One Latent Variable Suffices? A Theory-to-Code Guide to DeepSeek’s MLA for KV-Cache
Systems
·
en / zh
·
24 equations
Anatomy of DeepSeek 2
KV-Cache
2025-01-20
From Zero to Reasoning Hero: How DeepSeek-R1 Leverages Reinforcement Learning to Master Complex Reasoning
Post-training
·
en / zh
·
2 equations
Anatomy of DeepSeek 1
GRPO
2025-01-15
A Review on the Evolvement of Load Balancing Strategy in MoE LLMs: Pitfalls and Lessons
Systems
·
en / zh
·
31 equations
MoE
2024-12-15
Patching the Foundation Models: Pitfalls and Pains in Machine Unlearning
Post-training
·
en / zh
·
14 equations
Unlearning