Skip to content
Yihua Zhang
Articles
Posts
EN
中
◐
All
GRPO
4 articles
2025-08-08
From GRPO to DAPO and GSPO: What, Why and How
Post-training
·
en / zh
·
18 equations
The GRPO lineage 2
GRPO
2025-07-02
Re-understanding KL Approximation from an RL-for-LLM Lens: Notes on “Approximating KL Divergence
Post-training
·
en / zh
·
18 equations
The GRPO lineage 3
GRPO
2025-02-07
DeepSeek-R1 Dissection: Understanding PPO & GRPO Without Any Prior Reinforcement Learning Knowledge
Post-training
·
en / zh
·
16 equations
The GRPO lineage 1
GRPO
RLHF
2025-01-20
From Zero to Reasoning Hero: How DeepSeek-R1 Leverages Reinforcement Learning to Master Complex Reasoning
Post-training
·
en / zh
·
2 equations
Anatomy of DeepSeek 1
GRPO