Skip to content
Yihua Zhang
Articles
Posts
EN
中
◐
All
RLHF
2 articles
2025-02-11
Navigating the RLHF Landscape: From Policy Gradients to PPO, GAE, and DPO for LLM Alignment
Post-training
·
en / zh
·
88 equations
RLHF
2025-02-07
DeepSeek-R1 Dissection: Understanding PPO & GRPO Without Any Prior Reinforcement Learning Knowledge
Post-training
·
en / zh
·
16 equations
The GRPO lineage 1
GRPO
RLHF