Explain GRPO for LLM Reinforcement Learning and How It Differs from PPO
Company: ByteDance
Role: Applied Scientist
Category: Machine Learning
Difficulty: medium
Interview Round: Technical Screen
Explain GRPO (Group Relative Policy Optimization), a reinforcement learning algorithm used to fine-tune large language models, for example to improve their reasoning. Walk through one training step: what is sampled, how rewards become advantages, and what objective the policy optimizes. Then explain how GRPO differs from PPO as used in RLHF, and what that difference buys and costs.
```hint What the critic was for
In PPO, a learned value model tells you whether a response did better or worse than expected. Ask what cheaper reference point is available when you can sample many responses to the same prompt.
```
```hint One reward, many tokens
With outcome rewards, a whole response receives a single number. Decide how that number becomes a learning signal for each token, and what keeps the policy from drifting too far.
```
### Clarifying Questions
- Should the answer cover only outcome rewards (one score per response), or also step-level (process) rewards?
- Do the rewards come from a learned reward model, or from rule-based checks such as verifying a math answer?
- How much detail is expected: the full clipped objective with its KL term, or the intuition and the algorithm?
### What a Strong Answer Covers
- The setting: a policy LLM, a frozen reference model, and a reward source
- One training step: sampling a group of responses per prompt, scoring them, and normalizing rewards within the group into advantages
- The objective: the per-token probability ratio, clipping, the KL penalty toward the reference model, and how the terms are averaged
- The contrast with PPO: no value model, and the consequences for memory, stability and credit assignment
- Failure modes such as groups with identical rewards, biases introduced by the normalization, and reward hacking, with mitigations
### Follow-up Questions
- If every response in a group gets the same reward, what gradient does that prompt produce, and how would you handle such prompts?
- What biases do dividing by the group's reward standard deviation and by each response's length introduce?
- How would you choose the group size, and what does it trade off?
- GRPO adds the KL term directly to the loss rather than subtracting it from the reward. What is the difference?
Overview: An ML question asking you to explain Group Relative Policy Optimization (GRPO) for fine-tuning large language models with reinforcement learning. It tests your grasp of group-normalized advantages, the clipped objective with a KL penalty, the contrast with PPO and its value model, and common failure modes.