GRPO in NumPy: Group-Relative Advantages, Process Rewards and Shaped Rewards

Quick Overview

Implement the core of Group Relative Policy Optimization in a NumPy notebook: group-normalized advantages, the clipped objective with a KL penalty, per-token credit from process rewards, and shaped rewards. Tests fluency with array axes and broadcasting and an understanding of how reward normalization interacts with shaping terms.

GRPO in NumPy: Group-Relative Advantages, Process Rewards and Shaped Rewards

Company: Amazon

Role: Machine Learning Engineer

Category: Machine Learning

Difficulty: medium

Interview Round: Onsite

In a notebook, implement the core computations of Group Relative Policy Optimization (GRPO) with NumPy, then extend them to process rewards and to shaped rewards. For each prompt the policy samples a group of `G` completions; each completion is scored, and GRPO judges every completion relative to the other completions of the same prompt instead of training a separate value network. Assume the notebook provides these arrays for a batch of `P` prompts, with completions padded to length `L`: - `rewards`, shape `(P, G)`: the outcome reward of each completion; - `logp_new`, `logp_old`, `logp_ref`, shape `(P, G, L)`: log-probabilities of the sampled tokens under the policy being trained, the policy that generated the samples, and a frozen reference policy; - `mask`, shape `(P, G, L)`: 1 for completion tokens and 0 for padding. ### Clarifying Questions - Should the group standard deviation be the population or the sample estimate, and how should a group whose rewards are all equal be handled? - Is the loss averaged over each completion's tokens and then over completions, or over all tokens in the batch at once? - Which KL estimator should the penalty use, and what are the clip range and the KL coefficient? ### Part 1 — Group-relative advantages and the GRPO loss Compute each completion's group-relative advantage from the outcome rewards, assign it to every token of that completion, and compute the clipped GRPO objective with a KL penalty toward the reference policy, returned as a scalar loss. Vectorize: no Python loops over prompts, completions or tokens. ```hint Reduce, then broadcast back Every statistic here is taken over one specific axis. Keep that axis when you reduce, so the result lines up with the array you combine it with, and let the mask handle padding. ``` #### What This Part Should Cover - Group normalization along the correct axis, including the zero-variance case. - The probability ratio, clipping and minimum in the surrogate, the KL term, and correct masking. - Aggregation over tokens and completions, and the sign of the returned loss. ### Part 2 — Process rewards A process reward model now scores intermediate reasoning steps. Completion `i` has a list of step rewards and, for each step, the index of the token where that step ends. For one prompt's group, turn these into per-token advantages the way GRPO does under process supervision. ```hint What can a token affect A token can only influence the part of the completion that comes after it. Decide what that implies for which step rewards should count toward its advantage. ``` #### Clarifying Questions for this Part - Are step rewards normalized within each completion or across all steps in the group? - Should the completion's outcome reward also contribute to the token advantages? #### What This Part Should Cover - Normalization of the step rewards and the rule that maps them to per-token advantages. - A vectorized implementation for completions with different numbers of steps. - How process rewards change credit assignment compared with a single outcome reward, and how they can be gamed. ### Part 3 — Shaped rewards Add shaping terms to the reward, for example a bonus for following the required output format or a penalty tied to completion length, and show how they enter the advantage computation. ```hint Watch what normalization does Group normalization rescales whatever reward it is given. Check what a small shaping bonus turns into in a group where every completion has the same task reward. ``` #### Clarifying Questions for this Part - Which shaping terms are defined, and are they given per completion or per token? - Should shaping be combined with the task reward before or after group normalization? #### What This Part Should Cover - Where shaping enters relative to group normalization, and how its weight is set. - The interaction between shaping and normalization, including groups with identical task rewards. - Risks of shaping, such as reward hacking or a changed optimal policy, and how to detect them in training logs. ### What a Strong Answer Covers - Correct axes, broadcasting and masking throughout, verified on small hand-checked arrays. - Numerical care: an epsilon in normalization, ratios computed from log-probability differences, and a non-negative KL estimate. - A clear explanation of why GRPO needs no value network and what the group baseline costs. - Reasoned choices for every ambiguous convention: the standard deviation, loss aggregation, and where shaping is added. ### Follow-up Questions - Why can dividing by the group standard deviation bias training toward some prompts, and what changes if you drop it? - If every completion in a group receives the same reward, what does that group contribute to the gradient, and what would you do about it? - How would you choose the group size under a fixed sampling budget? - How does averaging the loss per completion treat long and short completions differently?

Overview: Implement the core of Group Relative Policy Optimization in a NumPy notebook: group-normalized advantages, the clipped objective with a KL penalty, per-token credit from process rewards, and shaped rewards. Tests fluency with array axes and broadcasting and an understanding of how reward normalization interacts with shaping terms.

|Home/Machine Learning/Amazon
Amazon logo
Amazon
Jan 3, 2026
mediumMachine Learning EngineerOnsiteMachine Learning
0
0

In a notebook, implement the core computations of Group Relative Policy Optimization (GRPO) with NumPy, then extend them to process rewards and to shaped rewards. For each prompt the policy samples a group of G completions; each completion is scored, and GRPO judges every completion relative to the other completions of the same prompt instead of training a separate value network.

Assume the notebook provides these arrays for a batch of P prompts, with completions padded to length L:

  • rewards , shape (P, G) : the outcome reward of each completion;
  • logp_new , logp_old , logp_ref , shape (P, G, L) : log-probabilities of the sampled tokens under the policy being trained, the policy that generated the samples, and a frozen reference policy;
  • mask , shape (P, G, L) : 1 for completion tokens and 0 for padding.

Clarifying Questions Guidance

  • Should the group standard deviation be the population or the sample estimate, and how should a group whose rewards are all equal be handled?
  • Is the loss averaged over each completion's tokens and then over completions, or over all tokens in the batch at once?
  • Which KL estimator should the penalty use, and what are the clip range and the KL coefficient?

Part 1 — Group-relative advantages and the GRPO loss

Compute each completion's group-relative advantage from the outcome rewards, assign it to every token of that completion, and compute the clipped GRPO objective with a KL penalty toward the reference policy, returned as a scalar loss. Vectorize: no Python loops over prompts, completions or tokens.

What This Part Should Cover Guidance

  • Group normalization along the correct axis, including the zero-variance case.
  • The probability ratio, clipping and minimum in the surrogate, the KL term, and correct masking.
  • Aggregation over tokens and completions, and the sign of the returned loss.

Part 2 — Process rewards

A process reward model now scores intermediate reasoning steps. Completion i has a list of step rewards and, for each step, the index of the token where that step ends. For one prompt's group, turn these into per-token advantages the way GRPO does under process supervision.

Clarifying Questions for this Part Guidance

  • Are step rewards normalized within each completion or across all steps in the group?
  • Should the completion's outcome reward also contribute to the token advantages?

What This Part Should Cover Guidance

  • Normalization of the step rewards and the rule that maps them to per-token advantages.
  • A vectorized implementation for completions with different numbers of steps.
  • How process rewards change credit assignment compared with a single outcome reward, and how they can be gamed.

Part 3 — Shaped rewards

Add shaping terms to the reward, for example a bonus for following the required output format or a penalty tied to completion length, and show how they enter the advantage computation.

Clarifying Questions for this Part Guidance

  • Which shaping terms are defined, and are they given per completion or per token?
  • Should shaping be combined with the task reward before or after group normalization?

What This Part Should Cover Guidance

  • Where shaping enters relative to group normalization, and how its weight is set.
  • The interaction between shaping and normalization, including groups with identical task rewards.
  • Risks of shaping, such as reward hacking or a changed optimal policy, and how to detect them in training logs.

What a Strong Answer Covers Guidance

  • Correct axes, broadcasting and masking throughout, verified on small hand-checked arrays.
  • Numerical care: an epsilon in normalization, ratios computed from log-probability differences, and a non-negative KL estimate.
  • A clear explanation of why GRPO needs no value network and what the group baseline costs.
  • Reasoned choices for every ambiguous convention: the standard deviation, loss aggregation, and where shaping is added.

Follow-up Questions Guidance

  • Why can dividing by the group standard deviation bias training toward some prompts, and what changes if you drop it?
  • If every completion in a group receives the same reward, what does that group contribute to the gradient, and what would you do about it?
  • How would you choose the group size under a fixed sampling budget?
  • How does averaging the loss per completion treat long and short completions differently?
Loading comments...