# What this piece explains
This brief summarizes how GRPO (Group Relative Policy Optimization) trains smaller language models using verifiable rewards. The goal is to show the training mechanics, a simple arithmetic example that makes the idea concrete, and the trade-offs you get when you rely on outcome checks rather than supervised step-by-step targets.
# GRPO in a paragraph
GRPO was introduced as an alternative to conventional PPO actor-critic setups to reduce memory and critic training requirements. Instead of training a separate value estimator (a critic) to provide a baseline, GRPO samples a group of responses for the same prompt and uses the group's mean and standard deviation of rewards to compute an advantage for each attempt. The optimizer then favors attempts with higher advantages and downweights lower ones.
# Why verifiable rewards matter here
Verifiable rewards are outcome checks that can be computed independently of another language model. For problems with a clearly checkable final answer — like arithmetic or other tasks where a programmatic verifier can confirm correctness — the verifier supplies the reward signal without providing a target sequence. That lets you train on question–answer pairs without supervised worked solutions.
# The arithmetic example (concrete)
Question: A shop has 6 boxes with 8 items in each box. It sells 6 items. How many items remain? Answer: 42.
A programmatic verifier can check that calculation (boxes * items_per_box - sold == 42) and give a binary reward (1 if correct, 0 if incorrect). If the model generates multiple attempts for the same prompt, those attempt-level rewards form the group used by GRPO.
Sample attempts and binary rewards:
- Attempt 1: 42 -> reward 1
- Attempt 2: 40 -> reward 0
- Attempt 3: 42 -> reward 1
- Attempt 4: 48 -> reward 0
- Scope of the signal: The outcome-based advantage applies to the whole sampled response, not to individual reasoning steps inside the response. If the final number is correct, the model is rewarded even if the internal reasoning is flawed.
- No learned critic needed: GRPO's group baseline avoids training a separate value network, reducing memory and engineering complexity relative to actor-critic PPO pipelines.
- Dependence on group variation: If every sampled attempt gets the same reward, the group standard deviation is zero and GRPO produces no ranking signal for that group.
# Practical trade-offs and cautions
- Reward design matters. A verifier that checks only final outcomes will not penalize incorrect or spurious reasoning that happens to produce the right answer occasionally. That can encourage brittle or shortcut behavior.
- Outcome checks are best where independent verification is straightforward (math, deterministic transforms, verifiable tool outputs). For open-ended reasoning, verifiers are harder to design and may require different architectures or human evaluation.
- The GRPO update provides a relative ranking among sampled attempts, so the quality of sampled attempts and the diversity within a group affect learning dynamics.
# Bottom line
GRPO offers a lower-memory, outcome-driven way to train small reasoning models when you can build reliable verifiers. It simplifies infrastructure by using group-level baselines, but the training outcome depends heavily on how you define and compute the reward.