🎯 · Models

GRPO

Reinforcement learning that scores a group of answers against their own average to sharpen reasoning.

In one line

GRPO fine-tunes a model by sampling several answers per prompt and reinforcing the ones that beat the group's average on a checkable task.

ConceptWhat it is

Group Relative Policy Optimization, GRPO, is a reinforcement-learning fine-tuning method that improves a model's reasoning by optimizing its policy against rewards, but without the separate value model (the critic) that PPO requires. Introduced in DeepSeek's work on math and reasoning models, it replaces the learned baseline with a group-relative one: for each prompt it samples a group of completions and judges each against that group's own average reward.

It exists because the critic in classic RLHF-style PPO is expensive to train and awkward to stabilize. By deriving the advantage signal directly from a batch of sampled answers, GRPO removes that whole component while keeping the familiar clipped policy-update machinery. It pairs naturally with verifiable rewards — a math answer you can check, or code that either passes its tests or does not — which is why it became a workhorse for eliciting long chain-of-thought reasoning.

How it worksThe mechanics

For a given prompt the current policy samples a group of completions; a reward function scores each one, typically a rule-based verifier that checks whether the final answer is correct, the code passes, or the output meets a required format. The group's mean reward serves as the baseline, and each completion's advantage is its reward minus that mean, normalized by the group's spread, so answers that beat their peers are reinforced and worse-than-average ones are suppressed. Those advantages weight a clipped, PPO-style gradient update, with a KL penalty pulling the policy back toward a frozen reference model to prevent drift. The loop repeats over many prompts until reasoning quality plateaus.

At a glanceSee it

GRPO diagram
GRPO diagram 1

GRPO's defining move — drop PPO's learned critic and use the group mean as the baseline, paid for by sampling more completions per prompt.

GRPO diagram 2

Where GRPO's rewards come from — rule-based verifiers or a learned reward model — and the zero-variance trap when every completion in a group scores the same.

When to use itWhere it fits

  • Reasoning tasks with automatically checkable answers such as math, competitive programming, or formal logic.
  • When you want RL-driven capability gains but cannot afford or stabilize a separate PPO critic.
  • Eliciting or lengthening chain-of-thought behavior on an already-competent base model.
  • Domains where you can write a cheap, reliable reward function instead of collecting human labels.

When NOT to use itLimits & anti-patterns

  • Open-ended or subjective tasks where correctness cannot be programmatically verified.
  • Teaching the model new facts or knowledge, since RL shapes behavior rather than adding information.
  • Small budgets or tight timelines, because generating many completions per prompt is compute-heavy.
  • When a simpler method like SFT or DPO already produces the behavior you need.

Trade-offsAdvantages & costs

Advantages
  • Drops the value model, cutting memory and removing a fragile, hard-to-tune part of PPO.
  • Strong, well-demonstrated gains on math and reasoning benchmarks.
  • Works with rule-based verifiable rewards, avoiding costly human preference labeling.
  • Reuses the mature clipped-objective and KL-penalty machinery teams already know from PPO.
Trade-offs & costs
  • Reward design is the hard part; a sloppy reward invites reward hacking and gamed shortcuts.
  • Sampling a full group of completions per prompt makes every training step compute-intensive.
  • Limited to domains with verifiable or otherwise trustworthy reward signals.
  • Still full RL, so runs can be unstable, hyperparameters are sensitive, and a retrain is needed to update behavior.

ExampleIn the real world

DeepSeek's R1 line of reasoning models is the best-known worked case. Starting from a capable base model, the team applied GRPO with simple rule-based rewards, one signal for whether the final math or code answer was correct and another for keeping the reasoning inside the required format. The model learned on its own to produce long, self-checking chains of thought, and accuracy on hard math benchmarks climbed over the course of training without any human-labeled reasoning traces.

ToolsHow to implement it

  • Hugging Face TRL GRPOTrainerthe standard open-source implementation of the GRPO loss.
  • verla scalable reinforcement-learning-for-LLMs library that supports GRPO-style training at scale.
  • OpenRLHFdistributed RLHF toolkit that offers GRPO alongside PPO.
  • Unslothmemory-efficient recipes that make GRPO runnable on modest single-GPU setups.

Cost & effortWhat it takes

Compute is high: every update samples a whole group of completions per prompt, so rollout generation dominates the bill, though dropping the critic saves meaningful memory versus PPO. The heaviest human effort goes into designing and hardening the reward function and its verifiers rather than labeling data. It modifies a single model in place, and any behavior change requires re-running the RL loop rather than a quick patch.

A living map of modern AI — kept current every morning