Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

RLVR and GRPO: train against a checkable outcome

By Anup Rai7 min readReviewed September 2026

Reinforcement learning with verifiable rewards (RLVR) updates a model using rewards computed by a checker, such as answer comparison, code tests, or a proof verifier. The model is the policy that produces actions or tokens. The checker scores observable outcomes under a defined task contract.

“Verifiable” describes the reward mechanism. It does not make the checker complete, immune to exploitation, or a proof that every reasoning step is sound. In an interview, explain the verifier before choosing an optimization algorithm.

RLVR: Reinforcement Learning with Verifiable Rewards

Suppose a code-generation model must implement small transformation functions. A training environment can run candidates against hidden tests in an isolated sandbox. The desired outcome is correct behavior on the specification, while the immediate reward measures passing those tests. The gap between those two is the central risk.

  1. Sample a task and its independent checking procedure.
  2. Generate candidate solutions from a recorded policy version.
  3. Score candidates within time, memory, and execution limits.
  4. Estimate which sampled actions performed better than a baseline.
  5. Update the policy using the selected RL objective.
  6. Evaluate on untouched tasks and inspect reward exploitation.
Architecture / visual model
flowchart LR T[Tasks with checking contract] --> P[Current policy rollouts] P --> V[Isolated verifier] V --> R[Rewards and advantages] R --> U[Policy update] U --> P U --> E[Independent evaluation and exploit checks]
Read diagram source
flowchart LR
  T[Tasks with checking contract] --> P[Current policy rollouts]
  P --> V[Isolated verifier]
  V --> R[Rewards and advantages]
  R --> U[Policy update]
  U --> P
  U --> E[Independent evaluation and exploit checks]

An outcome reward scores the final result. Process supervision scores intermediate steps. Long explanations, self-checking, or backtracking can emerge in some training settings, but are not guaranteed or required by the definition of RLVR. DeepSeek-R1 provides a documented open training report; do not attribute its exact algorithm to providers that have not published theirs.

GRPO

Group Relative Policy Optimization (GRPO) estimates a baseline from multiple completions for the same prompt rather than training a separate value model for that purpose. It was introduced in DeepSeekMath. This can save value-model training state, but still requires rollout generation, scoring, policy optimization, and any reference-model computations in the chosen recipe.

For a simple outcome-reward version, sample a group with rewards r₁ … r_G and use a normalized relative advantage:

Aᵢ = (rᵢ − mean(group_rewards)) / (std(group_rewards) + ε)

For rewards [1, 1, 0, 0], population mean is 0.5 and standard deviation is 0.5; advantages are approximately [1, 1, -1, -1]. Implementations may differ in normalization. Explain the comparison rather than memorizing an unexplained formula.

The original clipped policy objective uses probability ratios for the sampled token at its history, comparing the updated policy with the rollout policy. A sequence log-probability is a sum across tokens; exponentiating that sum is not the same thing as a per-token ratio. This distinction matters when discussing sequence-level variants.

A KL penalty to a reference policy can constrain drift, depending on the recipe. The rollout policy used in importance ratios and the reference policy used for regularization play different roles and need not be identical throughout training.

Diagnose the group signal

Group outcome Relative reward signal Consequence
Mixed success and failure Nonzero advantages The policy can favor better sampled outcomes
All correct Zero advantages in this outcome-only group Little task-learning signal from the relative reward term
All wrong Also zero advantages More sampling or better task selection may be needed
Reward varies for irrelevant formatting Nonzero but misdirected signal The model can improve reward without solving the task

When all advantages are zero, the policy-gradient reward term contributes no signal. An enabled KL or other loss can still produce gradients. Saying “the entire training step has no gradient” would be too broad.

Objective variants and their costs

Approach Main change to understand Tradeoff to test
Dr.GRPO Examines and changes normalizations associated with length and difficulty biases Behavior depends on data and the full training recipe
DAPO Uses changes including asymmetric clipping, dynamic sampling, and token-level loss Filtering uninformative groups can consume extra rollouts
GSPO Uses sequence-level importance sampling and sequence-based optimization Compare stability and efficiency in the target setting

Primary reports: Dr.GRPO, DAPO, GSPO. A length bias is a property of the objective and its normalization, not simply the fact that a sequence shares an outcome reward. These methods are alternatives to evaluate, not a universal ranking or a required replacement order.

Reward Design and Failure Modes

Failure Example Repair and remaining cost
Incomplete verifier Code hard-codes visible examples Hidden and adversarial tests; specification coverage still matters
Parser exploit Answer extraction accepts an unrelated matching number Strict output parsing and adversarial verifier tests
Correct result, flawed explanation Invalid intermediate step happens to reach the answer Evaluate steps when they are part of the required deliverable
Format dominates correctness Tagged output earns reward despite a wrong result Rebalance or remove shaping rewards; monitor each component
Length becomes a shortcut Extra tokens correlate with reward or exploit normalization Inspect objective, truncation, and success at matched token budgets
Benchmark leakage Training includes near copies of evaluation problems Source-aware splits and contamination checks

A learned reward model has its own exploitation surface, but a deterministic checker also has blind spots. There is no general theorem that every RLVR setup is more stable than every RLHF setup.

Does RL Add Capability or Sharpen Sampling?

Pass@1 measures success with one sampled attempt under a specified setup. Pass@k asks whether at least one of k attempts succeeds. A better pass@1 can reflect a higher probability of already-possible correct outputs. Success found only after very many samples is different from reliable single-attempt performance.

To investigate a capability claim:

  1. Compare the base and trained model using the same prompts, decoding rules, and evaluation checker.
  2. Report sampling counts and generated-token or compute budgets; large k costs more.
  3. Measure important difficulty slices and use multiple seeds or uncertainty estimates.
  4. Check final-answer validity and, when needed, intermediate reasoning.
  5. Include other base families before generalizing a result.

Finite sampling cannot prove that a base model had zero probability of producing a solution. Conversely, improved sampling efficiency is useful even without a claim of entirely new capability. The Spurious Rewards study reports gains from uninformative rewards for some Qwen settings that did not generalize to other model families. A base-versus-RL study and ProRL report different capability-boundary findings under different training setups. These are reasons to inspect experimental conditions, not evidence for one universal conclusion. Keep the interview answer grounded in the measured capability claim rather than declaring this debate settled.

Distillation: The Cheaper Path

Method Typical source of trajectories Supervision Cost to examine
Teacher-response SFT Teacher Target response tokens Generation, filtering, and supervised training
RLVR Current policy Checked outcomes Rollouts, verifier execution, and policy updates
On-policy distillation Current student Teacher feedback on student states Teacher scoring plus training infrastructure

Distillation can transfer useful task behavior to a smaller model without developing an RL system. It is not always cheaper after teacher queries and refresh work, but it is a valuable baseline. Continued training on relevant reasoning data before RL may also improve the starting policy. Treat “mid-training” as a stage whose objective and data must be specified, not a magic capability switch. A controlled study of pretraining, mid-training, and RL investigates their interaction using synthetic reasoning tasks; its conclusions should be interpreted within that experimental scope.

Practical Guidance

Start with an existing model and a measured task baseline. Consider RLVR when the task has a credible verifier, expected gains justify the infrastructure, and the policy can generate some informative successes. Tasks that are always solved or never solved may provide little group-relative signal, though curriculum and exploration choices can change that.

For a batch of 256 prompts, eight completions each, and 2,000 generated tokens per completion, one rollout batch produces 4,096,000 tokens before training. This illustrative arithmetic explains why rollout length and sampling dominate many cost discussions. Add verifier execution, retries, rejected groups, and checkpointing; do not estimate training cost from parameter count alone.

Interview practice

  1. What does RLVR verify? Whatever the checker measures under its contract. Hidden-test coverage and parsing are part of the design.
  2. What does GRPO remove? A separate learned value baseline in the described recipe, not all auxiliary memory or rollout cost.
  3. Why can all-wrong groups be uninformative? Equal outcome rewards yield zero relative advantages; regularization can still act.
  4. Why distinguish token and sequence ratios? They define different importance-weighting objectives and can have different stability properties.
  5. Does pass@1 improvement prove new capability? No. It establishes better one-attempt performance in the evaluated setup; stronger claims require further evidence.
  6. When is distillation the better proposal? When teacher supervision can achieve the task target with lower total development and operating cost.
  7. How do you close a reasoning-training design? State the verifier contract, measured quality gain, exploit checks, rollout budget, and release/rollback criteria.

Recall card and closing

Checker → samples → relative signal → update → independent evidence. The strongest answer explains both how reward improves behavior and how a model could earn that reward for the wrong reason. Continue with inference fundamentals to budget the resulting serving workload.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Quantization: budget memory without guessing quality
NEXT LESSONInference: follow the request before optimizing it →

Explore the diagram