Back to resources

RLHF vs RLVR vs RLCD: How AI Models Learn What a "Good Answer" Is

RLHF vs RLVR vs RLCD: How AI Models Learn What a "Good Answer" Is

Ask an AI model to explain recursion and you might get two answers. Both are correct. One is clear and the other is impossible to follow.

You know right away which one is better. The model doesn’t, unless someone teaches it.

That problem sits at the center of how modern AI assistants are trained. The answer isn’t just “more data.” It’s better feedback: someone or something has to tell the model what “good” looks like.

Three techniques do that in very different ways: RLHF, RLVR and RLCD. Here’s how each one works, where it breaks down, and why the difference matters.

First, a quick primer on reinforcement learning

A model tries an answer and gets a score, called a reward. Over many attempts, it’s updated so that high-scoring answers become more likely.

Prompt → Model → Response → Reward → Update → Better responses

The main difference between RLHF, RLVR and RLCD is where the reward comes from.

1. RLHF: Reinforcement Learning from Human Feedback

People compare model outputs and pick the better one.

Example: two correct explanations of recursion, one dense and one beginner-friendly. Annotators prefer the clearer one.

How it works:

  1. Fine-tune a base model on example answers (supervised fine-tuning)
  2. Have humans rank multiple responses to the same prompt
  3. Train a reward model to predict those preferences
  4. Optimize the language model against that reward (historically with PPO)

Keep in mind: PPO is one optimization algorithm. RLHF is the overall approach. The reward model is also only an approximation of human preferences, not an understanding of human values.

Limitations: labeling is expensive, annotators disagree, and the model can learn to exploit weaknesses in the reward model.

2. RLVR: Reinforcement Learning with Verifiable Rewards

A verifier, rather than a human, checks whether the output is correct.

Example: “Write a Python function that checks if a number is prime.” Run unit tests. If it passes, the reward is positive. If it fails, the reward is zero.

This works well for math (check the final answer), code (run the tests) and formal tasks (check the constraints). The reward is objective, cheap to compute and scalable.

The catch: a verifier can only reward what it can verify. If it checks only the final answer, flawed reasoning that happens to reach the right answer still gets full reward. RLVR improves verifiable outcomes. On its own, it doesn’t guarantee sound reasoning.

3. RLCD: Reinforcement Learning from Contrastive Distillation

The model creates its own preference data by generating contrasting answers.

Example principle: “Give clear and helpful answers.”

  • Positive prompt: produce a helpful, clear answer
  • Negative prompt: produce an unhelpful, confusing answer

The positive output is automatically labeled “preferred.” Many of these pairs train a preference model, which is then used for RL.

Compared with asking an AI judge to score two near-identical outputs, pushing the outputs apart on purpose tends to produce cleaner labels.

The catch: the label comes from how each answer was generated, not from anyone judging it. Sometimes the “positive” answer isn’t actually better. RLCD also still relies on humans to write the principles and prompts. It removes direct preference labeling, not people.

One task, three reward mechanisms

Task: “Write a SQL query for the top 5 customers by revenue.”

  • RLHF: humans judge readability and usefulness
  • RLVR: run the query on a test database and compare the result with the expected rows
  • RLCD: contrasting “clear SQL” and “messy SQL” prompts produce preference pairs

Now try “Explain how a database index works.” RLVR struggles here, because no simple check can tell whether an explanation is clear. These techniques aren’t interchangeable.

Common misconceptions

  • “PPO = RLHF.” No. PPO is one algorithm used inside RLHF.
  • “RLVR guarantees correct reasoning.” No. It rewards only what the verifier checks.
  • “RLCD needs no reward model.” No. It still trains a preference model.
  • “A reward means the model understands why an answer is good.” No. A reward is just a number to optimize.

RLHF vs RLVR vs RLCD compared side by side

The bottom line

  • RLHF asks humans which answer they prefer.
  • RLVR asks a verifier whether the answer meets a checkable objective.
  • RLCD builds preference data by contrasting answers generated under opposing instructions.

Modern pipelines often combine these methods. Knowing where a model’s reward comes from tells you a lot about what that model has actually been optimized to do.

Book a CallMore resources