RunLocalModel.com

LLM Alignment Explained: RLHF, DPO, GRPO, and RLAIF

By the RunLocalModel editorial team · Published September 30, 2026 · ~8 minute read

If you only read one paragraph Fine-tuning teaches a model to answer. Alignment teaches it which answer a person would rather read. The methods share that goal. They differ in who gives the grade. RLHF: people rank answers, a second model learns to imitate them, then PPO nudges the chatbot. Heavy. DPO: skip the second model and train straight on the ranked pairs. Cheaper, and what a lot of open models use. GRPO: write several answers to the same question and score them against each other. That is what DeepSeek-R1 uses. RLAIF: let an AI do the ranking, so you pay fewer people.
Quick answers
What problem is this solving?
Fine-tuning copies examples. Alignment chooses between two fluent replies: which one a person, a checker, or a judge would rather see.
What is RLHF + PPO?
People rank answers. A grader model learns the ranking. PPO then nudges the chatbot toward higher scores, with a leash so it does not wander off and cheat the grader.
Why did DPO take over so many open recipes?
It uses the ranked pairs directly. There is no second model to train and no online loop to keep from falling over.
Where does GRPO fit?
It compares answers inside one group, so you can skip a separate judge. DeepSeek-R1 uses it, often with a score a program can check.

This is the part of the teaching that decides whether you can rely on the model. The whole pipeline is in How LLMs Are Trained.

A model that answers can still answer badly

Fine-tuning teaches the shape of a reply. The examples are finite. Outside them, the model can be wordy, too eager to agree, casually unsafe, or confident about the weaker of two plausible answers. Preference data records the judgment: for this question, answer A is better than answer B. Every method below is a way to make A more likely than B. The engineering difference is who holds the red pen, and how many extra models you have to train.

RLHF + PPO: people, then a stand-in, then a nudge

Reinforcement learning from human feedback is the pipeline that made instruction-following assistants mainstream. InstructGPT wrote it down in three steps, building on earlier preference research and on PPO.

  1. People rank. For a question, the model writes several replies. People sort them.
  2. Train a stand-in grader. A second network learns to score a reply the way those people did. It is a compressed rater.
  3. Nudge the chatbot with PPO. The chatbot writes new replies, the stand-in scores them, and PPO updates the chatbot so the score goes up. A leash — the KL penalty, if you want the jargon — keeps it from drifting so far from the fine-tuned model that it finds a cheat the score likes and a person hates.

This works, and it is heavy. You pay people to rank, you train and host a second model, and you run a loop that keeps sampling while the chatbot is changing. The stand-in can be fooled. If the chatbot finds a quirk that scores high and reads badly, PPO will reinforce the quirk. The leash is there to limit that. Labs still use versions of this loop when they want the model to try replies that were not in the original ranked set.

DPO: the ranking goes straight into the model

Direct Preference Optimization starts from a math fact. The stand-in grader in the RLHF recipe can be written in terms of the chatbot itself. Once you do that, the update on a triple — question, preferred answer, rejected answer — is a straightforward training loss. You never build the grader. You never run PPO.

The data looks the same as RLHF’s ranking data. Someone, or something, has already said which reply is better. Training happens offline. The model does not have to write fresh answers during the update, which is why DPO runs are cheaper and easier to keep stable. That is why it became the default in a lot of open recipes. It still needs good pairs. A sloppy ranking teaches a sloppy preference. And because it never tries new replies while training, it can only reinforce differences that are already in the dataset. When you want the model to discover a strategy nobody wrote down, you want an online method: PPO, or GRPO below.

GRPO: score the group, skip the extra judge

Group Relative Policy Optimization, introduced in DeepSeekMath and used for DeepSeek-R1, is an online method with a smaller cast. For each question the model writes a group of answers. Each answer gets a score. What matters is how that score compares with the rest of the group, usually how far it sits from the group average. Answers above the group get a push up. Answers below it get a push down. There is no separate judge network learning what “average” looks like, and that judge is a large part of what makes PPO expensive.

The score can be a person’s preference. In the reasoning models that made GRPO famous, it is often a check a program can run: the final number matches, the unit tests pass, the reply used the expected format. Comparing inside the group is what lets them skip a learned grader. The checker is what tells the group who deserved to win. How that checker teaches a model to think before it answers is the subject of Reinforcement Fine-Tuning.

RLAIF: an AI writes the rankings

Human rankings dominate the budget of RLHF and of DPO. Reinforcement learning from AI feedback replaces most of them with a model judge. Anthropic’s Constitutional AI does this with a written constitution: a list of principles, a critique of a draft, and a revision. The RLAIF paper compared AI-written preferences with human-written ones and found the AI labels could carry a large share of the signal.

People still write the principles and still spot-check samples. The saving is in the volume of pairwise judgments. The risk moves with the judge. Whatever the judge likes too much — a tone, a length, a kind of hedge — becomes what the student model learns. A good setup treats the judge as something you test, not as a free source of truth.

Side by side

Method Who gives the grade Extra model Tries new answers while training
RLHF + PPO People, distilled into a grader model Yes, the grader, and usually a critic Yes
DPO Whoever ranked the pairs, applied directly No No
GRPO A score over a group of answers No separate judge Yes, a group per question
RLAIF An AI judge, often under a written list of rules The judge Depends on whether those labels then feed PPO, DPO, or something else

RLAIF is a way to get labels. DPO, PPO, and GRPO are ways to use a score. A lab can mix them: AI rankings plugged into DPO, or a checkable test plugged into GRPO.

What you can feel, and what you can run

On a model you download, this stage shows up as a steadier tone, a more consistent length, and refusals on some requests. The instruct tag usually means fine-tuning plus some preference stage. The model card is where you see whether that stage was DPO, PPO, or a house recipe.

If you are teaching a local model yourself, DPO on a few thousand carefully ranked pairs is the method that fits one GPU, especially combined with the add-ons in the fine-tuning guide. A full PPO stack — grader model, online sampling, the leash — is a lab system. GRPO is approachable when you already have an automatic checker, which is the setup reasoning-model training uses, and it is still a serious training run rather than an afternoon project.

Related guides on this site