Skip to content
StrataHub

Algorithms · 2022

RLHF

The technique that made ChatGPT feel helpful learned less from data and more from human thumbs-up and thumbs-down.

A language model trained only to predict the next word is good at continuing text but not at being helpful, honest, or safe. Reinforcement Learning from Human Feedback, or RLHF, was the method that bridged that gap and, in 2022, helped turn a raw model into ChatGPT.

The idea traces to a 2017 paper by Paul Christiano and colleagues at OpenAI and DeepMind, who trained agents from human preferences rather than a hand-coded reward. People compared pairs of behaviors and picked the better one, and the system learned to chase whatever those choices implied.

Applied to language models in OpenAI's 2022 InstructGPT work, the recipe has three stages. First, humans write and rank example responses. Second, those rankings train a reward model that predicts which answers people prefer. Third, the language model is fine-tuned with reinforcement learning to score highly against that reward model.

The result was a model that follows instructions, refuses obviously harmful requests, and stays on topic far more reliably than its untuned ancestor, even though it is often no larger. The alignment came from human judgment, not extra scale.

RLHF has limits. It inherits the biases and blind spots of its raters, it can teach a model to sound agreeable rather than be correct, and it is expensive to run. Variants like reinforcement learning from AI feedback try to reduce the human bottleneck.

Still, RLHF was the hinge between impressive research demos and products people actually wanted to talk to. It made the difference between a model that knew a lot and one that would help.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.