Reinforcement learning teaches an agent to act by rewarding good outcomes and penalizing bad ones over time.
ConceptWhat it is
Reinforcement learning, RL, trains an agent to choose actions in an environment by trial and error, receiving a reward signal that reinforces good decisions and discourages bad ones, rather than learning from a fixed labeled dataset.
It exists for problems that are fundamentally sequential and interactive, like game-playing, robotics, or tuning a language model's behavior, where the right action depends on the current state and future consequences, not a single static label.
How it worksThe mechanics
An agent observes the current state of an environment, takes an action, receives a reward and a new state, and updates its policy, its strategy for choosing actions, to increase the reward it accumulates over many episodes of trial and error.
At a glanceSee it
The RL algorithm landscape as a decision tree — first whether you can model the environment, then whether to learn action values, the policy directly, or both at once in actor-critic.
How RL aligns a language model — RLHF turns human preference rankings into a training signal, forking between a PPO-plus-reward-model path and the simpler direct-preference path.
When to use itWhere it fits
- Sequential decision problems where actions affect future states, like robotics or game AI.
- Optimizing a policy over time with a clear reward signal, like ad bidding.
- Fine-tuning language models on human preference signals, as in RLHF.
- Environments where simulation makes trial and error cheap and safe.
When NOT to use itLimits & anti-patterns
- Problems with a fixed labeled dataset and no sequential decision structure, where supervised learning is simpler.
- Real-world settings where trial and error is costly, dangerous, or slow.
- Situations lacking a clear, well-shaped reward signal to learn from.
Trade-offsAdvantages & costs
Advantages
- Handles sequential, long-horizon decision problems directly.
- Can discover strategies humans would not have designed.
- Learns purely from interaction, no labeled dataset required.
- Directly optimizes for a defined objective over time.
Trade-offs & costs
- Often needs huge numbers of trials to learn well.
- Reward design is hard, and a poorly shaped reward causes bad behavior.
- Training can be unstable and hard to reproduce.
- Risky or expensive to train in the real world versus simulation.
ExampleIn the real world
DeepMind's AlphaGo and its successor AlphaZero used reinforcement learning through millions of self-played games to reach superhuman performance at Go and chess.
ToolsHow to implement it
- Stable-Baselines3widely used library of standard RL algorithms.
- Gymnasiumthe standard environment interface for RL research.
- Ray RLlibscalable RL training for production and research.
- TRL by Hugging Faceapplies RL techniques to fine-tune language models.
Cost & effortWhat it takes
Training cost can be very high due to the sheer number of trial-and-error episodes needed; simulation environments reduce real-world risk and cost; engineering effort for reward design is substantial.
What changedWhat changed here
Updated this page A study of reinforcement learning agents inferring hidden rules from trial-and-error feedback, with analysis of transfer and generalization.
Three kinds of claim, strongest first. Signal runs every morning.