In plain terms
A child learns to ride a bicycle without a manual. They try, wobble, fall, adjust, and after enough attempts the body knows what works. Nobody gave them the correct answer for each moment; they were given a goal and felt the consequences. Reinforcement learning trains software the same way: it is told what counts as success, left to try millions of times, usually in a simulation, and it keeps what scores well.
Why it matters
It suits problems that are sequences of decisions with a measurable outcome and a safe place to practise: games, robotics, routing, pricing, control of industrial equipment, and increasingly the training of language models to reason and to complete tasks. It can find strategies no person thought of. It is also demanding: it needs very many trials, so a simulator is almost always required, and the system optimises exactly what is rewarded, which is rarely exactly what was meant. A badly chosen reward produces behaviour that scores well and serves nobody.
Example
A parcel company trains the controller for its sorting robots in a simulated warehouse. The reward is one point for every parcel delivered to the right chute and minus five for a collision. In the first thousand runs the robots mostly collide. After two million runs they sort 14% more parcels an hour than the hand-written rules did. They had also learned to avoid heavy parcels, which earned no extra points, and the reward had to be corrected.
Most often confused with
RL vs. Supervised learning
Supervised learning is shown the correct output for every input. Reinforcement learning is never shown a correct action; it receives a score, often long after the decisions that earned it, and has to work out which of them deserved the credit. Use supervised learning when you can supply right answers. Reinforcement learning fits when you can only say how good an outcome was.
Origin: The foundations of the modern field were laid by Andrew Barto and Richard Sutton from the early 1980s; they received the 2024 Turing Award for this work.
Under the hood
The setting is an agent interacting with an environment: at each step it observes a state, chooses an action under its policy, and receives a reward and a new state; the aim is to maximise cumulative reward. Central difficulties: the trade-off between exploring new actions and exploiting known good ones, assigning credit when rewards arrive late, and sample efficiency. Main families: value-based methods such as Q-learning, policy-gradient methods such as PPO, and model-based methods that learn or are given a model of the environment. Deep reinforcement learning uses neural networks for the policy; its best-known result is AlphaGo, which defeated the Go champion Lee Sedol in 2016. Practical issues: the gap between simulation and reality, reward hacking (also called specification gaming), and safe exploration. In language models it appears as RLHF and as training on tasks whose results can be checked automatically, an approach used in training reasoning models. “Agent” here is the textbook term and predates today's LLM agents.