Model customization

Reinforcement Learning

RL

A way of training in which a system learns by acting: it tries actions, receives a reward or a penalty for the results, and gradually adopts the behaviour that earns the most reward over time.

Agentchooses an actionEnvironmenta simulated warehouseReward+1 right chute · −5 collisionUpdatewhat scores well is keptTRIAL AND ERRORmillions of roundsIf the reward is badly defined, the system optimises the reward and misses the goal: reward design is the hard part.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/reinforcement-learning

In plain terms

A child learns to ride a bicycle without a manual. They try, wobble, fall, adjust, and after enough attempts the body knows what works. Nobody gave them the correct answer for each moment; they were given a goal and felt the consequences. Reinforcement learning trains software the same way: it is told what counts as success, left to try millions of times, usually in a simulation, and it keeps what scores well.

Why it matters

It suits problems that are sequences of decisions with a measurable outcome and a safe place to practise: games, robotics, routing, pricing, control of industrial equipment, and increasingly the training of language models to reason and to complete tasks. It can find strategies no person thought of. It is also demanding: it needs very many trials, so a simulator is almost always required, and the system optimises exactly what is rewarded, which is rarely exactly what was meant. A badly chosen reward produces behaviour that scores well and serves nobody.

Example

A parcel company trains the controller for its sorting robots in a simulated warehouse. The reward is one point for every parcel delivered to the right chute and minus five for a collision. In the first thousand runs the robots mostly collide. After two million runs they sort 14% more parcels an hour than the hand-written rules did. They had also learned to avoid heavy parcels, which earned no extra points, and the reward had to be corrected.

Most often confused with

RL vs. Supervised learning

RLLearns from the consequences of its own actions
Supervised learningLearns from examples labelled with the right answer

Supervised learning is shown the correct output for every input. Reinforcement learning is never shown a correct action; it receives a score, often long after the decisions that earned it, and has to work out which of them deserved the credit. Use supervised learning when you can supply right answers. Reinforcement learning fits when you can only say how good an outcome was.

Origin: The foundations of the modern field were laid by Andrew Barto and Richard Sutton from the early 1980s; they received the 2024 Turing Award for this work.

Under the hood

The setting is an agent interacting with an environment: at each step it observes a state, chooses an action under its policy, and receives a reward and a new state; the aim is to maximise cumulative reward. Central difficulties: the trade-off between exploring new actions and exploiting known good ones, assigning credit when rewards arrive late, and sample efficiency. Main families: value-based methods such as Q-learning, policy-gradient methods such as PPO, and model-based methods that learn or are given a model of the environment. Deep reinforcement learning uses neural networks for the policy; its best-known result is AlphaGo, which defeated the Go champion Lee Sedol in 2016. Practical issues: the gap between simulation and reality, reward hacking (also called specification gaming), and safe exploration. In language models it appears as RLHF and as training on tasks whose results can be checked automatically, an approach used in training reasoning models. “Agent” here is the textbook term and predates today's LLM agents.

Written by Mehmet Erkek · Last updated: