Phase 8 · AI in the Real WorldModule 30~40 min read

Reinforcement Learning

Learning by trial and reward: an agent acts in an environment, collects rewards, and discovers a strategy — the paradigm behind game-playing AI and robotics.

What you'll learn

The third great family of learning (after supervised and unsupervised) is reinforcement learning (RL): learning by trial and reward. An agent takes actions, receives rewards, and gradually discovers a strategy. It's how AI mastered Go and Atari, how robots learn to walk — and part of how ChatGPT was aligned.

By the end of this module you'll be able to:

  • Define agent, environment, state, action, and reward
  • Explain what a policy and a value function are
  • Describe the exploration–exploitation tradeoff
  • Sketch how Q-learning and deep RL work

Agents, environments & rewards

RL has a simple loop: an agent observes the state of its environment, takes an action, and receives a reward plus a new state. Its only goal is to maximise total reward over time. Crucially, nobody gives it the right answers — it must discover good behaviour from the reward signal alone.

States, actions & policies

A policy is the agent's strategy: which action to take in each state. A value functionestimates how good each state is (the reward you can expect from it onward). Learning the values lets the agent derive a good policy — always move toward higher value. Watch an agent learn a gridworld: values spread out from the goal, then it follows them home:

An agent learning a gridworld
Value iteration, then acting
+1−1

reach +1, avoid −1 · reward −0.04 per step

1/11In reinforcement learning an agent acts in an environment for reward. Here: reach +1, avoid the −1 pit.
Values (green tint) spread from the +1 goal; arrows show the learned policy; then the agent follows it to the goal.

Explore vs exploit

A central dilemma: should the agent exploit the best action it knows, or explore a new one that might be even better? Too much exploiting and it gets stuck in a mediocre habit; too much exploring and it never cashes in. Good RL balances the two — often by acting randomly a small fraction of the time and greedily the rest.

Value functions & Q-learning

Q-learning learns the value of each state-action pair — the "Q value" — from experience, nudging its estimates toward the reward received plus the value of where it lands. Over many episodes those estimates converge, and acting greedily on them gives an excellent policy. No model of the environment required — just trial, error, and reward.

Deep RL in the world

When states are too many to tabulate (a screen of pixels, a robot's sensors), we replace the table with a neural network — deep reinforcement learning. This powered DeepMind's Atari agents and AlphaGo/AlphaZero, which taught themselves superhuman play. RL also drives robotics and recommendation systems, and — as RLHF (Module 27) — it's how large language models are tuned to human preferences. The same reward-driven loop, everywhere.

Recap & quick check

Key takeaways

  • RL: an agent takes actions in an environment to maximise cumulative reward, with no labelled answers.
  • A policy maps states to actions; a value function estimates how good a state (or action) is.
  • The exploration–exploitation tradeoff balances trying new actions against using known-good ones.
  • Q-learning learns state-action values from experience and acts greedily on them.
  • Deep RL uses neural networks for huge state spaces — powering AlphaGo, robotics, and RLHF for LLMs.

Quick check

1. In reinforcement learning, what does the agent try to maximise?

2. A 'policy' is…

3. The exploration–exploitation tradeoff is about…

4. How is RL connected to modern LLMs?

We've built powerful systems. The last two modules ask the essential questions: is it safe and fair, and how do we ship it? Next up: Module 31 — Ethics, Bias, Safety & Alignment.