What you'll learn
How does AI turn the words "an astronaut riding a horse" into a photorealistic image? Two ideas made it possible: the adversarial game of GANs, and the denoising process of diffusionmodels — the technology behind Stable Diffusion, DALL·E, and Midjourney.
By the end of this module you'll be able to:
- Explain why generating images is hard
- Describe how a GAN's two networks compete
- Understand the diffusion denoising process
- Say how text-to-image guidance works
The image generation problem
A small 256×256 colour image has nearly 200,000 numbers, and only a vanishingly tiny fraction of all possible pixel combinations look like a real photo. Generation means learning that razor-thin distribution of "realistic images" well enough to produce brand-new points on it. Two approaches cracked it.
GANs: a two-player game
A Generative Adversarial Network pits two networks against each other. The generator creates fake images; the discriminator tries to tell fakes from real. As each improves, so must the other — an arms race that drives the generator toward strikingly realistic output:
Generator
Makes fake images from noise, trying to fool the judge. Gets better at faking.
compete
⇄
Discriminator
Judges real vs fake. Gets better at spotting fakes — pushing the generator to improve.
GANs produce sharp images but are notoriously tricky to train (the two networks must stay balanced). Around 2021, diffusion models largely overtook them.
Diffusion: learning to denoise
Diffusion models take a wonderfully different route. In training they repeatedly add noise to real images until nothing remains, and learn to reverse that — to remove a little noise at a time. To generate, they start from pure static and denoise, step by step, until an image emerges. Watch it happen:
step 0: pure random noise
Key idea
Text-to-image & guidance
To follow a prompt, the denoiser is conditioned on text: the caption is encoded (using the same embedding and attention ideas from Phase 6) and steers every denoising step toward an image that matches the words. And rather than denoise millions of pixels directly, latent diffusion (Stable Diffusion) works in a compressed latent space — far cheaper, which is why image generation can run on a single GPU.
Recap & quick check
Key takeaways
- Realistic images are a tiny slice of all pixel combinations; generation must learn that distribution.
- A GAN pits a generator against a discriminator; their competition yields realistic images.
- Diffusion models learn to remove noise, then generate by denoising from pure static step by step.
- Text-to-image conditions the denoiser on an encoded prompt to steer the result.
- Latent diffusion denoises in a compressed space, making generation efficient.
Quick check
1. In a GAN, what does the discriminator do?
2. How does a diffusion model generate an image?
3. How does a diffusion model follow a text prompt?
4. What makes latent diffusion efficient?
We've covered how AI perceives and creates. The final phase looks at AI that acts — and at using all of this responsibly. Next up: Module 30 — Reinforcement Learning.