Phase 7 · Generative AI & LLMsModule 27~38 min read

How LLMs Are Trained

The pipeline that turns raw text into a helpful assistant: massive pretraining, supervised fine-tuning, and reinforcement learning from human feedback.

What you'll learn

A model that predicts internet text isn't yet a helpful assistant — it will happily continue a scam email or ramble. Turning a raw model into something like ChatGPT or Claude takes a multi-stage pipeline. This module walks through it.

By the end of this module you'll be able to:

  • Explain pretraining and its scale
  • Describe supervised fine-tuning
  • Say what RLHF adds and why
  • Choose between fine-tuning, prompting, and RAG

Pretraining

The first and by far most expensive stage is pretraining: next-token prediction over a massive chunk of the internet, books, and code — trillions of tokens. From this alone the model absorbs grammar, facts, translation, and a surprising amount of reasoning. The result is a base model: knowledgeable, but untamed. Training a frontier model this way can cost millions of dollars and thousands of GPUs running for weeks.

The training pipeline

Three stages turn raw prediction into a helpful, safe assistant:

From internet text to helpful assistant

1 · Pretraining

Predict the next token on a huge slice of the internet. Learns language, facts, reasoning.

→ a raw 'base' model

2 · Supervised fine-tuning

Train on curated example dialogues written by humans.

→ follows instructions

3 · RLHF

Humans rank responses; the model is tuned to prefer the better ones.

→ helpful, safe assistant

Pretraining gives knowledge; fine-tuning teaches it to follow instructions; RLHF aligns it with human preferences.

Supervised fine-tuning

Next, supervised fine-tuning (SFT) trains the base model on high-quality example conversations — prompts paired with ideal responses, written by humans. This teaches the model the format of being a helpful assistant: answer the question, follow the instruction, use a helpful tone.

RLHF

Finally, reinforcement learning from human feedback (RLHF) polishes behaviour. Humans compare pairs of model responses and pick the better one; those preferences train a reward model, and the LLM is then optimised (with reinforcement learning, Module 30) to produce responses the reward model rates highly. This is where helpfulness, honesty, and safety are dialled in — the difference between a raw text engine and an assistant you can actually talk to.

Note

RLHF is why the same underlying model can feel wildly different across products: the base knowledge is similar, but the fine-tuning and preference data shape its personality, refusals, and style.

Fine-tune, prompt, or RAG?

You usually won't train your own LLM. To adapt an existing one to your needs, you have three tools, from cheapest to most involved:

  • Prompting: just ask well (Module 28). Instant, no training.
  • RAG: retrieve relevant documents and put them in the prompt — grounds the model in your data.
  • Fine-tuning: further-train the model on your examples — for a consistent style or specialised task.

Recap & quick check

Key takeaways

  • Pretraining does next-token prediction on trillions of tokens, producing a knowledgeable but untamed base model.
  • Supervised fine-tuning teaches the model to follow instructions using human-written example dialogues.
  • RLHF uses human preference rankings to make the model helpful, honest, and safe.
  • Pretraining is hugely expensive; fine-tuning and alignment are comparatively cheap.
  • To adapt an LLM: prompt (cheapest), add RAG for your data, or fine-tune for a specialised style/task.

Quick check

1. What happens during pretraining?

2. What does RLHF primarily add?

3. Which is the cheapest way to adapt an existing LLM to your task?

4. A 'base model' straight out of pretraining is…

Most of us interact with LLMs by using them, not training them. Next: how to get great results from a model you didn't build. Next up: Module 28 — Using LLMs Well.