Phase 5 · Grounding & KnowledgeModule 20~32 min read

Working with Long Context & Documents

Modern models can read whole books at once — but bigger isn't automatically better. Learn placement, the 'lost in the middle' problem, and summarization strategies.

What you'll learn

Modern models can read a whole book in a single prompt — hundreds of thousands of tokens. That's liberating, but it comes with a catch: bigger context is not automatically better. Where you place information, and how much you include, changes how well the model actually uses it.

By the end of this module you'll be able to:

  • Explain why simply stuffing the window full can backfire
  • Describe the "lost in the middle" effect and design around it
  • Choose between summarizing, stuffing, and retrieving
  • Decide when long context beats RAG — and when it doesn't

Bigger isn't automatically better

A huge context window is a capability, not a strategy. Filling it with everything you have introduces three costs: more tokens (higher latency and price, per Module 3), more distraction (irrelevant text competes for the model's attention), and a greater chance the key fact gets buried. Include what helps; leave out what doesn't.

Lost in the middle

Research and practice both show a consistent pattern: models attend most strongly to the beginningand end of a long prompt, and can overlook facts stranded in the middle. A critical instruction buried halfway through a 100-page document is the easiest thing for the model to miss.

The “lost in the middle” effect
Start
Middle
End

Recall of a fact by its position in a long prompt — high at the ends, lowest in the middle.

A fact's recall depends on where it sits in a long prompt — strongest at the start and end, weakest in the middle.

Where to put what matters

Use that curve deliberately:

  • Put the instruction at the very start, and restate it at the very end of a long document.
  • Lead with the most important source rather than burying it among less relevant material.
  • Keep the question close to the content it's about, not separated by pages of filler.
  • Cut ruthlessly. The best defense against "lost in the middle" is a smaller middle.

Key idea

Placement is a prompt-engineering decision. Even with a giant window, treat the start and end as prime real estate and put your instructions and most important facts there.

Summarize, stuff, or retrieve

Three ways to fit a large body of text into a useful prompt:

StrategyHowBest when
StuffPut the whole thing in the context windowThe document is modest and you need every detail
SummarizeCondense first, then work from the summaryThe full text is too big or mostly irrelevant
Retrieve (RAG)Pull only the relevant chunks (Module 19)A large or changing corpus; you need citations
These combine: summarize each document, then retrieve the summaries; or retrieve chunks, then stuff them.

Long context vs. RAG

With big windows, why not skip retrieval and paste everything? Sometimes you should — for a single moderate document, long-context "stuffing" is simpler and keeps full detail. But for large or frequently changing corpora, RAG is usually cheaper (you send fewer tokens), more focused (less distraction), and easier to cite. Many production systems use both: retrieve the relevant documents, then let a long-context model read them in full.

Tip

Default heuristic: if it fits comfortably and it's all relevant, stuff it. If it's large, mostly irrelevant, or changes often, retrieve. Measure — don't assume the biggest prompt is the best one.

Recap & quick check

Key takeaways

  • A large context window is a capability, not a strategy — stuffing it full adds cost, distraction, and burial risk.
  • 'Lost in the middle': models attend most to the start and end of a long prompt and can miss facts in the middle.
  • Place instructions and key facts at the start and end, keep the question near its content, and cut filler.
  • Fit big text via stuffing, summarizing, or retrieving — and combine these as needed.
  • Stuff a single relevant document; retrieve for large, mostly-irrelevant, or changing corpora — often use both.

Quick check

1. Why can filling a huge context window backfire?

2. What is the 'lost in the middle' effect?

3. Given 'lost in the middle', where should a key instruction go in a long document prompt?

4. When is RAG usually preferable to stuffing everything into long context?

Managing what goes into the window — instructions, retrieved facts, history — is becoming a discipline of its own. Next up: Module 21 — Context Engineering.