Phase 5 · Grounding & KnowledgeModule 21~32 min read

Context Engineering

Prompting grows up: deliberately managing everything in the model's window — instructions, memory, retrieved facts, and tools — as a limited, valuable resource.

What you'll learn

As apps grow beyond single prompts — chatbots with memory, agents with tools, RAG over big corpora — the skill shifts from writing a prompt to curating everything in the model's window. That discipline is context engineering: deciding, for each turn, exactly what the model should see and what it shouldn't.

By the end of this module you'll be able to:

  • Explain how context engineering extends prompt engineering
  • Treat the context window as a limited budget to allocate deliberately
  • Carry memory across turns without letting history explode
  • Use compaction to keep long sessions inside the window

From prompt to context

Prompt engineering asks "what words do I write?" Context engineering asks "what should be in the window at all?" In a real application, the final context is assembled at runtime from many sources — a system prompt, tool definitions, retrieved documents, conversation history, and the user's message. Everything you learned still applies; now you're also the editor deciding what makes the cut.

The window as a budget

Recall Module 3: the context window is finite, and everything shares it. Context engineering treats those tokens as a budget to allocate. Each item earns its place by improving the answer more than the space (and distraction) it costs.

Everything competes for one budget

System prompt

Role, rules, output format

Tools & schemas

Definitions of what the model can call

Memory / history

Relevant earlier turns (often compacted)

Retrieved facts

Only the chunks this turn needs

User input

The current request

Room for output

Reserved space for the reply

Each turn, you decide how much of the window goes to rules, tools, memory, retrieved facts, the input, and the reply.

Key idea

The guiding question of context engineering: "What is the smallest set of information that lets the model do this task well?" More context is not more help — the right context is.

Memory across turns

A chat that remembers can't just keep appending every message — it would overflow the window (Module 3) and drift into "lost in the middle" territory (Module 20). Instead, apps engineer memory: keep recent turns verbatim, store older facts as a running summary, and pull in only what's relevant to the current message.

Memory approachWhat it keepsTrade-off
Full historyEvery message verbatimSimple, but overflows and gets costly
Sliding windowThe last N turns onlyCheap, but forgets earlier context
Running summaryA condensed recap + recent turnsCompact, but detail is lost to the summary
Retrieved memoryOnly past turns relevant to nowFocused, but needs a retrieval step
Real assistants blend these: a summary of the old, verbatim recent turns, and retrieval for specifics.

Compaction & summarization

Compaction is the key technique for long sessions: when the conversation grows large, replace a block of older turns with a concise summary that preserves the essentials (decisions made, facts established, the user's goal) and drops the chatter. The session continues with far fewer tokens and the important thread intact.

Tip

Summarize for the purpose, not generically. "Summarize the key decisions and open questions so we can continue the task" keeps what the next turn needs; a vague summary can quietly drop the one detail that mattered.

Context engineering in real apps

Assembling a good context each turn usually means:

  • A stable system prompt with role, rules, and format (written once, reused every turn).
  • Only the tools relevant to the task — every tool definition costs tokens and attention.
  • Retrieved facts for this turn, not the whole knowledge base.
  • Compacted memory — a summary of the old plus verbatim recent turns.
  • Clear placement — key material at the start and end (Module 20).

Note

This is the foundation agents are built on. An agent runs many turns, each accumulating tool results and observations; without disciplined context engineering, its window fills with noise and quality collapses. We put it to work in Phase 6.

Recap & quick check

Key takeaways

  • Context engineering curates everything in the model's window, not just the wording of one prompt.
  • The window is a token budget; each item must earn its place by helping more than it costs in space and distraction.
  • The core question: what is the smallest set of information that lets the model do this task well?
  • Manage memory by blending a running summary of old turns with verbatim recent turns and retrieval for specifics.
  • Compaction replaces old turns with a purpose-built summary so long sessions stay inside the window.

Quick check

1. How does context engineering extend prompt engineering?

2. What's the guiding question of context engineering?

3. Why can't a memory-enabled chat just append every message forever?

4. What is compaction?

You can now ground and curate context expertly. Time to leave the chat box entirely and drive models with code — where prompting powers real software. Next up: Phase 6, Module 22 — Prompting Through the API.