What you'll learn
As apps grow beyond single prompts — chatbots with memory, agents with tools, RAG over big corpora — the skill shifts from writing a prompt to curating everything in the model's window. That discipline is context engineering: deciding, for each turn, exactly what the model should see and what it shouldn't.
By the end of this module you'll be able to:
- Explain how context engineering extends prompt engineering
- Treat the context window as a limited budget to allocate deliberately
- Carry memory across turns without letting history explode
- Use compaction to keep long sessions inside the window
From prompt to context
Prompt engineering asks "what words do I write?" Context engineering asks "what should be in the window at all?" In a real application, the final context is assembled at runtime from many sources — a system prompt, tool definitions, retrieved documents, conversation history, and the user's message. Everything you learned still applies; now you're also the editor deciding what makes the cut.
The window as a budget
Recall Module 3: the context window is finite, and everything shares it. Context engineering treats those tokens as a budget to allocate. Each item earns its place by improving the answer more than the space (and distraction) it costs.
System prompt
Role, rules, output format
Tools & schemas
Definitions of what the model can call
Memory / history
Relevant earlier turns (often compacted)
Retrieved facts
Only the chunks this turn needs
User input
The current request
Room for output
Reserved space for the reply
Key idea
Memory across turns
A chat that remembers can't just keep appending every message — it would overflow the window (Module 3) and drift into "lost in the middle" territory (Module 20). Instead, apps engineer memory: keep recent turns verbatim, store older facts as a running summary, and pull in only what's relevant to the current message.
| Memory approach | What it keeps | Trade-off |
|---|---|---|
| Full history | Every message verbatim | Simple, but overflows and gets costly |
| Sliding window | The last N turns only | Cheap, but forgets earlier context |
| Running summary | A condensed recap + recent turns | Compact, but detail is lost to the summary |
| Retrieved memory | Only past turns relevant to now | Focused, but needs a retrieval step |
Compaction & summarization
Compaction is the key technique for long sessions: when the conversation grows large, replace a block of older turns with a concise summary that preserves the essentials (decisions made, facts established, the user's goal) and drops the chatter. The session continues with far fewer tokens and the important thread intact.
Tip
Context engineering in real apps
Assembling a good context each turn usually means:
- A stable system prompt with role, rules, and format (written once, reused every turn).
- Only the tools relevant to the task — every tool definition costs tokens and attention.
- Retrieved facts for this turn, not the whole knowledge base.
- Compacted memory — a summary of the old plus verbatim recent turns.
- Clear placement — key material at the start and end (Module 20).
Note
Recap & quick check
Key takeaways
- Context engineering curates everything in the model's window, not just the wording of one prompt.
- The window is a token budget; each item must earn its place by helping more than it costs in space and distraction.
- The core question: what is the smallest set of information that lets the model do this task well?
- Manage memory by blending a running summary of old turns with verbatim recent turns and retrieval for specifics.
- Compaction replaces old turns with a purpose-built summary so long sessions stay inside the window.
Quick check
1. How does context engineering extend prompt engineering?
2. What's the guiding question of context engineering?
3. Why can't a memory-enabled chat just append every message forever?
4. What is compaction?
You can now ground and curate context expertly. Time to leave the chat box entirely and drive models with code — where prompting powers real software. Next up: Phase 6, Module 22 — Prompting Through the API.