Phase 6 · Tools, Agents & the APIModule 22~32 min read

Prompting Through the API

Leave the chat box behind. Send prompts programmatically: the messages format, system prompts, key parameters, and reading a response in Python.

What you'll learn

Everything so far works in a chat box — but real products call models programmatically, through an API. That's where prompting becomes software: you control the full message list, the parameters, and what happens to the response. This module maps the chat concepts you know onto the API you'll build with.

By the end of this module you'll be able to:

  • Explain what the API gives you that a chat box doesn't
  • Construct the messages array with system, user, and assistant roles
  • Set the key request parameters that shape a response
  • Read a response and use streaming and stop sequences

Chat UI vs. the API

A chat interface hides most of the machinery: it supplies a system prompt for you, remembers the conversation, and picks default parameters. The API hands you all of it. That's more responsibility — you manage the history and settings yourself — but it's the only way to embed a model in an app, run a prompt a million times, or build the tools and agents in the rest of this phase.

The messages array

Instead of one text box, the API takes a list of messages, each tagged with a role (Module 4). A system prompt sets standing behavior; user and assistant messages alternate to form the conversation. To give the model memory, you resend the relevant history each call — the API itself is stateless.

api_call.py
# Prompting through an API (generic, provider-neutral shape).
response = client.messages.create(
    model="your-model",
    system="You are a concise assistant. Answer in one sentence.",
    messages=[
        {"role": "user",      "content": "What is a context window?"},
        {"role": "assistant", "content": "It's the max tokens a model can consider at once."},
        {"role": "user",      "content": "And what happens if I exceed it?"},
    ],
    temperature=0.3,
    max_tokens=200,
)

print(response.output_text)
A system prompt plus an alternating user/assistant history. Note the assistant turn is included so the model 'remembers' the exchange.

Key idea

The API is stateless: it doesn't remember previous calls. Any memory your app has is memory you send in the messages each time — which is exactly why context engineering (Module 21) matters here.

Key request parameters

A handful of parameters control most behavior. The essentials:

ParameterWhat it does
modelWhich model answers — trades capability against speed and cost
systemThe system prompt: role, rules, format (set once per request)
messagesThe conversation history the model reasons over
temperatureRandomness of sampling (Module 16) — low for consistent, high for creative
max_tokensCaps the length of the reply (and its cost)
stopSequences that end generation early when produced
Every prompting technique from earlier phases ultimately becomes values in a call like this.

Reading the response, streaming & stop

The response object carries the generated text plus useful metadata — notably a token usagecount (your cost, per Module 3) and a stop reason telling you why generation ended.

  • Streaming delivers the reply token by token as it's generated, so a UI can show text appearing live instead of waiting for the whole response.
  • Stop sequences end generation the moment a chosen string appears — handy for structured output or to prevent the model from running past the part you want.
  • Check the stop reason. If it's "max tokens," the answer was cut off — raise the cap or ask for something shorter, rather than trusting a truncated reply.

Watch out

A truncated response (hit max_tokens) can look complete but be missing its ending — including a closing brace that breaks your JSON parser. Always check the stop reason before trusting output in a pipeline (Module 17).

Recap & quick check

Key takeaways

  • The API exposes what a chat box hides: the full message list, parameters, and control over the response.
  • You send a list of role-tagged messages — a system prompt plus alternating user/assistant turns.
  • The API is stateless: any memory is history you resend each call, so context engineering applies directly.
  • Key parameters: model, system, messages, temperature, max_tokens, stop — the home of every earlier technique.
  • Use streaming for live UIs and stop sequences to bound output; always check the stop reason for truncation.

Quick check

1. What does the API give you that a chat box doesn't?

2. The API is 'stateless.' What does that mean for memory?

3. Which parameter caps how long the reply can be?

4. Why check the response's stop reason?

The API lets you send prompts from code. The next step turns the model from a talker into a doer — by giving it tools it can call. Next up: Module 23 — Tools & Function Calling.