You ship a chat feature. It works in testing. Then a user holds a long conversation, and the assistant forgets what they said at the start. Or the API returns an error about a prompt being too long. Or the monthly bill is higher than your estimate, and nobody can say why.
All 3 problems have one cause: the context window. It is not a setting on the model that you can ignore. It is a budget, and your backend is the one spending it.
This guide explains what a context window is, what counts against it, what happens when you go over, and the patterns that keep a real application inside it.
What a Context Window Is
A context window is all the text a language model can look at while it writes a response. Anthropic's documentation describes it as "all the text a language model can reference when generating a response, including the response itself", and compares it to working memory. It is separate from the model's training data, which is what the model learned before you ever called it.
The window is measured in tokens, not words. A token is a chunk of text, often a short word or part of a longer one. As a working rule, treat a token as roughly 3 to 4 characters of English.
Here is the part that surprises most backend developers: the model does not remember your last request. Each call starts from nothing. If you want the model to know what was said earlier, your code has to send it again. The context window is the size of the envelope you are allowed to put that history in.
How Big Are Context Windows Today
Sizes have grown fast, and they differ by model. As of October 8, 2026, the figures from the vendors' own documentation are:
- Anthropic lists a 1M-token window for its current flagship models, including Claude Opus 5.5 and Claude Sonnet 5.5, with up to 128k output tokens in a single request. Other Claude models have 200k tokens. Source: Claude context windows documentation, checked October 8, 2026.
- OpenAI lists 1.05M tokens of context and 128K max output tokens for its GPT-6 models, including GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna. Source: OpenAI models documentation, checked October 8, 2026.
- Ollama, which runs open-weight models on your own machine, defaults to 4k tokens under 24 GiB of GPU memory, 32k from 24 to 48 GiB, and 256k at 48 GiB or more. Source: Ollama context length documentation, checked October 8, 2026.
These numbers will change. Treat them as a snapshot and read the current page for the model you call. The lesson holds even when the numbers move: a window of 1M tokens is large, but it is not free, and it is not infinite.
What Counts Against the Window
Developers often count only the user's latest message. The window holds much more than that. Everything below competes for the same space:
- The system prompt. Your instructions, rules, and persona. They are sent on every request.
- The message history. Every earlier user and assistant message you include.
- Tool definitions and tool results. The description of each function the model can call, plus whatever those functions return.
- Retrieved documents. Passages you pull from a database or a vector store and paste into the prompt.
- Images and files. Documents, PDFs, and images also turn into tokens.
- The output. What the model writes back counts too, including any reasoning tokens when thinking is enabled.
Anthropic's documentation confirms the last point: output the model generates counts toward the window. So the real question is never "does my input fit?" It is "does my input plus the longest answer you will allow fit?"
What Happens When You Go Over
Providers handle the overflow differently, so read the docs for the one you call. On the Claude API, if the input alone is too long, you get a 400 invalid_request_error that says the prompt is too long. On newer Claude models, a request is accepted when input plus max_tokens exceeds the window, and generation simply stops with the stop reason model_context_window_exceeded when it reaches the limit.
The second case is the dangerous one, because nothing fails. The user just gets an answer that stops halfway. Your code should check the stop reason on every response and treat a context stop as a real event: log it, and decide whether to retry with less input.
There is also a quieter failure that never throws an error. Anthropic's documentation notes that accuracy and recall degrade as token count grows, a problem it calls context rot. A prompt can fit comfortably and still produce worse answers because it is stuffed with text the model does not need. More context is not the same as better context.
Why It Is a Budget, Not a Limit
Think about the window the way you think about a database connection pool or a request timeout. It is a finite resource, and your design decides how it is spent.
Every token in the window costs money, because providers charge per token for input and for output. Every token also adds latency, because the model has to process it before it answers. A conversation that grows by 500 tokens a turn does not just get slower and pricier on turn 40. It pays for the entire history again on every single turn, because you resend it each time.
For an engineer building from Lagos, Nairobi, or Accra, this is a budget decision as much as a technical one. Most hosted APIs are priced in dollars, and a feature that quietly doubles its token use doubles its cost against a local revenue base. Context discipline is the cheapest optimisation you will ever make.
Self-hosting changes the trade. With an open-weight model on your own hardware, there is no per-token charge, but your GPU memory sets the context size, as the Ollama defaults above show. You swap a dollar cost for a hardware limit. Neither is better. Know which one you are paying.
Patterns for Staying Inside the Window
Most production systems combine 3 patterns. Start with the simplest and add the others when you need them.
Trim the History
Keep the most recent messages and drop the oldest when the history exceeds a token budget. This is the cheapest fix and handles most chat features. The system prompt is kept separate, so it is never trimmed.
Create the file trim-history.js and add this function:
// trim-history.js
const estimateTokens = (text) => Math.ceil(text.length / 4);
export function trimHistory(messages, budget = 6000) {
const kept = [];
let used = 0;
// Walk from newest to oldest and stop when the budget is spent.
for (let i = messages.length - 1; i >= 0; i--) {
const cost = estimateTokens(messages[i].content) + 4;
if (used + cost > budget) break;
kept.unshift(messages[i]);
used += cost;
}
return kept;
}
The function loops from the newest message backward and adds messages until the budget is gone, then returns what fit in the original order. The + 4 covers the small overhead each message carries. The 4-characters-per-token estimate is a safety margin, not an exact count. For exact numbers, use your provider's token counting tool before you send.
Call it right before the model request, and reserve space for the answer by subtracting your max_tokens and your system prompt from the total window first.
Summarise What You Drop
Trimming loses facts. If a user said their name or a requirement 30 messages ago, it is gone. To keep the important parts, ask the model to summarise the older messages into a short paragraph, then send that summary in place of the raw history. You pay for one extra call, and you save tokens on every call after it.
Some providers now offer this on the server side. Anthropic describes server-side compaction, in beta for newer Claude models, as its main strategy for long-running conversations. It summarises earlier turns so the conversation can continue past the limit. Building your own summary step still works everywhere and gives you control over what is kept.
Retrieve Instead of Stuffing
Do not paste a whole knowledge base into the prompt. Store documents outside the model, find the few passages that match the question, and put only those in the window. This pattern is called retrieval-augmented generation (RAG). It keeps the prompt small, keeps cost flat as your data grows, and helps with context rot because the model sees less noise. Our guide to RAG, embeddings, and vector stores shows how to build the retrieval side.
Where the History Lives
Trimming and summarising only work if you stored the full history somewhere first. Keep every message in your database and decide at request time which ones go into the window. The database holds everything. The window holds what this request needs.
That separation is the core design idea. If you are designing the storage side, our chat application database design guide walks through a schema for messages and conversations. For the bigger picture of where context handling sits in an AI system, see the 6 layers every AI backend needs.
How to Debug a Context Problem
When an AI feature misbehaves, check the context before you blame the model. Log the token counts for system prompt, history, retrieved text, and output on every request. Log the stop reason. Most "the model forgot" bugs turn out to be a trimming rule that dropped the wrong message, and most "the model is wrong" bugs are a retrieval step that supplied the wrong passage. We go deeper on this in how to debug AI backend systems.
If you want to go from reading about these patterns to building them in a structured order, the AI Engineering course covers LLM API integration, prompting, and retrieval step by step, with the backend skills these patterns rest on.
Summary
A context window is the total text a model can see while it answers, and that includes your system prompt, the history, tool results, retrieved documents, and the answer itself. The model remembers nothing between requests, so your backend decides what goes in the window every time.
Treat it as a budget. Every token costs money and adds latency, and a window that fits is not the same as a window that works well. Keep the full history in your database, trim or summarise what you send, retrieve documents instead of stuffing them, and check the stop reason on every response. Start with trimming today, measure your token counts, and add the other patterns when the numbers tell you to.
Frequently Asked Questions
Is a Bigger Context Window Always Better?
No. A bigger window raises the ceiling, but you still pay for every token you put in it, and every extra token adds latency. Anthropic's documentation also warns that accuracy and recall degrade as the token count grows, which it calls context rot. A small, well-chosen prompt usually beats a large, noisy one.
Does the Output Count Toward the Context Window?
Yes. The window holds everything the model can reference while it answers, including the answer itself. Your system prompt, the whole message history, tool results, and documents count as input, and every token the model generates counts too. That is why a request can fail or stop early even when your input alone looked small enough.
What Happens When You Exceed the Context Window?
It depends on where you cross the line. If the input alone is too long, the Claude API returns a 400 invalid_request_error that says the prompt is too long. On newer Claude models, a request whose input plus max_tokens goes over the limit is still accepted, and generation stops with the stop reason model_context_window_exceeded when it hits the limit. Other providers have their own error shapes, so read the docs for the one you call.
Is the Context Window the Same as Memory?
No. The context window is working memory for a single request. The model does not remember anything between requests unless your backend sends it again. Long-term memory is something you build: a database of past messages, a summary, or retrieved documents that you add back into the next prompt.
How Do I Count Tokens Before Sending a Request?
Use the provider's own counter. Anthropic offers a token counting API that estimates usage before you send a message, and most providers publish a tokenizer for their models. A rough rule of thumb is 4 characters per token for English text, which is fine for a safety margin but not for billing.
Can I Run a Model With a Large Context Window on My Own Machine?
Yes, but memory is the limit. Ollama, a free tool for running open-weight models locally, picks a default context length based on your GPU memory: 4k tokens under 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k at 48 GiB or more. A larger context needs more memory, so check how much you have before raising it.



