Skip to content

Context Window Explained: Why Size Matters and Where It Breaks

Context window size shapes what an LLM can reason about: GPT-4.1 lists 1,047,576 tokens, Gemini 2.0 Flash holds 1M, and o3 caps at 200K. Learn why exceeding it fails hard.

Google Gemini AI developer platform
Google Gemini · Credit: Google

A context window is the fixed-size token buffer that a language model reads in full before generating any output. Everything outside that buffer is invisible to the model: an earlier conversation turn, a paragraph past the cutoff, or a document that overflows by a single token might as well not exist. The size of this buffer governs what a Transformer can reason about, what it forgets, how much each API call costs, and where the application has to compensate with engineering. Recent vendor documentation puts the spread at roughly three orders of magnitude, from a few thousand tokens on an embedding endpoint to about a million on the largest chat models. The exact ceilings, model by model and with sources, appear in the comparison table below. The takeaway is that a single vendor can ship several different limits at once, and each one is a different failure mode when an input crosses the line.

What a Context Window Actually Is

A black outline of a human head in profile, with a white network of connected circles inside it, on an orange background.
Credit: Anthropic

A context window is the total number of tokens a model can hold in its working memory during a single inference call, covering everything the model sees before producing output. The constraint comes from the Transformer architecture itself: attention computes pairwise relationships across the input, and the model's positional embeddings and KV cache are sized for a fixed maximum sequence length. Tokens are not words. They are subword units produced by a tokenizer such as OpenAI's cl100k_base or Google's SentencePiece variant, and the ratio of English words to tokens hovers near 0.75. A 1,000-word prompt typically consumes around 1,300 input tokens, a number that grows when the text is code, JSON, or non-Latin script.

For practitioners coming from how large language models process and predict tokens, the window is best understood as the slot of attention a model can pay at once, not as a memory the model keeps between calls. Each request to a stateless API resends the full conversation, and the model re-reads it from scratch. Background on the surrounding architecture lives in the explainer on how modern generative AI systems work.

Token
A subword unit produced by the tokenizer. Roughly four characters of English text per token.
Context length
The maximum number of tokens the model accepts in a single request, set by the model architecture and the published API limit.
Working memory
The portion of the buffer occupied by the system prompt, prior turns, retrieved documents, tool outputs, and the assistant's reply in progress.

How Tokens Are Counted Inside the Context Window

The context window budget is consumed by several token categories simultaneously, and understanding their accounting is essential for building reliable applications. OpenAI's completions API reference is explicit on the central rule: "the token count of your prompt plus max_tokens cannot exceed the model's context length" (OpenAI). That single sentence drives most of the application-level discipline around the window. A 200K-token model with a 100K max output cap leaves 100K for everything else, and that residual is shared among system instructions, retrieved chunks, prior turns, tool definitions, and tool call results.

  1. System prompt and developer instructions. Fixed overhead on every call. A long system prompt is a standing tax against the remaining budget.
  2. Conversation history. All prior user and assistant messages, resent on each turn unless the application truncates, summarizes, or uses a server-side conversation state feature.
  3. Retrieved context. Documents pulled by retrieval-augmented generation, tool outputs, and any inline files. This is typically the largest variable cost.
  4. Tool and function schemas. JSON definitions for every function the model can call. A dozen tools can easily run several thousand tokens.
  5. Reserved output budget. The model must hold space for the response it has not yet generated. The max_tokens parameter caps that reservation.

Because the limit is a sum, raising max_tokens shrinks the room available for input tokens. Engineers building agent loops learn this the hard way: a long tool-output history can leave less than a single sensible reply's worth of headroom inside a nominally generous window.

Context Window Sizes Across Current Models

Context window sizes vary substantially across current production models, and the exact figures come from vendor API documentation rather than marketing summaries. The headline numbers most teams use in planning are the ones below, each sourced from the producing vendor's own model page. The token limit a model advertises is the ceiling, not a guaranteed reasoning depth across the full span: longer prompts often see attention quality degrade well before the hard cap.

ModelContext window (tokens)Max output tokensSource
GPT-4.11,047,57632,768OpenAI
o3200,000100,000OpenAI
Gemini 2.0 Flash1,000,000vendor-documentedGoogle
text-embedding-3-small8,191n/a (embedding)OpenAI

Two patterns stand out in this snapshot. First, embedding models live in a different regime. The 8,191-token cap on text-embedding-3-small with the cl100k_base encoding forces application code to chunk long documents before vectorization, a constraint that does not relax with chat-model scaling. Second, reasoning models trade raw window size for inference depth: o3's 200K ceiling is small next to GPT-4.1, but its 100K max output budget makes room for the long internal reasoning traces the model produces before answering. Google describes Gemini as "the first model capable of accepting 1 million tokens" in its long-context documentation, a milestone framing that anchors the current shape of the market.

What Happens When a Prompt Exceeds the Context Window

When a prompt exceeds the context window, the model cannot silently expand its capacity: the request either fails with an error or the input is truncated, and both outcomes degrade output quality. The default behavior depends on the API, the SDK, and the model family. OpenAI's embedding cookbook is the clearest worked example: text-embedding-3-small has a context length of 8,191 tokens, and "going over that limit causes an error" (OpenAI). Chat endpoints behave similarly when prompt tokens plus max_tokens cross the ceiling.

  1. Hard error. The most common production failure. The API returns a 400 with a context-length-exceeded message, and the caller is expected to shorten the input or lower max_tokens.
  2. Head or tail truncation. Some client libraries drop the oldest messages or trim the tail of a long document. The model answers, but it answers about a different question than the user asked because the missing material happened to contain the constraint.
  3. Silent quality drop. Even within a generous token limit, attention over the full span is not uniform. Long-context studies consistently show degraded retrieval of facts placed in the middle of the prompt, the lost-in-the-middle effect.
  4. Tool-call thrash. Agent loops that summarize prior tool output into the next request can compound rounding errors, dropping cited identifiers that the next step needs.

The defensive posture is to count tokens before sending, reserve a known max_tokens floor for the answer, and decide explicitly how to shed material when the budget is tight. Truncation should be a documented policy, not an SDK default the team has not read.

Reasoning Tokens and Their Effect on Context Window Budget

Reasoning models introduce a third token category that occupies context window space alongside input and output tokens, compressing the effective budget available for the prompt and response. OpenAI's reasoning guide states the rule directly: "reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens" (OpenAI). On models such as o3, that hidden cost can be substantial. The 100,000 max output tokens that o3 advertises is the cap on visible response plus invisible reasoning combined.

  • Hidden consumption. Reasoning tokens are produced before the visible answer and counted against the same budget.
  • Billed as output. Their cost-per-token matches the output tier, not the cheaper input tier, so a hard problem becomes a hard bill.
  • Variable per request. The same prompt can produce a short trace on one call and a far longer trace on the next, depending on how the model decides to reason.
  • Affects max_tokens math. If a caller reserves only a small output budget for the answer and the model spends most of it on reasoning, the user sees a truncated reply.

Building reliably against reasoning models means treating the reasoning budget as a planning variable rather than a hidden detail. Teams that watch usage in production usually set max_completion_tokens with explicit headroom for the trace, then measure typical reasoning lengths per task type. The deeper mechanics of inference-time reasoning, including the trade-offs of test-time compute, are covered in the explainer on how reasoning models use chain-of-thought and test-time compute.

Prompt Caching: Reducing the Cost of Large Context Windows

Prompt caching lets developers store frequently reused context between API calls so the model does not re-process identical input tokens on every request. The economics matter at scale. A very large system prompt sent on every one of tens of thousands of daily requests multiplies into billions of input tokens of repeated work. Anthropic, which exposes prompt caching directly on its API, states that the feature "enables developers to cache frequently used context between API calls" and reports cost reductions of up to 90% and latency reductions of up to 85% for long prompts (Anthropic). Those are vendor-claimed maxima rather than guaranteed averages, but the directional saving is real for any workload with a stable preamble.

  • What gets cached. The serialized prefix of the prompt, typically the system message and any large reference documents that do not change between calls.
  • How hits are billed. Cached input tokens are charged at a discounted rate relative to fresh input tokens. The write to the cache itself usually costs a small premium on first use.
  • Time-to-live. Caches expire on a vendor-defined window. A workload with bursty traffic earns fewer hits than a steady stream of requests.
  • What it does not change. Prompt caching reduces the cost and latency of reusing context. It does not raise the token limit. A cached prefix still consumes its full share of the window.

The mental model is closer to a CDN than to a memory upgrade. Caching pays off when the same long prefix is reused across many calls. It does not help a one-shot summarization of a large document, and it does not let the model attend to more material than the architecture supports.

Workarounds When the Context Window Is Not Enough

When the context window cannot hold all necessary input, three main architectural patterns extend an application's effective memory beyond the raw token limit. Each pattern trades a different cost against a different failure mode, and most production systems combine them. Google's long-context documentation lists summarization of large corpuses, question answering, and agentic workflows as the canonical use cases that the patterns below address (Google).

  1. Retrieval-augmented generation (RAG). The application indexes a corpus into an embedding store, retrieves the top relevant chunks at query time, and inserts only those chunks into the prompt. RAG keeps the prompt small and adapts to a knowledge base that updates faster than the model can be retrained. The trade-off is recall: a question whose answer spans many distant chunks may not be served by a top-k retrieval.
  2. Hierarchical summarization. Long documents are summarized in passes. Section summaries roll up to chapter summaries, and the model reasons over the condensed structure. This pattern fits codebases, legal filings, and meeting transcripts where the source is too long for any window but the relevant signal compresses well.
  3. Sliding window with state. The application keeps a running summary of prior turns, drops the oldest verbatim messages once they age out, and reinjects the summary on each call. Useful for long-running chat sessions and agents where conversation length grows without bound.

None of these patterns is a substitute for a longer raw window when the question genuinely depends on cross-referencing material across an entire long document. They are substitutes for a longer window when the task is selective: most reads of a knowledge base touch a small fraction of it, and only a fraction of that fraction matters for the current answer. The right architectural choice depends on whether the bottleneck is retrieval quality, repeated cost, or attention depth across the full prompt. For deeper background on the model-side constraints that drive these patterns, see the explainer on architectural limits of large language models.

References

Frequently Asked Questions

What is a context window in an AI language model?

A context window is the maximum number of tokens a language model can process in one inference call. Everything outside that limit is invisible to the model: earlier conversation turns, documents, or instructions beyond the token ceiling are simply not available to inform the response, per Anthropic and Google's API documentation.

What happens when a prompt exceeds the context window limit?

Exceeding the context window causes a hard error or silent truncation depending on the API and SDK. OpenAI's completions API documentation states that the combined token count of the prompt plus max_tokens cannot exceed the model's context length, so the request will fail if that ceiling is crossed without truncation handling in the application layer.

Do reasoning tokens count against the context window budget?

Yes, reasoning tokens occupy space in the context window even though they are not visible via the API. OpenAI's reasoning guide confirms that these tokens are billed as output tokens and consume context window capacity, which means the effective budget for the visible prompt and response shrinks on reasoning models.

What is prompt caching and how does it relate to context window cost?

Prompt caching stores frequently reused context on the server so the model does not re-process the same input tokens on every request. Anthropic states that prompt caching can reduce costs by up to 90% and latency by up to 85% for long prompts, making large context windows more economical for applications that reuse a stable system prompt or knowledge base.

Share this guide

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.