Context Window Exceeded: Budget Prompts Before Retrying
Fix context-window errors by budgeting history, tools, retrieved text and output. Preserve relevant evidence instead of blindly truncating conversations.
Published 2026-09-29 · Updated 2026-09-29 · KeepRouter Editorial · 5 minute read

A context-window error means the request exceeds a model or endpoint limit. Count the complete serialized request, including instructions, history, tool definitions and retrieved content, then reserve room for output according to the model's rules. Retrying the same oversized payload normally reproduces the same error.
For a long-running assistant, the visible user question may be tiny while the actual request is large. Fixing that gap requires inspecting how the application builds context, not asking the user to shorten a sentence that contributes little to the total.
Distinguish the limits you are configuring
A model may have a context window, a maximum output size and endpoint-specific input restrictions. Client libraries can also apply their own limits. Claude's context-window documentation and Gemini's token guide describe model-specific accounting; do not assume one formula covers every API unchanged.
| Limit | Meaning | Common misunderstanding |
|---|---|---|
| Context window | Capacity available to the model's request/response process | The whole amount is always available for input |
| Maximum output | How much generation is allowed | Raising it increases input capacity |
| Client-side budget | What the application chooses to send | It changes the provider's model limit |
| File or media limit | Operation-specific input restriction | A large token window bypasses it |
Read the limit for the exact model revision and route. A similarly named model on another platform may have a different supported window or request format. A gateway alias is not sufficient evidence of identical limits.
Build a context ledger
For a simplified hypothetical 16,000-token budget, suppose instructions and tools use 2,000 tokens, history uses 5,000, retrieved passages use 6,000 and the new question uses 500. Input totals 13,500. Reserving 2,000 for output and 1,000 as application headroom would exceed the budget by 500.
context_budget = 16_000
parts = {"instructions_tools": 2_000, "history": 5_000,
"retrieval": 6_000, "question": 500}
output_reserve = 2_000
headroom = 1_000
available = context_budget - sum(parts.values()) - output_reserve - headroom
print(available) # -500: reduce input before callingThis is a planning worksheet, not a tokenizer or a published model limit. Use the service's token-counting method where available and account for how it treats media and reasoning. Character count divided by a fixed constant is only a rough estimate, particularly across languages and code.
Remove repetition before removing evidence
Start with duplicated system instructions, repeated tool descriptions and retrieval passages included more than once. These often consume budget without adding information. Then inspect old tool outputs: a full build log may have been useful once, while a concise error summary and file location are sufficient for the next step.
Do not remove the evidence needed to resolve the current question. If a user asks about one clause in a document, keep that clause and its source identity. Replacing the whole document with a vague summary can fit the request while making the answer less reliable.
For retrieval, cap both the number of passages and their combined size. Ten short passages and ten long passages are not equivalent. Prefer relevance and non-duplication over simply taking the first results until a character limit is reached.
Compact conversation state deliberately
A useful summary preserves decisions, unresolved questions, identifiers, constraints and references needed for the next action. It should not turn a tentative suggestion into an approved decision. Keep application-owned facts outside the model-generated summary when they control permissions, balances or workflow state.
Test a summary on a later question that requires an earlier constraint. For example, if the user said “do not edit generated files,” the next task should still obey that rule after compaction. A shorter context that loses the constraint is not a successful optimization.
When tools are involved, keep protocol-valid call/result pairs. Removing only the assistant call while leaving its result can create a different API error. Review the tool-loop guide before implementing arbitrary message deletion.
Use a larger model window when the task needs it
A larger supported window is useful when relevant material truly cannot be reduced without harming the task. It is less useful when an application resends redundant history forever. Compare correctness and cost on the same evidence rather than treating a larger advertised number as an automatic quality upgrade.
Long-context evaluation should include facts near the beginning, middle and end, conflicting passages and a question that the material cannot answer. Keep the expected evidence outside the prompt and inspect whether the model uses it correctly. Fitting a document into a request does not establish reliable retrieval of every fact inside it.
Turn deterministic failures into actionable responses
When a request is too large, return a clear application state: reduce history, narrow the document set or choose a suitable model. Do not retry it in a transient-error loop. Keep a redacted size breakdown so future failures can be diagnosed without storing the customer's full content.
Use the model catalog to inspect supported limits, the RAG budget guide to quantify repeated context, and the cost calculator before expanding the production window. Preserve a small regression fixture that verifies both size handling and the task's important facts.
Frequently asked questions
Will increasing max output tokens fix context errors?
Usually it does not increase input capacity and may require more reserved space. Check the exact model's accounting and reduce the complete request size.
Can I just delete the oldest messages?
Only if important constraints and protocol relationships remain intact. Preserve required facts and valid tool-call/result pairs, and test the resulting conversation.
Does a larger window guarantee better answers?
No. It changes capacity. Evaluate whether the model finds and uses the relevant evidence across the actual document set, including conflicting or absent facts.
Sources reviewed
Article last reviewed 2026-09-29