What Is a Context Window?

Models & Architecture Also known as: context length

A context window is the amount of text an AI model can consider at one time, measured in tokens. It covers everything the model reads for a single response, including the prompt, any provided data, and the conversation so far. Anything beyond the window is not seen.

How a context window works

Everything sent to a model for one response has to fit inside its context window, and text is measured in tokens, which are chunks of words. When a conversation or document exceeds the window, something has to be dropped or summarized, because the model cannot see past the limit. Larger context windows let a model take in more at once, such as a long document or a lengthy conversation, without losing the earlier parts.

Why the context window matters

The context window matters because it sets a hard ceiling on how much a model can work with at a time. Exceed it and the model loses the earliest information, which can make it forget instructions given at the start of a long conversation or miss parts of a long document. It is also a cost factor, since more tokens in the window usually means more to process. Managing what goes in the window is the job of context engineering.

Frequently asked questions.

The stuff we hear most on the first call. Don't see yours? Book a 30-minute conversation.

What is measured in a context window?
Tokens, which are chunks of text roughly the size of a word or part of a word. The window counts the prompt, any data, and the conversation history together.
What happens when you exceed the context window?
The model cannot see the text beyond the limit, so earlier information is dropped or has to be summarized, which is why long conversations can lose their earliest details.

Want to put AI
to work?

We work with leadership teams to find the right opportunities, define the strategy, and build the systems that move the business forward.