What Is Latency?
Latency is the delay between making a request and getting a response. For an AI system, it is the time from sending a prompt to receiving the model's answer. Lower latency means a snappier, more responsive experience, while high latency shows up as a noticeable wait for output.
How latency works
Latency is the total time a request spends being handled, from the moment it is sent to the moment a response comes back. For an AI model, that includes sending the input, the model doing the computation, and returning the result. Larger models generally take longer, because more computation stands between the prompt and the answer.
With text-generating models, one measure matters especially: time to first token, the wait before the first piece of the response appears. Because these models produce output gradually, a fast first token makes a system feel responsive even when the full answer takes a moment to finish.
Why latency matters for AI
Latency shapes how usable an AI system feels and how much it costs to run at scale. A slow response frustrates users and can make an otherwise capable system feel broken, so latency is a core target when moving from a working model to a production one. It also trades off against other goals: larger models are often more capable but slower, so teams balance quality, speed, and cost for the task at hand.
Related terms
Frequently asked questions.
The stuff we hear most on the first call. Don't see yours? Book a 30-minute conversation.
What is latency in AI?
What causes high latency in AI models?
What is time to first token?
Want to put AI
to work?
We work with leadership teams to find the right opportunities, define the strategy, and build the systems that move the business forward.