What Is Latency?

Latency is the delay between making a request and getting a response. For an AI system, it is the time from sending a prompt to receiving the model's answer. Lower latency means a snappier, more responsive experience, while high latency shows up as a noticeable wait for output.

How latency works

Latency is the total time a request spends being handled, from the moment it is sent to the moment a response comes back. For an AI model, that includes sending the input, the model doing the computation, and returning the result. Larger models generally take longer, because more computation stands between the prompt and the answer.

With text-generating models, one measure matters especially: time to first token, the wait before the first piece of the response appears. Because these models produce output gradually, a fast first token makes a system feel responsive even when the full answer takes a moment to finish.

Why latency matters for AI

Latency shapes how usable an AI system feels and how much it costs to run at scale. A slow response frustrates users and can make an otherwise capable system feel broken, so latency is a core target when moving from a working model to a production one. It also trades off against other goals: larger models are often more capable but slower, so teams balance quality, speed, and cost for the task at hand.

Frequently asked questions.

The stuff we hear most on the first call. Don't see yours? Book a 30-minute conversation.

What is latency in AI?
Latency in AI is the delay between sending a request to a model and receiving its response. It is usually measured in milliseconds or seconds and reflects how quickly the system reacts.
What causes high latency in AI models?
Mainly model size and the amount of computation per request, along with hardware limits, network delays, and long inputs or outputs. Bigger models and longer responses take more time.
What is time to first token?
It is the wait before a text model produces the first piece of its response. A short time to first token makes a system feel responsive, since output then streams in rather than arriving all at once.

Want to put AI
to work?

We work with leadership teams to find the right opportunities, define the strategy, and build the systems that move the business forward.