Real-Time Voice AI: The Latency Budget
A voice agent is a budget problem before it is a speech problem. You have roughly 800 milliseconds between the end of a sentence and the start of the reply, and nine places to spend it. This page is where they go and how to defend them.
Conversation runs on a clock that predates your product. Across ten languages the median gap between a question and its answer is about 100 ms. An agent that takes 859 ms is not a bit slow. It is outside the range human turn-taking has ever been measured in.
The budget is not dominated by the model. In the worked example below, deciding that the user stopped talking costs 295 ms, the language model costs 220 ms, and moving audio around costs 245 ms. Most teams optimise the middle one.
Every lever here is a trade against a different failure. Shorten the endpointer and you interrupt people. Shrink the jitter buffer and the recogniser eats concealment noise. Drop the cascade for a native speech model and you lose the transcript you were logging.
01The clock you are running against
Before any engineering, one measurement. It is the only number on this page you cannot negotiate with.
Stivers and colleagues timed the gap between a yes/no question and its answer in ten languages drawn from five continents. Every language produced a unimodal distribution with its mode between 0 and 200 ms. The cross-linguistic median was +100 ms. Language means ran from +7 ms in Japanese to +469 ms in Danish.
Your agent competes against that distribution, and 469 ms was the slow end of it. A turn that takes 859 ms sits at roughly eight times the cross-linguistic median and nearly double the slowest language mean measured.
Telephony arrived at a compatible number from the other direction. ITU-T Recommendation G.114 says that below 150 ms of one-way mouth-to-ear delay most applications feel almost transparent, and that above 400 ms the delay is unacceptable for general network planning.
So the working line comes from arithmetic and needs no benchmark. Two one-way trips at the ITU's unacceptable limit come to 800 ms, and the human distribution is almost entirely below 500 ms. Somewhere around 800 ms of turn latency, a reply stops landing anywhere a conversational gap has been observed to land.
OpenAI's GPT-4o system card reports audio responses in as little as 232 ms with an average of 320 ms, and explicitly positions that as human-comparable. Kyutai's Moshi paper gives a theoretical 160 ms, decomposed as an 80 ms Mimi frame plus 80 ms of acoustic delay, and about 200 ms in practice on an L4.
Both are native speech models. A 2026 tutorial that built a careful self-hosted cascade from streaming ASR, vLLM and streaming synthesis measured 755 ms time to first audio, best case 729 ms. That gap between 320 and 755 is what the rest of this page is about.
02Where the milliseconds go
Nine stages, measured from the user's last phoneme to the first sample of the reply reaching their ear. Read the long bar first.
Each row, and where its number came from:
| Stage | ms | Where the number comes from |
|---|---|---|
| Capture + Opus encode | 25 | One 20 ms frame plus the 5 ms of look-ahead RFC 6716 says the SILK layer needs for noise shaping. |
| Network, up | 30 | One direction of a 60 ms round trip, inside the 20 to 200 ms range ElevenLabs' latency docs quote. |
| Jitter buffer | 80 | NetEq's starting target delay in the WebRTC source, before it adapts to the measured arrival spread. |
| Endpointing | 295 | LiveKit Turn Detector v1's mean latency at a 10% false-cutoff operating point on eot-bench. |
| ASR final transcript | 24 | NVIDIA's reported median time to final transcription for cache-aware streaming Nemotron ASR. |
| LLM first token | 220 | Assumed. This one you must measure yourself. Page 31 is about what sets it. |
| TTS first audio | 75 | ElevenLabs' stated Flash model inference time for short inputs, excluding network and player buffering. |
| Network, down | 30 | Symmetric with the uplink. |
| Playout buffer | 80 | The same jitter mechanism on the return path. |
The budget rests on one assumption and eight sourced figures, and totals 859 ms. The shape matters more than the exact sum, so push on the numbers yourself.
Try it Price three optimisations against the same budget
import numpy as np
# One turn, end of the user's speech to first agent audio.
# Every row is sourced or flagged as an assumption on the page.
stage = ["capture + Opus encode", "network up", "jitter buffer",
"endpointing", "ASR final transcript", "LLM first token",
"TTS first audio", "network down", "playout buffer"]
ms = np.array([25, 30, 80, 295, 24, 220, 75, 30, 80], float)
total = ms.sum()
print(f"budget {total:.0f} ms")
for i in np.argsort(-ms)[:4]:
print(f" {stage[i]:<21} {ms[i]:5.0f} ms {ms[i]/total:5.1%}")
moved = ms[[0, 1, 2, 7, 8]].sum()
print(f"\naudio moving, not thinking {moved:.0f} ms")
print(f"LLM first token {ms[5]:.0f} ms")
def edit(i, v):
x = ms.copy(); x[i] = v; return x.sum()
print(f"\nhalve the LLM {edit(5, 110):.0f} ms")
print(f"endpointer at 5% cutoff {edit(3, 543):.0f} ms")
print(f"fixed 800 ms silence {edit(3, 800):.0f} ms")
749. Loosening the endpointer from a 10% to a 5% false-cutoff rate moves it to 1107, which erases that win three times over. A plain 800 ms silence timer lands at 1364. The turn-taking policy has more leverage than the model, and it is a config value.Tool calls and retries add to this. A function call in the middle of a turn inserts a second full model round trip, and nothing about the first one streams across it. One HTTP retry on any hop costs more than the entire budget above.
Both are invisible at p50 and dominate p95. Latency here is a distribution, and the complaints come from its tail.
03Endpointing: deciding that you stopped
The largest single item in the budget is a decision and not a computation, and getting it right is the highest-leverage work in voice AI.
Voice activity detection is cheap. Silero VAD processes a 30 ms chunk in under a millisecond on one CPU thread, from a JIT model of about two megabytes. Knowing whether audio contains speech is a solved, nearly free problem.
Endpointing is the hard one: has this person finished, or are they thinking? A pause of 300 ms means both things, and the audio is identical.
Fixed thresholds, and why they cannot win
The default policy is a silence timer. Wait $T$ milliseconds of non-speech, then commit the turn. It has exactly two behaviours, and they are the two failure modes:
- False cut-off. A mid-sentence hesitation lasts longer than $T$, so the agent interrupts. The user has to start over, and the recovery costs several seconds.
- Dead air. The user finished, and you pay $T$ milliseconds of silence before anything happens. Every single turn pays it.
The trade is direct and it has no free parameter. Raising $T$ reduces cut-offs and raises dead air by exactly the same amount, because dead air is $T$.
Semantic endpointing moves the curve
The fix is to look at what was said as well as whether sound is present. "My account number is four two" is unfinished after 400 ms of silence. "That's all, thanks" is finished after 150 ms. A small language model reading the partial transcript can tell them apart.
LiveKit's Turn Detector v1, released in June 2026, distils a Qwen2.5-7B teacher into a 0.5B student for CPU inference, and reports numbers on an open benchmark called eot-bench:
| Latency budget | LiveKit v1 | Deepgram Flux | Other reported |
|---|---|---|---|
| 300 ms | 9.9% false cut-off | 12.9% | ultraVAD 27.7% |
| 600 ms | 4.5% false cut-off | 9.9% | Soniox 5.5% |
Read that as a curve and not a ranking. The same post gives the inverse view: a 10% false-cutoff target costs 295 ms of mean latency, and a 5% target costs 543 ms. Halving your interruption rate costs 248 ms on every turn, which is the exchange rate you are buying.
Try it The endpointing trade, at matched error rates
import numpy as np
rng = np.random.default_rng(0)
N = 60000
is_end = rng.random(N) < 0.5 # the user really has finished
# Mid-turn hesitations in ms. A real turn end never resumes.
pause = np.where(is_end, np.inf,
rng.lognormal(np.log(240), 0.6, N))
hes = ~is_end
def fixed(T):
"""Wait T ms of silence, then commit. Dead air is always T."""
return (hes & (pause > T)).sum() / hes.sum()
# A semantic endpointer reads the partial transcript at 120 ms and
# commits early when confident. The separation below is invented;
# only the shape of the curve is the point.
score = rng.normal(np.where(is_end, 1.7, 0.0), 1.0)
def semantic(c, T=800.0):
early = score > c
bad = (hes & early & (pause > 120)) | (hes & ~early & (pause > T))
wait = np.where(early, 120.0, T)
return bad.sum() / hes.sum(), wait[is_end].mean()
grid = np.arange(50, 3000, 5.0)
rates = np.array([fixed(t) for t in grid])
print(" false cut | fixed silence | semantic | saved")
for c in (0.4, 1.0, 1.6, 2.2):
r, wait = semantic(c)
t_match = grid[np.argmin(np.abs(rates - r))] # same error rate
print(f" {r:6.2%} | {t_match:6.0f} ms | {wait:5.0f} ms |"
f" {t_match - wait:5.0f} ms")
125 to 146 ms across the four rows, because it can commit before the silence has elapsed. Second, and more important, the curve is steep in both columns: going from 16.16% to 7.22% false cut-offs costs 148 ms of extra dead air. Neither policy escapes that. You are choosing a point on a curve, and no setting is correct.Start generating on the partial transcript while the endpointer is still undecided, and cancel if the user keeps talking. You pay for speculative model turns you throw away, and in exchange the LLM's 220 ms runs concurrently with the endpointer's 295 and does not follow it.
It only works if cancellation is cheap and complete. A cancelled turn that still emits audio, or still writes to conversation state, is worse than the latency it saved.
That covers the budget, the one term that dominates it, and the trade that sets that term, which together make up the listening half of a turn.
The rest of the page is the other three problems: being interrupted, the accuracy you pay for streaming, and the fact that all of this is running over a lossy UDP connection you do not control.
04Barge-in, and the state it corrupts
Interruption is usually specified as "stop the audio". It is three separate problems, and the third one is where the bugs live.
Hearing the user through your own voice
The microphone picks up the loudspeaker. Without cancellation, the agent's own output is the loudest thing in the input, so the VAD fires on the agent and the agent interrupts itself.
Acoustic echo cancellation solves this by subtracting a filtered copy of the reference signal, which is why the playback path has to hand its samples to the AEC. RFC 7874 says WebRTC endpoints SHOULD include an AEC and deliberately mandates no particular algorithm.
The hard case is double-talk: both parties speaking at once, which is exactly what barge-in is. The adaptive filter must not diverge while it is being trained on a mixture. The ICASSP 2022 AEC Challenge scored entries on subjective quality across single-talk and double-talk conditions plus a speech recognition word acceptance rate, which is the right pairing: residual echo that a human tolerates can still wreck a transcript.
Audio you cannot take back
When you decide to stop, some of your speech has already left. Frames handed to the playout buffer are gone from your control. With the 80 ms buffer from section 02 plus a 20 ms frame, assume up to about 100 ms of speech reached the ear after your cancel.
This is a small number with a large consequence, because the agent said a word you believe it did not say.
On barge-in, the naive implementation stops playback and writes the full generated response into the conversation history. The model now believes it delivered a sentence the user never heard, so it never repeats it and answers as if the information landed.
The fix is to truncate the assistant turn at the point where audio stopped. That requires alignment between generated text and emitted audio, which is why streaming synthesis should return timing marks as well as bytes. Without them you are guessing where to cut.
Deciding it was an interruption at all
"Mhm", "right" and "yeah" are backchannels. They mean keep going, but a VAD classifies them as speech and stops the agent, which is why over-eager barge-in feels as broken as none at all.
Separating a backchannel from a real interruption needs the words, because energy alone does not separate them, so it lands on the same model that does semantic endpointing. In a cascade, this means your barge-in decision depends on your ASR having already produced a partial transcript for a 200 ms utterance.
Cancellation also has to reach every component. Stopping the synthesiser while the language model keeps generating leaves an orphaned response that arrives after the user's next question and answers the previous one.
05Streaming ASR: what a chunk costs
Three knobs decide the accuracy you pay for arriving early. Two of them turn out not to cost you turn latency at all.
A batch recogniser conditions on the whole utterance. A streaming one must commit to words while audio is still arriving, using everything before plus a short window after. That window has several names: lookahead, right context, emission delay, chunk latency.
NVIDIA published a clean version of the exchange rate for its cache-aware streaming Nemotron models in January 2026. Raising chunk latency from 0.16 s to 0.56 s drops word error rate from 7.84% to 7.22%.
Working it out, 400 ms of extra delay buys 0.62 WER points, about 0.16 points per 100 ms. The trade is unusually mild, which is the interesting part.
Chunk delay is spent while the user is still speaking. It overlaps the utterance and does not append to it. Only the last chunk's delay lands after the final word, and even that mostly overlaps the endpointer's wait.
So the ASR row in section 02 is 24 ms and not 160. The useful question is how much of the lookahead is still unpaid when the endpointer fires, and usually the answer is none of it.
Why you cannot just chunk an offline model
Slicing audio into windows and running a batch model on each looks like streaming and is not. Pages 13 and 34 cover why the accuracy degrades. The cost that matters here is compute.
Sliding windows overlap, so the encoder repeatedly re-processes audio it has already seen, and recomputes attention over context it already had. A cache-aware streaming encoder keeps the activations instead and processes each frame once.
NVIDIA reports the size of that difference as concurrency: at a 320 ms chunk on an H100, 560 simultaneous streams against a 180-stream baseline, roughly three times. For a voice product, concurrency per GPU is the unit cost of a conversation, so this is a pricing decision wearing an architecture costume.
The third knob is the one nobody sets
Left context matters too. A streaming encoder that keeps unbounded history grows its state per stream, and memory per stream is what caps concurrency. Bounding it costs a little accuracy on long turns and is frequently the difference between one GPU and four.
06Streaming TTS: only the first chunk matters
Total synthesis time is almost irrelevant. Time to the first audible sample is the entire metric, and the biggest term in it is not in the synthesiser.
Once playback starts, generation only has to keep up with real time. A real-time factor below 1.0 means the buffer never drains. Everything else is time to first audio, and it decomposes further than most teams look.
ElevenLabs' own latency documentation is unusually candid about this. Flash model inference is around 75 ms for short inputs. Network round trips are 20 to 200 ms. A 500 ms player buffer is common. The model is the small term.
The chunking decision that costs half a second
The synthesiser cannot start until it has text, and the text arrives token by token. What you choose to wait for is therefore a latency decision made upstream of the TTS.
Take the arithmetic: at 40 tokens per second, one token is 25 ms. Waiting for a complete 20-token sentence costs 500 ms after time to first token. Cutting at the first clause boundary, say 6 tokens, costs 150 ms.
This is a 350 ms difference from a text-splitting rule, larger than anything you will win by changing synthesiser. It is also why "we stream the TTS" is an ambiguous claim worth interrogating: streaming audio out while buffering text in to a sentence boundary gives you the latency of a non-streaming system.
Prosody is planned over a phrase. Synthesise "I can help with that" and "if you give me the account number" as independent chunks and you get two falling intonations glued together, which sounds like two sentences.
The usable compromise is a short first chunk for speed and longer chunks after, since the first one buys the time for everything behind it. You are trading the naturalness of one phrase for the responsiveness of every turn.
Where the vocoder sits in this
The last stage turns frames into samples, and it cannot start from a single frame. Convolutional vocoders need enough frames to fill their receptive field, plus some lookahead, before the first sample exists. Chunks are also spliced with an overlap from the previous chunk so the seam is not audible.
Whether that stage is your bottleneck depends entirely on where you run. A 2025 comparison of low-latency neural vocoders found that on CPU-only systems the per-chunk streaming overhead, data movement and parameter loading, dominates raw arithmetic at low latencies. The same work reports that a single frame of lookahead recovers nearly all of the quality of a non-causal model.
On a server GPU the picture inverts: the autoregressive token generation in front of the vocoder is the slow part, and the vocoder is a rounding error. Claiming one universal bottleneck is how teams end up optimising the wrong stage.
Every stage of the cascade, priced. At this point the obvious question is whether the cascade is the right shape at all, given that two of its serial stages exist only to convert between audio and text.
Then the layer underneath all of it, which is a UDP connection with a loss rate you do not control.
07Cascade or speech-to-speech
The 2026 trade is real and it is not close to settled. Both options give up something specific, and the specifics are what decide it.
A cascade is ASR, then a text model, then synthesis. A native speech model consumes and emits audio, with no text bottleneck in the loop. Page 34 has the model-by-model view; this is the engineering consequence.
| Cascade | Speech to speech | |
|---|---|---|
| Reported turnaround | 755 ms measured in a 2026 self-hosted build | Moshi 200 ms on an L4; GPT-4o 320 ms average |
| Intermediate state | A readable transcript at every hop | Activations. No transcript unless you add an ASR pass |
| Swapping a component | Independent, one at a time | All or nothing |
| Prosody, emotion, overlap | Discarded at the transcript | Preserved end to end |
| Barge-in | A distributed cancel across three services | Native, because it never had a turn state |
| Guardrails and tools | Text-level, mature, easy to constrain | Harder to constrain mid-stream |
Be precise about what the cascade loses, because "information loss" undersells it. The transcript destroys hesitation, sarcasm, emphasis and emotional state, and it also destroys timing. Overlapped speech becomes either one sequence or two, and which one is a decision your diarizer made.
The cascade also serialises three separate warm-up latencies, and gives you three services that can independently rate-limit, retry or go down mid-turn.
What speech-to-speech loses is mostly observability. You cannot log what you never rendered, so transcripts, retrieval keys, quality evals, compliance review and text guardrails all need a parallel ASR pass that you were trying to avoid running.
The 2026 tutorial cited above is useful precisely because it tried both. Their cascade of streaming ASR, vLLM and streaming synthesis measured 755 ms. An end-to-end Qwen3-Omni reached 702 ms through a hosted API, and 146 seconds run locally.
That last figure is the real state of play. Native speech models are faster when somebody else operates them at scale, and mostly not self-hostable yet at conversational speed. Regulated and self-hosted deployments choose the cascade for reasons that are only partly about the model.
08Transport: why WebRTC, and what loss does
245 ms of the budget is audio in transit or sitting in a buffer. The protocol choice underneath that is not incidental.
Why not HTTP
TCP guarantees ordered delivery, so a lost segment is retransmitted and every later segment waits behind it. This is head-of-line blocking, and it is the wrong guarantee for a live conversation.
Audio that arrives 200 ms late is worthless; its moment has passed. It is better to conceal 20 ms of missing sound than to stall the next 200 ms waiting for it. WebRTC carries audio over SRTP on UDP, so a lost packet leaves a hole and does not stall the stream.
WebSockets over TCP remain fine for the control plane and for server-to-server hops on a reliable network. They are the wrong choice for the last mile to a phone on a train.
Opus, and the 25 ms you pay before sending anything
RFC 7874 makes Opus and G.711 mandatory for WebRTC endpoints, and requires decoders to support all Opus modes. Opus is not a preference here, it is the floor.
RFC 6716 gives its structure. The SILK layer handles 10 to 60 ms frames and needs 5 ms of extra look-ahead for noise shaping. The CELT layer handles 2.5 to 20 ms frames with 2.5 ms of overlap look-ahead. A 20 ms frame plus SILK's 5 ms is the 25 ms in the budget table, spent before a single packet is on the wire.
Opus also carries in-band forward error correction: a low-bitrate copy of the previous frame rides inside the next packet, so one isolated loss is recovered without a retransmission and without a round trip. It covers exactly one frame back, so a burst of four defeats it. Mozilla's 2016 experiments found the effect subtle at low loss rates and very noticeable at high ones.
The jitter buffer is a latency knob
Packets sent every 20 ms do not arrive every 20 ms. A buffer absorbs the spread by holding audio back, which is why it appears twice in the budget.
WebRTC's NetEq pulls 10 ms of audio at a time and adapts continuously. Its delay manager keeps a histogram of packet inter-arrival times in 20 ms buckets and targets roughly the 97th percentile, starting from 80 ms. When the buffer runs long or short it time-stretches the audio, accelerating or expanding, and it does not drop or stall.
Try it Buffer depth against concealed packets, and where the knee is
import numpy as np
rng = np.random.default_rng(1)
# 20 ms Opus packets. One-way delay is a 30 ms floor plus a heavy
# right tail: most packets are early, a few are very late.
n = 400000
sent = np.arange(n) * 20.0
jitter = rng.gamma(2.0, 7.0, n) # ms
arrive = sent + 30.0 + jitter
print("buffer concealed added delay")
for depth in (0, 20, 40, 60, 80, 120):
deadline = sent + 30.0 + depth
late = (arrive > deadline).mean()
print(f"{depth:4} ms {late:8.3%} {depth:8} ms")
q97 = np.quantile(jitter, 0.97)
print(f"\n97th percentile of the jitter: {q97:.0f} ms")
print("NetEq sets its target near that quantile, then adapts.")
22.020% to 2.187%. Another 40 ms takes it to 0.013%, for the same 40 ms of delay. The 97th percentile of this jitter distribution is 37 ms, which is exactly where the collapse happens. So NetEq tracks a high quantile of measured arrivals and does not use a constant: a fixed buffer is either wasting delay or shredding audio, and which one depends on a network you do not control.What concealment does to the recogniser
A concealed packet is not silence. The decoder invents plausible audio to bridge the gap, tuned to sound acceptable to a human ear. Your ASR was not trained on it.
A 2022 study over G.722 telephony found the behaviour is a knee and not a slope: a clean-trained recogniser held up to about 10% packet loss and then degraded roughly in proportion, while a model trained on network-distorted speech pushed that knee out to about 15%. Training on the distortion you will see moves the cliff.
There is a cheaper version of the same idea. Dissen and colleagues showed at Interspeech 2024 that a front-end adaptation network in front of a frozen Whisper, trained against the recogniser's own loss, recovers much of the loss-induced error without retraining the ASR.
Published word error rates come from files. Your audio came through a codec, a jitter buffer, a concealment algorithm and possibly a transcode at a carrier boundary down to 8 kHz.
If you have never measured your own WER on audio captured at the far end of your own transport, you do not know your recognition accuracy, only somebody else's.
09Build this
The budget in section 02 is a worked example with one assumed row. Yours will look different, and the difference is the only thing worth acting on.
You are not building an agent. You are building the measurement that tells you which stage to work on, and then checking whether the arithmetic predicted what happened.
- Put a monotonic timestamp at nine points on every turn: last speech frame in, endpoint decision, final transcript, first LLM token, first TTS byte, first byte on the wire, first sample played. Ship them with the turn id.
- Collect 200 real turns, and avoid staff turns. Plot p50 and p95 per stage as a stacked bar, widths proportional to milliseconds, the same way the diagram above is drawn.
- Find the gap between p50 and p95 for each stage. One stage will own nearly all of it. That stage is your actual problem, and it is usually not the one that is largest at p50.
- Log every pause with its duration and whether the user resumed. You now have a labelled endpointing dataset for free.
- Sweep the silence threshold over that logged data offline and plot false cut-offs against dead air. Pick a false-cutoff target first, then read off the threshold.
- Deploy that one change and re-measure, then compare the movement in total latency against what step five predicted.
10What breaks
Almost none of these are model failures, and almost all of them show up as "it feels laggy".
- p50 is fine and p95 is three seconds. One slow hop, a cold start or a retry. Users remember the tail, and a mean hides it completely.
- The endpointer was tuned in an office. Staff speak in fluent complete sentences to a demo. Customers hesitate, read numbers aloud and pause to find documents.
- Barge-in stops the audio but not the model. The orphaned generation arrives later and answers the previous question.
- The transcript records what was generated and not what was heard. After every interruption the model believes it delivered a sentence that was cut off mid-word.
- The synthesiser buffers to a sentence boundary. Streaming audio out while blocking on text in gives you non-streaming latency behind a streaming API.
- The jitter buffer is a constant. Tuned on a wired desk connection, then deployed to mobile networks, where it is now too shallow.
- Echo cancellation fails on a speakerphone. The agent hears itself, the VAD fires, and it interrupts itself in a loop.
- The first tool call doubles the turn. Nothing streams across a function call, so the user gets the full budget twice with no audio in between.
- Backchannels stop the agent. Someone says "mhm" and the agent goes quiet, which reads as being ignored.
- WER was measured on clean files. The production path adds a codec, concealment and possibly an 8 kHz transcode that no benchmark includes.
11Where you meet this in the wild
Same budget, very different constraints on which lines you can touch.
8 kHz narrowband, a carrier path you cannot instrument, and echo control happening somewhere you do not own. The transport lines of the budget are fixed, so endpointing and the model turn are all you can move.
Semantic endpointing, aggressive speculative generation, and WER measured on your own recordings.
WebRTC end to end, so you control the codec, the buffer and the AEC. The catch is that you also own the AEC, and laptop speakerphones in open-plan offices are the worst case for double-talk.
Tune the jitter buffer per session, ship timing marks with synthesised audio.
Network latency goes to zero, which removes about 220 ms. In exchange the vocoder and the recogniser now share one CPU, and concurrency stops mattering while single-stream latency starts to.
Small streaming ASR, short lookahead, a vocoder chosen for per-chunk overhead.
Not a place to assume a server-class bottleneck profile.
Constant background noise, multiple speakers, and people who change their order mid-sentence. Endpointing carries the product, because a false cut-off is a wrong order and not just a slow reply.
Long thresholds, target speaker extraction, explicit confirmation.
12Interview questions
BeginnerWhy is 800 ms the wrong number for a voice agent?
Because human conversation does not work at that speed. Stivers and colleagues measured the gap between a question and its answer across ten languages and found a unimodal distribution in every one, with modes between 0 and 200 ms and a cross-linguistic median of about 100 ms. Even the slowest language mean, Danish at 469 ms, is far below 800.
Telephony reaches a compatible conclusion from the other side. ITU-T G.114 treats one-way mouth-to-ear delay under 150 ms as essentially transparent and over 400 ms as unacceptable for planning.
A turn that takes 800 ms is therefore not slightly slow. It is outside the range in which conversational gaps have been observed, which is why it reads as a machine rather than as hesitation.
BeginnerWhat is endpointing and why is it not the same as voice activity detection?
Voice activity detection answers whether a frame of audio contains speech. It is cheap, essentially solved, and runs in well under a millisecond per chunk on a CPU.
Endpointing answers a different and much harder question: has this person finished their turn, or are they pausing mid-thought. The acoustic evidence for those two cases is identical, because both are silence.
A fixed silence threshold is the naive policy, and it has exactly two behaviours: cutting people off when a hesitation runs long, and adding its full value as dead air to every completed turn. Semantic endpointing improves on it by reading the partial transcript, since the words distinguish an unfinished account number where the silence cannot.
IntermediateYour agent takes 900 ms per turn. The LLM's time to first token is 200 ms. Where do you look, and in what order?
At endpointing first, because it is usually the largest single item and it is invisible in any model metric. Published operating points put a semantic endpointer at roughly 295 ms of mean latency for a 10% false-cutoff rate and 543 ms for 5%, so the difference between two reasonable settings exceeds the entire model turn.
Second, at whether text is buffered to a sentence boundary before synthesis starts. At 40 tokens per second that silently adds around 500 ms for a 20-token sentence.
Third, at transport. Two network legs plus two jitter buffers is commonly 200 ms or more, and nobody looks at it.
Only then at the recogniser, and specifically at whether its chunk delay genuinely overlaps the user's speech rather than being paid after the last word. The general rule: the terms nobody measures are larger than the term everybody optimises.
IntermediateWhy can a streaming ASR model afford lookahead almost for free in a voice agent?
Because chunk delay overlaps the utterance instead of appending to it. A recogniser emitting with 160 ms of right context is 160 ms behind the audio at every instant, but the user is still talking, so that lag costs no turn latency at all.
Only the final chunk's delay lands after the last word, and even that runs concurrently with the endpointer deciding whether the turn is over.
NVIDIA's cache-aware streaming numbers show the accuracy side is also mild: moving chunk latency from 0.16 s to 0.56 s improved word error rate only from 7.84% to 7.22%, about 0.16 points per 100 ms.
So lookahead is the wrong knob to fight over. The right question is how much of it is still unpaid at the moment the endpointer commits.
IntermediateWhy does a voice agent use WebRTC rather than a WebSocket?
Because TCP's ordering guarantee is the wrong guarantee for live audio. A lost segment is retransmitted and every later segment queues behind it, so a single loss stalls the stream by at least a round trip.
Audio that arrives 200 ms late has missed its playout slot and is useless, so concealing 20 ms of missing sound is strictly better than waiting. WebRTC carries media over SRTP on UDP, which turns a loss into a hole rather than a stall, and pairs it with an adaptive jitter buffer that time-stretches audio to absorb arrival spread.
It also brings the rest of the real-time stack. RFC 7874 makes Opus mandatory and says endpoints should include acoustic echo cancellation, which you need for barge-in. WebSockets remain reasonable for control messages and for server-to-server hops on a reliable network.
DeepWalk through implementing barge-in properly. What breaks in the naive version?
Three problems, and the third is the one that ships broken. First, you must hear the user through your own output, which requires acoustic echo cancellation with the playback signal as reference. The hard condition is double-talk, both parties speaking at once, which is precisely what barge-in is, and what the ICASSP AEC challenges exist to measure.
Second, you must stop, and stopping is not instantaneous. Frames already handed to the playout buffer will be heard. With an 80 ms buffer and 20 ms frames, assume roughly 100 ms of speech reaches the ear after you cancel.
Third, and most often wrong, you must repair conversation state. The naive implementation writes the full generated response into history, so the model believes it delivered a sentence the user never heard and will not repeat the information.
The fix is to truncate the assistant turn at the point audio actually stopped, which requires the synthesiser to return timing marks aligning text to emitted samples. Cancellation must also reach the language model, or an orphaned generation arrives later and answers the previous question.
Finally, backchannels like "mhm" should not trigger any of this, and telling them apart needs the words rather than the energy.
DeepYou are choosing between a cascade and a speech-to-speech model in 2026. Argue both sides with specifics.
The cascade's case is operational. Every intermediate is a string, so you can log it, filter it, apply mature text guardrails, constrain tool calls, run evals and hand transcripts to compliance. Each component swaps independently, which matters in a field where the best recogniser and the best synthesiser come from different labs.
Its costs are concrete: three serial models, three services that can each fail mid-turn, and a transcript that destroys prosody, emotion and the timing of overlapped speech. A careful 2026 self-hosted build of exactly this shape measured 755 ms to first audio.
The speech-to-speech case is latency and fidelity. Moshi reports 160 ms theoretical, an 80 ms frame plus 80 ms of acoustic delay, and about 200 ms in practice on an L4. GPT-4o's system card claims a 320 ms average. Interruption is close to native because a full-duplex model never maintained a turn state to unwind.
Its cost is observability. There is no transcript unless you run a parallel recogniser, which reintroduces the component you removed, and constraining generation mid-stream is harder.
The practical tiebreaker is deployment. The same 2026 comparison found an end-to-end model at 702 ms through a hosted API and 146 seconds run locally, so self-hosted and regulated deployments still mostly pick the cascade.
DeepHow does packet loss reach your word error rate, and what would you do about it?
It reaches it through concealment rather than through silence. When a packet is lost the decoder synthesises plausible audio to bridge the gap, tuned to be inoffensive to a human ear, and your recogniser was trained on neither the artefact nor the discontinuity.
The damage is a knee rather than a slope. A study over G.722 telephony found a clean-trained recogniser holding up to around 10% loss before degrading roughly in proportion, while a model trained on network-distorted speech pushed that knee to about 15%.
The mitigations stack. At the transport layer, enable Opus in-band forward error correction, which carries a low-bitrate copy of the previous frame in the next packet and recovers isolated losses with no retransmission and no round trip. It covers only one frame back, so bursts defeat it.
Size the jitter buffer from measured arrival statistics rather than a constant, since a shallow buffer converts jitter into loss.
At the model layer, either fine-tune on audio degraded the way your production path degrades it, or put a front-end adaptation network ahead of a frozen recogniser, as Dissen and colleagues did with Whisper at Interspeech 2024.
Above all, measure word error rate on audio captured at the far end of your own transport. A benchmark number from clean files describes a system you are not running.
13Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.