The Voice Agent Latency Problem
Voice agents fail in production when turn latency crosses the threshold where a human perceives the system as non-interactive. Cascaded ASR-plus-LLM-plus-TTS stacks typically land between 0.4 and 2.2 seconds of round-trip latency, while speech-to-speech models are vendor-claimed to achieve roughly 0.3 to 0.8 seconds. OpenAI’s Realtime API documentation describes sub-second end-to-end latency for its speech-to-speech sessions, and AssemblyAI’s streaming endpointing docs explain how endpointing decisions add 150 to 800 milliseconds before synthesis even begins. The Moshi paper reports a theoretical interaction latency of roughly 160 milliseconds for its full-duplex architecture. For production engineers, the takeaway is that latency is the first-class constraint that drives every downstream architecture decision. OpenAI Realtime API guide · AssemblyAI streaming/endpointing · Moshi paper
Transport Choices: WebRTC, WebSocket, and Telephony
Production voice agents require a transport layer that handles audio streams with predictable latency, and the three dominant options — WebRTC, raw WebSocket, and telephony via SIP/RTP — each impose distinct constraints on codec negotiation, jitter buffering, and NAT traversal. The right choice depends on whether your users connect from a browser, a mobile SDK, or a PSTN phone number, and each path introduces a different set of failure modes that your architecture must accommodate. Below we break down the mechanics of each transport and then compare them side by side. For deeper context on transport trade-offs, see our broader coverage on the NiteAgent blog.
WebRTC: Media-First Transport
WebRTC is a media-first transport that handles capture, encoding, encryption, and network traversal as a single integrated pipeline. The browser getUserMedia API exposes echo cancellation, noise suppression, and automatic gain control as configurable constraints, letting you tune the capture pipeline before audio ever leaves the client. Opus is mandatory-to-implement for WebRTC audio per RFC 7874, and RTP (RFC 3550) defines the packet timestamping and jitter measurement that receivers use to drive an adaptive playout buffer. ICE, STUN, and TURN negotiate a media path between the client and the agent’s media server across NATs and firewalls. This integration is WebRTC’s greatest strength: you get a production-grade media stack without building those pieces yourself. MDN WebRTC API · RFC 7874 WebRTC codecs · RFC 3550 RTP · MDN WebRTC protocols
WebSocket: Message Framing Without Media
WebSocket, defined in RFC 6455, provides TCP-based message framing between client and server but deliberately leaves media handling to the application layer. There is no codec negotiation, no built-in jitter buffer, and no echo cancellation — you must implement or integrate each of those yourself. TCP head-of-line blocking means that a single lost packet stalls all subsequent frames until retransmission completes, which can freeze the audio stream at the worst moment. WebSocket is the correct choice when you are sending pre-encoded audio chunks between your own client and server, or when you are bridging to a speech-to-speech API like OpenAI’s Realtime WebSocket endpoint. It is the wrong choice when you need browser-based echo cancellation or adaptive jitter buffering without writing that code yourself. RFC 6455 WebSocket · OpenAI Realtime over WebSocket
Telephony via Twilio Media Streams
Telephony integration routes PSTN or SIP calls through a carrier’s media infrastructure before your application ever sees an audio frame. Twilio Media Streams bridges that carrier audio to a bidirectional WebSocket using audio/x-mulaw;rate=8000, which means you receive 8 kHz G.711 mu-law audio — narrowband by modern standards. SIP, defined in RFC 3261, handles call setup and signaling, while RTP carries the media between the carrier and Twilio’s infrastructure. Critically, there is no client-side acoustic echo cancellation in this path: the telephone network itself provides line-level echo suppression, but your server must handle any residual echo from the far-end. This transport is essential when your users call a phone number, but the narrowband codec and fixed sample rate constrain your ASR and TTS model choices. Twilio Media Streams · RFC 3261 SIP · Twilio Programmable Voice
Transport Comparison Table
| Dimension | WebRTC | WebSocket | Telephony (Twilio Media Streams) |
|---|---|---|---|
| Media path | SRTP (client ↔ media server) | Application-managed | Carrier RTP → WebSocket bridge |
| Codec negotiation | SDP offer/answer, Opus mandatory | None (manual) | Fixed G.711 mu-law 8 kHz |
| Jitter buffer | Built-in RTP buffer | None (TCP ordering) | Carrier-managed |
| Head-of-line blocking | No (UDP/SRTP) | Yes (TCP) | Yes (bridged WebSocket leg) |
| Echo cancellation | Browser getUserMedia AEC |
None | Network-level only |
| NAT traversal | ICE/STUN/TURN | Not needed (client-initiated TCP) | Not needed (PSTN) |
Cascaded STT + LLM + TTS Architecture
A cascaded voice agent architecture chains three separate stages — automatic speech recognition, a large language model, and text-to-speech synthesis — into a serial pipeline where the output of each stage feeds the next. This design is swappable and debuggable: you can replace your ASR provider without touching your LLM or TTS stack, and you can inspect the transcript at each boundary. The cost is serial latency plus an additional vendor network hop at each stage. Streaming ASR begins emitting partial transcripts before the utterance is complete, but the final endpointing decision still requires a silence window that AssemblyAI’s streaming documentation exposes as a tunable end-of-turn policy. LLM time-to-first-token (roughly 150 to 800 milliseconds) and TTS time-to-first-byte (50 to 300 milliseconds) are engineering estimates that depend on model size, provider, and load. Each stage is a separate billing surface: per-minute STT, per-token LLM inference, and per-character or per-second TTS. The cascade remains the default for teams that need model-level control and observability into each stage. AssemblyAI streaming/endpointing · Deepgram docs · ElevenLabs Agents
Speech-to-Speech and Full-Duplex Architectures
Speech-to-speech architectures collapse the cascaded pipeline into a single model that maps input audio directly to output audio, eliminating the intermediate text representation and the serial latency it introduces. OpenAI’s Realtime API provides a speech-to-speech session with native function calling, voice-activity detection modes, and a response.cancel command — available over WebRTC, WebSocket, and SIP transports. The API supports server_vad for standard endpointing and semantic_vad with an eagerness parameter that controls how aggressively the model commits to a response. On the research side, Kyutai’s Moshi is a full-duplex multi-stream speech-text model operating at 24 kHz audio with 12.5 Hz frame rate, meaning 80-millisecond frames. The Moshi paper reports roughly 160 milliseconds of theoretical interaction latency and roughly 200 milliseconds experimental — figures that are paper-reported and should be validated independently before committing to an architecture. Speech-to-speech models are typically priced in audio tokens rather than separate STT-plus-TTS line items, which changes the cost model significantly. OpenAI Realtime API guide · OpenAI Realtime over WebSocket · Moshi paper · Moshi GitHub
Barge-In Handling and Turn Management
Barge-in — the ability for a user to interrupt the agent mid-response — is a core production requirement because without it the agent feels rigid and unresponsive. Implementing barge-in correctly requires detecting when the user starts speaking, canceling the agent’s in-flight generation and synthesis, and beginning a new turn without perceptible gaps. The mechanics span two layers: low-level voice activity detection that identifies speech energy, and higher-level semantic models that determine whether the speech constitutes an intentional interruption. Both layers must be fast enough that the user perceives the agent as having “stopped talking” within a few hundred milliseconds.
VAD and Endpointing
Voice activity detection determines whether an audio frame contains speech, and endpointing uses VAD output to decide when an utterance is complete. Endpointing windows typically range from 150 to 800 milliseconds of silence before the system treats the turn as finished, per AssemblyAI’s documentation. LiveKit provides a built-in VAD module as part of its agents framework, and Silero VAD is a widely used open-source alternative, offering a pre-trained model that runs efficiently on CPU. The choice of endpointing threshold is a product decision: a shorter threshold makes the agent feel faster but increases false triggers on pauses, while a longer threshold reduces false triggers but adds perceived latency. Production systems typically expose this as a tunable parameter rather than hard-coding a single value. LiveKit VAD · Silero VAD · AssemblyAI streaming/endpointing
Semantic VAD and Interruption Commands
Semantic VAD goes beyond energy detection to classify whether incoming speech represents an intentional interruption or a backchannel acknowledgment like “mm-hmm.” OpenAI’s Realtime API exposes this through semantic_vad with an eagerness parameter that controls how readily the system treats user speech as an interruption. Once an interruption is detected, the response.cancel command halts the agent’s generation and in-flight synthesis, and the system begins processing the new user input. LiveKit provides a turn detector that combines VAD with a semantic model to produce turn-level events, giving the application a higher-level signal than raw VAD frames. Interruption handling is therefore not a single component but a pipeline: VAD detects energy, semantic classification confirms intent, and the cancel command propagates through every downstream stage. OpenAI Realtime API guide · LiveKit turn detector · LiveKit turns
Latency Budget and Optimization Levers
A production voice agent’s total turn latency is the sum of every stage in the pipeline, and optimizing it requires understanding where each millisecond is spent. The budget below reflects realistic ranges for a cascaded stack, and every lever you pull to reduce one stage either shifts latency to another stage or trades quality for speed. The goal is not to minimize any single component but to keep the total within the range that humans perceive as conversational.
A Practical Budget
| Stage | Typical Range |
|---|---|
| Endpointing (silence detection) | 150–800 ms |
| LLM time-to-first-token | 150–800 ms |
| TTS time-to-first-byte | 50–300 ms |
| Uplink (client to server) | 20–150 ms |
| Downlink (server to client) | 20–150 ms |
| Total cascaded | ~0.4–2.2 s |
The endpointing range comes from AssemblyAI’s streaming documentation; the LLM and TTS figures are engineering estimates that vary with model size, provider, and load, and the network legs are typical WebRTC/RTP transit estimates rather than measured values. Summed at the low and high ends, the components land near 0.4 to 2.2 seconds — the baseline that speech-to-speech architectures aim to undercut. AssemblyAI streaming/endpointing · OpenAI Realtime API guide · RFC 3550 RTP
Optimization Levers
The most impactful lever is aggressive endpointing combined with semantic VAD: by reducing the silence window from 800 milliseconds to 150 milliseconds and using semantic_vad with high eagerness, you cut the largest single variable in the budget. Streaming TTS — where synthesis begins as soon as the first LLM tokens arrive rather than waiting for the complete response — collapses the LLM-TTS serial gap. Opus encoding at adaptive bitrates minimizes uplink and downlink transit time, and jitter-buffer tuning on the RTP receiver prevents the buffer from adding unnecessary latency during stable network conditions. Finally, immediate generation cancel via response.cancel ensures that interrupted turns do not waste synthesis cycles on audio the user will never hear. Each lever has a quality trade-off, and production teams should measure the perceptual impact of each change rather than applying all levers simultaneously. OpenAI Realtime API guide · RFC 6716 Opus · RFC 3550 RTP
Cost Model and Interruption Waste
Production voice agent costs accumulate across every stage of the pipeline, and the billing model differs fundamentally between cascaded and speech-to-speech architectures. A cascaded stack meters STT by the minute, LLM inference by the token, and TTS by the character or generated-second, plus per-minute telephony charges for PSTN legs. Speech-to-speech models are priced in audio tokens that encompass both input and output, collapsing three billing surfaces into one but making per-stage cost attribution harder. The critical cost lever is interruption waste: if the TTS stage has already begun streaming audio when the user barges in, that partial synthesis is typically billed even though the client never plays it — though cancellation semantics vary by provider, and some stop billing on stream cancel. This makes barge-in rate a direct cost lever: a system that interrupts on 30 percent of turns can waste roughly 30 percent of its billed TTS output on audio that is never heard. Teams should track interruption rate alongside latency as a first-class production metric. OpenAI model pricing · Deepgram pricing · ElevenLabs pricing · Twilio Voice pricing
FAQ
Should I use WebRTC or WebSocket for voice agents?
Use WebRTC when your client is a browser or mobile app that needs built-in echo cancellation, codec negotiation, and NAT traversal. Use raw WebSocket when you are bridging pre-encoded audio between your own infrastructure or connecting to a speech-to-speech API like OpenAI’s Realtime WebSocket endpoint. MDN WebRTC API · RFC 6455 WebSocket
What voice-agent latency is acceptable in production?
Cascaded stacks land between roughly 0.4 and 2.2 seconds of round-trip latency, which is acceptable for many use cases but feels sluggish for fast-turning dialogue. Speech-to-speech architectures target 0.3 to 0.8 seconds, which approaches the latency of natural human conversation. Measure your users’ tolerance empirically rather than assuming a universal threshold. AssemblyAI streaming/endpointing · OpenAI Realtime API guide
How does barge-in actually work?
Barge-in requires voice activity detection to identify when the user starts speaking, a semantic classifier to confirm the speech is an intentional interruption, and a cancel command that halts the agent’s in-flight generation and synthesis. The system then begins processing the new user input as a fresh turn, ideally within a few hundred milliseconds of the interruption. LiveKit VAD · OpenAI Realtime API guide
Why do canceled turns still cost money?
Because TTS synthesis is billed per character or per generated-second, and most providers bill audio already synthesized and streamed before a barge-in cancel arrives — though cancellation policies vary. The STT and LLM stages may also have consumed resources before the cancel propagated, so an interrupted turn can incur partial charges across multiple billing surfaces. OpenAI model pricing · Deepgram pricing
The Bottom Line
The right voice agent architecture depends on your latency budget, your users’ connection path, and your tolerance for vendor abstraction. Start with WebRTC for browser clients, use Twilio Media Streams for PSTN, and adopt a speech-to-speech model or a tightly tuned cascade only after you have measured your actual latency budget. The cascaded stack remains the default for teams that need stage-level observability and model swapability; speech-to-speech is the right choice when sub-second latency is a hard requirement and you can accept a single-vendor pipeline. In every case, treat barge-in rate, interruption waste, and endpointing threshold as production metrics you track from day one. For a deeper comparison of vendor implementations, see our NiteAgent blog and NiteAgent arena pages.
How This Guide Was Built
This guide is based on official documentation, pricing pages, and primary specs — no hands-on testing was performed. Vendor-reported figures are attributed as such throughout, and every external link points to official documentation, a pricing page, or a primary specification. For implementation-specific guidance, consult the LiveKit Agents docs, Pipecat docs, Vapi docs, and Retell docs, each of which provides framework-level patterns that complement the architecture decisions described here.
📖 Related Reads
- ToolBrain — tool reviews, LLM comparisons, and AI workflow guides
- NoCode Insider — AI workflow automation with no-code tools, agents, and APIs
Cross-links automatically generated from NiteAgent.
← Back to all posts


