PolyAI Dialog-RSN-1 launches for lower-latency voice agents
PolyAI released Dialog-RSN-1, a voice model that hears audio directly and targets faster, less awkward AI phone calls.
By Dominic Okoye · Staff Writer
· 3 min read
PolyAI Dialog-RSN-1 is the company’s new real-time voice conversation model for AI-driven phone calls, built to reduce lag and make automated agents respond more like a human operator. PolyAI did not disclose pricing, customer names or financial terms for the launch, but the product is aimed at a concrete pain point in call-center AI: latency and poor handling of spoken context.
The company said Dialog-RSN-1 can process audio directly inside the model rather than relying on a front-end speech recognition system to turn speech into text before a large language model generates a reply. Speech output remains separate, with a text-to-speech layer handling the voice a caller hears.
What is PolyAI Dialog-RSN-1?
Dialog-RSN-1 is a voice dialog AI model that perceives the caller’s audio itself, including cues that are often stripped out of a transcript. In PolyAI’s telling, that lets the model account for tone, cadence, accent, mood and background noise when deciding how to respond.
That architecture is meant to address common failure modes in automated calls. PolyAI said the model can use context to understand words that might be misread in a transcript, such as a caller spelling a name aloud or saying a brand name that resembles a common word. The company also said direct audio processing helps the system avoid cutting off the caller because it can better infer turn-taking from the rhythm of speech.
How fast does PolyAI say the model is?
PolyAI said Dialog-RSN-1 has latency between 280 and 500 milliseconds, and said the system can respond in under 300 milliseconds. The company compared that with OpenAI’s GPT-realtime-2.1, which it said often responds between 860 and 1,900 milliseconds.
The benchmark claim matters because conversational timing is one of the areas where voice agents still feel mechanical. Human speakers typically leave roughly 200 to 300 milliseconds between turns, according to the figures cited by PolyAI, and longer delays can make it harder for callers to interrupt, correct a misunderstanding or move the conversation on.
The split between Dialog-RSN-1 and text-to-speech is also a product decision, not just an engineering detail. PolyAI said keeping output voice generation separate gives customers more control over the voice, accent, intonation, emotional quality and cadence without retraining the underlying conversation model.
PolyAI contrasted that approach with voice systems such as GPT realtime and Gemini Live, where the output voice is part of the model. The company said a separate TTS layer can make voice customization easier to manage at scale and may let companies control speech generation costs by choosing their own deployment setup or hardware.
How PolyAI built the model
PolyAI said it created Dialog-RSN-1 by post-training open-weight multimodal models with supervised fine-tuning and reinforcement tuning, using in-house training data. The company said its pipeline is broadly model-agnostic and that it evaluated Gemma, GPT-OSS, Qwen and Mistral.
The stated engineering target was sub-300 millisecond delay when served on A100 graphics processing units, using either an 8 billion-parameter dense model or a 30 billion-parameter sparse model. PolyAI did not provide third-party validation of the latency figures, deployment costs or comparative accuracy data, so buyers will still need to test the system under their own call volumes, accents and telephony conditions.
This story draws on original reporting from SiliconANGLE.