PolyAI launches new real-time voice conversation model to make AI-driven calls more human
Voice assistant and conversational AI agent developer PolyAI Ltd. today announced the release of Dialog-RSN-1, a voice dialog artificial intelligence model capable of directly perceiving and responding to audio.
Voice AI companies have been working to reduce delay and increase accuracy in voice AI agents, enabling them to sound more human on phone calls. In many cases, the chain of software needed to make this happen usually starts with speech recognition and processing, then the large language model for text generation, and finally speech output.
The company’s new LLM does speech recognition and processing all within the model, but speech output is split out. This allows the AI model to completely “hear” and understand the voice conversation all at once, without another model in front of it that cannot pick up tone, cadence, or understand the full context of the conversation.
According to PolyAI, this new model provides an experience closer to another human listening to the call. Having the audio directly parsed by the AI, it can pick up on non-speech cues that might change the path of the conversation.










