Google’s Gemini 3.8 Live Updates
Google has launched Gemini 3.8 Live, enhancing its capabilities to allow continuous responses even while a tool completes its work. The Extended Thinking version enables developers to fine-tune reasoning processes, offering improved interaction in applications. This advancement is a part of Google’s ongoing commitment to enhancing artificial intelligence and interactive tools for developers.
NVIDIA’s NemotronLabs VoiceChat Success
NVIDIA has unveiled research on NemotronLabs VoiceChat, which features a novel combination of full-duplex speech and tool calling ability, showcasing an impressive 82.5% tool-selection F1 score on the Full-Duplex-Bench 3.0. This development represents significant progress in voice AI technology, enhancing the interaction capabilities of voice-driven applications and establishing NVIDIA’s authority in the domain.
Meta’s Muse Expansion
Meta is expanding its business calling service through Muse, now accessible to more beta users within the United States. This rollout prioritizes individuals who previously expressed interest in the service, aiming to improve communication efficiency in business settings and catering to the evolving demands of enterprise communication.
Nuance Labs and Investment Boon
Nuance Labs has announced a successful $50 million Series A funding round led by Lightspeed. This funding aims to develop a new model capable of processing audiovisual input and output simultaneously. The research preview of this model is anticipated later in the year, indicating an exciting advancement in the intersection of voice technologies and audiovisual media.
Speech and Voice AI Innovations
Several companies are making strides in speech recognition and voice AI:
– Instinct’s Concierge is currently assisting early users with tasks like making phone reservations and managing billing inquiries.
– Speechmatics has released Agent STT, which uses the Linden 1 model to improve recognition of names and account details.
– Deepgram’s new Nova-3 Pharma model targets the healthcare sector, specializing in pharmaceutical terms and drug names, enhancing clarity in medical communications.
– DeepL now offers real-time translation for meetings, maintaining the unique voice of each speaker across 12 languages, signaling a robust integration of voice technology in professional environments.
These advancements collectively highlight the rapid development and specificity of voice AI applications, catering to various industries and enhancing user experiences.

