Gemini 3.8 Live supports 97 languages, allowing users to continue conversations without waiting for tasks to complete. The model enables tasks to run in the background while maintaining real-time interaction, and supports real-time visual context, allowing users to ask questions via voice and enabling the model to understand what is currently being viewed. Gemini 3.8 Live Extended Thinking allows AI to "think while speaking," ensuring that complex tasks do not interrupt the conversation; it ranked first in independent evaluations by Artificial Analysis. Google stated that audio generated by its AI products is "detectable" and will include invisible SynthID watermarks.
Google has expanded the battlefield for its Gemini models from text and code to real-time voice.
On September 15 local time, Google announced the launch of two real-time audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as its most advanced real-time conversational models to date. These models not only facilitate more natural and fluid voice interactions but also enable tool invocation, visual information processing, and execution of complex tasks without interrupting the conversation.
Specifically, Gemini 3.8 Live focuses on large-scale, low-cost natural voice applications, supporting real-time visual understanding and multilingual dialogue. Gemini 3.8 Live Extended Thinking further incorporates deep reasoning capabilities, enabling it to "think" while speaking and complete multi-step tasks in the background as users continue their interaction. Google stated that the latter ranked first in Artificial Analysis's Speech-to-Speech Quality Index with a score of 82.6.
This signifies that Google is attempting to evolve real-time voice AI from a mere "voice-enabled chatbot" into a voice agent capable of understanding, reasoning, invoking tools, and completing tasks.
From "voice chat" to "conversing while executing tasks": A major upgrade in real-time audio capabilities
The core change in Gemini 3.8 Live is that voice interaction is no longer limited to a turn-based exchange of "you speak one sentence, AI responds with one sentence."
Google stated that the new model can asynchronously execute tool and API calls in the background while maintaining continuous conversation. Users do not need to wait for tasks to finish before continuing to speak; the model can keep interacting while related tasks run in the background.
For example, in multi-step tasks, the model can first respond to the user, then invoke relevant tools in the background, and finally integrate the results into the ongoing conversation once the tasks are completed.
This capability is particularly suitable for voice agents. Traditional voice applications often require chaining multiple components, such as speech recognition, language models, and speech synthesis, where latency in any single link can degrade user experience. In contrast, Google emphasizes that Gemini 3.8 Live is a native speech-to-speech model capable of directly performing reasoning and task execution around real-time voice interactions. According to the Google Developers Blog, this architecture reduces the complexity associated with traditional "cascade-style" voice systems.
Meanwhile, the new model supports real-time visual context. Users can not only ask questions via voice but also enable the model to understand what it is seeing, thereby achieving more context-aware voice interactions. Scenarios demonstrated by Google include real-time employee training, troubleshooting, and enabling AI to participate in chess games using visual information.
Seamless switching across 97 languages targets global voice applications.
Multilingual capability is also a major selling point of Gemini 3.8 Live.
Google stated that 3.8 Live supports 97 languages and can automatically identify and switch languages during conversations while maintaining consistent voice performance.
Logan Kilpatrick, Software Lead at Google AI Studio, also noted that the new model supports 97 languages and incorporates capabilities such as asynchronous tool calling.
For enterprises, this means developers can build voice agents covering different markets using a single model, without needing to design complex voice processing systems separately for each language.
The Google Developers Blog also highlighted the new model's ability to process alphanumeric information, including verification codes, claim numbers, and technical data. Such information is particularly important in voice scenarios for customer service, finance, healthcare, and enterprise services, as traditional voice systems are often prone to errors in recognizing numbers, letters, and specialized terminology.
Regarding pricing, Google positions Gemini 3.8 Live and 3.8 Live Extended Thinking as competitive frontier models. When developers call the Live API, the price for audio input is $0.005 per minute, and for audio output, it is $0.018 per minute.
Extended Thinking allows AI to "think while speaking," ensuring complex tasks do not interrupt the conversation.
If 3.8 Live addresses "how to make AI speak more naturally," then 3.8 Live Extended Thinking addresses "how to enable AI to perform complex tasks while speaking."
Google states that Extended Thinking enables deeper reasoning while maintaining real-time conversation. The model can even use voice prompts such as “Let me check...” to inform users that it has begun processing the task, thereby avoiding prolonged periods of silent waiting.
More importantly, the model is capable of executing multi-step tasks in the background while maintaining the foreground conversation.
Use cases demonstrated by Google include completing multi-step bookings via asynchronous function calls, generating React components in real time based on user voice inputs and sketches, and directly creating business plans and marketing kits through natural language instructions.
This further reinforces Google’s previously emphasized agentic approach, wherein the value of AI lies not merely in providing answers, but in understanding user intent, invoking tools, processing information, and ultimately completing tasks.
Ranked first in independent evaluations, Gemini 3.8 Live targets production-grade voice agents
Google also leveraged third-party evaluation results to endorse its new models.
According to data disclosed by Google, Gemini 3.8 Live with Extended Thinking ranked first in Artificial Analysis’s Speech-to-Speech Quality Index with a score of 82.6; achieved a completion rate of 68.6% in the τ-Voice agent task tests, reaching 35.1% in Sierra’s τ-Voice-banking test; and scored 97.7% in the Big Bench Audio test.

The standard version of Gemini 3.8 Live ranked second in the Speech Agent Arena, with an emphasis on cost efficiency and scalability for deployment. Google stated that this model is particularly suitable for developers and enterprises building production-grade voice agents.
Google has not positioned these two models simply as consumer chat products, but is targeting both the developer and enterprise markets.
Currently, both models are available to developers via the Gemini API and Google AI Studio; enterprise users can participate in private previews through Gemini Enterprise. General users can access related capabilities through Search Live and Gemini Live.
Meanwhile, Gemini 3.8 Live Extended Thinking has also been integrated into Google Workspace, making it available for products such as Docs, Gmail, and Keep.
AI voice technology is being fully deployed, but Google also emphasizes "detectability."
As generative AI increasingly enters scenarios such as voice calls, customer service, search, and office workflows, the transparency of AI-generated audio has become a key focus of Google’s latest release.
Google stated that all audio generated by its AI products will include SynthID invisible watermarks. These watermarks are directly embedded in the audio output to help identify AI-generated content and reduce the risk of generative AI being used to disseminate misinformation.
This also means that while Google is promoting the adoption of voice AI in broader production environments, it is simultaneously addressing issues such as whether users know they are interacting with AI and whether the audio they hear is AI-generated.
From Search to Workspace, and further to developer APIs and enterprise agent platforms, the scope of Gemini 3.8 Live has clearly expanded beyond that of a simple voice assistant.
For Google, the significance of this upgrade may lie not merely in launching two new audio models, but in further enhancing Gemini’s multimodal capabilities across text, images, video, code, and real-time voice, while directly linking voice interactions to tool invocation and agent execution.
In other words, the next phase of AI interaction that Google is betting on may no longer involve "opening a chat window to ask questions," but rather speaking as one would with a human colleague, while allowing the AI to handle tasks in the background.
Editor/Stephen