Gemini 3.8 Live keeps talking while its tools work in the background
Google’s new voice models can continue a conversation while calling APIs, with an Extended Thinking version that narrates progress during longer tasks.
OddBrief EditorialAI-assisted, human-reviewed
AIKey facts
- Models
- Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking
- Capability
- Conversation can continue during background tool calls
- Languages
- 97 supported within a conversation
- Reported score
- 82.6 on Google’s Speech-to-Speech Quality Index
- Caveat
- Benchmark figures are reported by Google
Google's latest Gemini voice model can call a tool without putting the conversation on hold. Gemini 3.8 Live is designed to keep listening and speaking while an API or external service works in the background, a change aimed at making voice agents feel less like a sequence of loading screens.
Google introduced two versions: Gemini 3.8 Live, optimized for scale and cost, and Gemini 3.8 Live Extended Thinking, aimed at tasks that require more reasoning. The company published the announcement on September 15 and updated it on September 17.
Tools move off the critical path
Voice agents often pause after a user asks for information that requires a search, database call or action in another service. Google says the new Live model can continue the exchange during that wait. An assistant might clarify a travel preference while checking availability, rather than going silent until the tool returns.
The Extended Thinking version takes the idea further. It can reason and speak at the same time, providing progress updates while working through a complex task. That behavior could make a long operation easier to follow, though it also raises a design challenge: useful transparency can quickly become distracting narration.
Both models support 97 languages in a single conversation, according to Google. They are being distributed through the Gemini API as well as Google's consumer and workplace products, including Search, Workspace and the Gemini app.
Google leads with benchmark gains
Google reports a score of 82.6 on its Speech-to-Speech Quality Index and 68.6 percent on the τ-Voice benchmark. It also cites 35.1 percent on a banking scenario within Sierra's τ-Voice tests and 97.7 percent on Big Bench Audio.
Those results suggest progress in natural speech and task completion, but they remain benchmark measurements presented by the model maker. Real-world performance will depend on accents, noisy environments, tool reliability and how well a product recovers from interruptions or mistaken actions.
The announcement does not provide a single failure rate for background tool use or explain how often the model speaks before it has enough evidence. Developers will need to set clear rules for when the agent should keep talking, wait or ask permission.
Voice agents become concurrent systems
The deeper shift is architectural. A conversational model that speaks, listens, reasons and calls tools concurrently is no longer a simple request-response interface. It behaves more like an operating system coordinating several processes in real time.
That can make interactions faster and more human. It can also make errors harder to track because the model may be talking about one step while another is still running. Product teams will need logs, interruption controls and strong confirmation gates for actions with consequences.
Gemini 3.8 Live turns latency into conversation. Whether that feels helpful will depend less on how much the model says than on whether it knows when silence is the better answer.
Sources
- Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended ThinkingGoogleprimary source


