Wednesday, September 16, 2026
Google Launches Gemini 3.8 Live and Extended Thinking Voice Models

Google Launches Gemini 3.8 Live and Extended Thinking Voice Models



Google launches Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two native speech-to-speech models designed to converse while executing tools and reasoning through background tasks. Both are available to developers through the Gemini API and Google AI Studio.

 

The September 15 release separates fast, high-volume voice interactions from more complex workflows that need planning and several tools. Google is also distributing the models through Search Live, Gemini Live and selected Workspace products, while enterprise access begins in private preview.

 

The Gemini 3.8 Live launch introduces five concrete capabilities:

  • Background API calls while conversation continues
  • Near-real-time visual understanding
  • Automatic switching across 97 languages
  • Configurable reasoning for complex voice tasks
  • SynthID watermarking on generated audio

 

More on This Story

 

Google Launches Gemini 3.8 Live With Two Operating Profiles

Gemini 3.8 Live is optimized for low-latency conversations, direct commands and tools that respond quickly. Google positions it for customer-service triage, language practice, voice search, interactive storytelling and device control where natural turn-taking matters more than prolonged deliberation.

 

Gemini 3.8 Live Extended Thinking targets tasks that require several reasoning steps or slower external systems. Examples in Google's developer documentation include technical diagnostics, travel planning across multiple booking services, code tutoring and workflows that would otherwise leave an awkward pause while a tool finishes.

 

The two models use separate endpoints. Standard Live has a fixed reasoning profile and supports blocking or non-blocking tools. Extended Thinking offers low, medium and high reasoning settings, requires asynchronous tool declarations and uses an interaction-status signal to indicate whether a longer task is still active.

 

That architecture lets Extended Thinking speak short acknowledgments and narrate progress while it continues processing. A booking assistant, for example, can tell a caller that it is checking options while parallel functions search flights, query hotels and compare results.

 

Background Tool Calls Keep Voice Conversations Moving

The central product change is asynchronous function calling. Gemini can launch an API or software tool in the background, continue streaming audio and incorporate the result when it arrives. Conventional voice agents often stop speaking while their application waits for an external service.

 

Both models also accept visual context, enabling an agent to discuss what a camera or screen is showing. Google demonstrated the system interpreting live scenes, troubleshooting with visual input and turning spoken feedback plus rough sketches into working React components.

 

Google says the models can automatically detect and move among 97 supported languages during one conversation. The developer release also emphasizes alphanumeric precision for information such as confirmation codes, claim numbers and technical identifiers, where a single recognition error can invalidate an entire transaction.

 

Incremental content updates allow structured data to merge with a continuing audio response. That feature is intended for situations where new information arrives during a call, rather than forcing the agent to discard its reply and restart after every tool result.

 

Independent Tests Put Extended Thinking at the Top Overall

Google reports an 82.6 score for Extended Thinking on the Artificial Analysis Speech-to-Speech Index, 68.6% on the agentic τ-Voice benchmark and 97.7% on Big Bench Audio. The composite index weighs reasoning, agent performance, human preference and task completion equally.

 

Artificial Analysis' current leaderboard confirms Extended Thinking holds the highest composite score among models with complete results. Standard Gemini 3.8 Live scores 76.0 overall and has stronger conversational dynamics, but substantially lower agentic performance than the Extended Thinking configuration.

 

The independent table also prevents an overly broad conclusion. Other models lead individual measures: StepAudio 3 Realtime ranks higher for speech reasoning, while Qwen Audio 3.0 Realtime Plus leads conversational dynamics and has much lower published input-audio pricing.

 

Benchmarks are controlled tests rather than guarantees for every call center, tutor or assistant. Accent variation, noisy audio, tool reliability, prompt design and the time needed by external services can change production performance even when the underlying model ranks well.

 

Gemini API Pricing Starts at Fractions of a Cent Per Minute

Google's developer announcement estimates audio input at $0.005 per minute and audio output at $0.018 per minute. The footnote derives those figures from token prices of $3 per million input tokens and $12 per million output tokens.

 

Actual costs will depend on conversation length, generated speech, reasoning, tool activity and application design. A cheap minute of audio can still support an expensive workflow if the agent invokes paid databases, booking systems or other models behind the scenes.

 

Developers can access both endpoints through the Live API and AI Studio. Google lists Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents among integration partners that handle real-time media infrastructure.

 

For consumers, standard 3.8 Live is rolling out in Search Live. Extended Thinking is entering Gemini Live and is available to Google AI Pro and Ultra subscribers in Workspace Docs, while Gmail and Keep availability extends across Google AI subscriptions.

 

Gemini Enterprise access remains a private preview rather than general availability. Google says support for its customer-experience enterprise product is coming later, so organizations should not treat every announced surface as immediately open to all accounts.

 

All generated audio carries Google's imperceptible SynthID watermark. That provides a technical provenance signal for detecting AI-created or edited speech, although its effectiveness depends on compatible detection tools and whether the audio survives later compression or modification.

 

The launch moves voice agents closer to continuous collaboration rather than a strict ask-wait-answer loop. Its practical test will be whether asynchronous tools and live reasoning remain accurate during messy real-world conversations, where callers interrupt, change languages and expect completed actions rather than a polished voice alone.

THEFLGHT
author

THEFLGHT

Elevating narratives from the heart of London's intellectual epicentre.

0 Comments:

Leave a Reply

AI Agent Data Breach: Spain Discloses Its First Report
Google Launches Gemini 3.8 Live and Extended Thinking Voice Models
Meta Launches Meta One AI Subscriptions Across Instagram, WhatsApp and Facebook
Axelera Launches Europa AI Chip for Dell and Supermicro Systems
Einride Autonomous Truck Launches Into Daily Lidl Service in Germany
DeepSeek IPO Plan Taps Yan Wentao as First CFO