Voice, Video, and Reasoning in One: Google’s Gemini 3.8 Live with Extended Thinking
Google released two new speech-to-speech voice models, joining a suddenly crowded field of models built to power voice agents.
Google released two new speech-to-speech voice models, joining a suddenly crowded field of models built to power voice agents.
What’s new: Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech-to-speech models that can serve as real-time voice agents. The Gemini 3.8 Live model is built for scale and cost-efficiency while 3.8 Live Extended Thinking is intended for high-complexity tasks. Both models are designed to reduce response latency during live dialogue.
How it works: Google released few details about the model’s architecture and training data, but did say that the models are based on Gemini 3 Pro. Gemini 3.8 Live can take in image and video input, supports 97 languages, and can execute tools and API calls in the background while continuing a conversation. Live Extended Thinking can also reason and think simultaneously. The models’ knowledge cutoff is January 2025.
- In speech-to-speech benchmarks from Artificial Analysis, a company that evaluates AI models, Gemini 3.8 Live Extended Thinking generally ranked above its less powerful sibling, 3.8 Live, in various tests, but users preferred 3.8 Live for some tasks. The Live Extended Thinking model is ranked first in the Speech to Speech Index benchmark, with a score of 82.6%. It was fourth in AA’s Speech Reasoning benchmark – which evaluates an audio model’s ability to answer reasoning-based questions – behind StepAudio 3 Realtime and two Qwen models. It ranked in the top spot in the Tau Voice benchmark, with a score of 68.6 percent. In comparison, Gemini 3.8 Live is somewhat lower in the Speech to Speech Index, at 76 percent – the fifth spot – and significantly lower in the Agentic Performance Benchmark, at 30.1 percent. However, 3.8 Live ranked second in the Speech Agent Arena Leaderboard, where people hold blind live voice conversations and pick the model they prefer for tasks like booking a dentist appointment. It also performed better than Live Extended Thinking in certain benchmarks that measured conversational dynamics and the percentage of correctly completed tasks.
- The models are available to developers through the Gemini API and app, Google Cloud Vertex / Vertex AI, and Google AI Studio. Gemini 3.8 Live is also used in Search Live, and Gemini 3.8 Live Extended Thinking powers voice experiences in the Gemini app and Google Workspace products including Docs and Gmail.
- Google says that all audio generated by its AI products is watermarked with SynthID, a technique for detecting AI-generated content.
- Both models are available for free, with some limitations, in certain products like the Gemini app, Google Search, and the Gemini API. Paid tier queries are not used to improve the model, while the free tier is. The paid tier is priced per million tokens; the input price is $0.75 for text, $3.00 ($0.005/min) for audio, and $1.00 (or $0.002/min) for images or video. For output, the models are priced $4.50 per million tokens of text and $12.00 or $0.018/min of audio. Artificial Analysis found that Gemini 3.8 Live costs $0.84 per hour of input audio — the cheapest model in the Index — and a significantly lower cost than Live Extended Thinking at $3.50 per hour.
Behind the news: Traditionally, voice agents have relied on a pipeline that converts speech to text, sends it to a reasoning model, and converts the response back to speech. This approach has historically been seen as more accurate and easier for developers to control because LLMs generally reason in text. However, each step can add latency, which detracts from the user experience. OpenAI recently released a speech-to-speech model called GPT-Live-1, a voice model developers can use to build voice-enabled apps. This model can listen and speak at the same time, allowing it to handle interruptions and respond more naturally, while a separate reasoning model runs in the background.
Why it matters: It’s correct to call Gemini 3.8 Live a voice model since its primary language input is voice rather than text, but its ability to understand and reason over image and video input makes it more versatile. Putting voice and video input together is particularly powerful for helping users use AI in real-time to deal with anything on screen: web interfaces, games, multi-application workflows, etc. This differentiates it from GPT-Live-1 and other models that work strictly with voice and audio.
We’re thinking: Voice-driven agents are particularly exciting for developers because they offer a new real-time interface for building applications. Speech is a natural input format that most people can use regardless of their experience using a computer. It’s particularly useful for mobile and automotive interfaces where text input is less convenient. It’s up to developers to create applications that democratize users’ access to AI.