Google introduces Gemini 3.8 Live for real-time voice AI applications

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. They are designed for voice applications with image, video, and text inputs.

On September 15, 2026, Google announced Gemini 3.8 Live and the Gemini 3.8 Live Extended Thinking variant, two models designed for real-time voice AI applications. They support continuous voice conversations and can also accept image, video, and text inputs in addition to audio.

The new releases expand Google’s offering for voice-agent developers. The models are available through the Gemini API and Google AI Studio. According to the company, availability in other Google products is being introduced gradually and may vary by region, subscription type, or enterprise preview.

Gemini 3.8 Live for voice AI applications

The basic Gemini 3.8 Live model is designed for low-latency, continuous interactions. It is therefore not limited to sending a recording once for transcription or generating a response. The interface is intended for applications in which the user and model converse while the AI also receives context from a camera, video, or text.

Google also lists asynchronous calls to functions and external tools for the model. This option is intended to allow the voice dialogue to continue while a task is performed in the background. In practice, this could involve scenarios where an agent needs to work with a service or function connected by the developer during a conversation.

The Extended Thinking variant adds configurable reasoning for more complex, multi-step tasks. Google thus distinguishes it from the model focused primarily on fast, continuous conversation. However, the specific reliability of this processing in production use will require independent testing, especially when external tools are used autonomously.

Multiple inputs in one interface

According to the model card, the models can work with audio, images, video, and text. Combining these formats could be important for applications that respond not only to a spoken request but also to a situation captured by a camera or content displayed to the user.

For developers, it is significant that voice dialogue, visual context, model responses, and tool calls are part of a single interface. Voice interfaces are thus moving beyond simple speech-to-text conversion toward agents intended to perform multi-step tasks during a conversation.

Google also states that audio generated by its AI products is watermarked with SynthID. This is a mechanism intended to identify content created by Google’s artificial intelligence.

Limitations listed in the documentation

In the model card, Google openly lists several limitations. The models may hallucinate, meaning they may produce incorrect or unsupported responses. Slowdowns and timeouts may occur during use. The knowledge cutoff is January 2025, so without additional current context, the model may not know about events after that date.

In enterprise deployments, these limitations are particularly relevant when a voice agent has access to external tools. Developers will need to independently verify response accuracy, behavior when a tool execution fails, and the suitability of checks before actions are executed.

Google presents its own evaluations of the models’ capabilities, but claims regarding benchmark leadership, price competitiveness, or user preference have not yet been independently verified in the available materials.

What comes next

It will be important to see when the new features appear outside developer channels, such as in the Gemini app, Search Live, and Google Workspace, and in which countries. Developers will also be watching pricing, usage limits, and conditions for Vertex AI and enterprise deployments.

Another question will be how independent comparisons assess latency, accuracy, and success on multi-step tasks against competing voice models. Real-world deployment will also show whether cases of hallucinations, failures of connected tools, or security incidents emerge.

Sources

  • Google Blog / Google DeepMind – Confirms the announcement, main features, gradual availability channels, and the stated labeling of audio using SynthID.
  • Google Blog / Google DeepMind – Confirms availability through the Gemini API and Google AI Studio, developer features, and indicative audio pricing cited by Google.
  • Google DeepMind Model Card – Confirms the models’ inputs and outputs, distribution channels, known limitations, and Google’s safety evaluation.
  • Google AI for Developers – Confirms the technical nature of the Live API as a low-latency interface for continuous voice and visual interactions.

Verified and updated: 09/16/2026 06:19

Sharing