Categories
Technology

Google adds voice, avatars to Gemini

Gemini gains expressive speech, voice cloning and real-time avatars for natural AI conversations

Google is making its Gemini artificial intelligence platform more conversational with new voice and avatar capabilities that allow AI-generated characters to speak, express themselves and interact with users in real time.

The latest rollout includes Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Live with Live Avatar. Together, the updates expand Gemini beyond text and give developers new ways to build voice-based and visual AI experiences.

The new technology is aimed at making interactions with artificial intelligence sound and feel more natural. Instead of simply converting written text into speech, Google’s latest models can control the style and delivery of a voice, generate conversations between multiple speakers and, in some cases, reproduce a voice from a short recording.

The most notable addition to Google’s text-to-speech technology is voice cloning.

Gemini 3.8 Flash TTS can create a synthetic version of a person’s voice using a short audio sample. Reports on the rollout say the system can clone a voice from about 30 seconds of recorded speech.

Developers can also describe the voice they want through natural-language instructions. They can specify characteristics such as the speaker’s style, tone, accent or role, allowing them to create voices suited to different applications.

The technology could be used for digital assistants, audiobooks, games, video content and virtual characters. A company, for example, could create a branded AI assistant with a consistent voice, while content creators could generate narration without recording every line themselves.

Google has also designed the models to produce more expressive speech. The AI can adjust elements such as tone, pacing and delivery, helping conversations sound less robotic.

Gemini’s new text-to-speech capabilities are not limited to individual sentences.

The models can generate dialogue involving multiple speakers, allowing developers to create AI conversations with different voices. This could make the technology useful for podcasts, educational content, interactive stories and other applications where dialogue is central.

The system is also designed to handle conversational flow more naturally. Rather than producing isolated audio clips, developers can use the technology to create a complete exchange between AI-generated speakers.

Google is offering Gemini 3.8 Flash-Lite TTS as a lighter option for applications that need speech generation at scale. This could benefit developers building voice agents and services that need to produce large amounts of audio.

The new models are being made available through Google’s developer tools, including Google AI Studio and the Gemini API.

Google is also changing how its conversational AI looks.

With Gemini 3.8 Live with Live Avatar, users can interact with an animated AI character that speaks and responds during a live conversation.

The avatar can synchronise its mouth movements with generated speech and use facial expressions while responding. This creates a visual layer on top of Gemini’s existing conversational abilities.

The system is designed for near real-time interaction, allowing users to communicate with an AI character rather than simply receiving a text or audio response.

Google is initially focusing Live Avatar on enterprise applications. Businesses could use the technology for customer support, digital guides, training, product demonstrations and interactive kiosks.

Companies can select existing avatars or create their own digital characters with specific appearances and voices. The technology can be integrated into websites, mobile applications and other interactive experiences.

The ability to create realistic voices and faces also raises concerns around synthetic media and identity.

Voice cloning can make it difficult to distinguish between genuine recordings and AI-generated audio, particularly when a person’s voice is reproduced from a short sample.

Google says its AI systems include safeguards designed to address these risks. The company is also using SynthID, its watermarking technology for identifying AI-generated content, with the Live Avatar system.

The safeguards are particularly relevant as AI-generated voices and faces become increasingly realistic and accessible to developers.

The latest updates show Google’s broader push to make Gemini a multimodal AI platform.

The company’s AI assistant has already expanded beyond text to work with images, audio and other forms of information. The latest developments focus on making the output itself more natural, particularly when Gemini is used for conversation.

The new text-to-speech models give developers greater control over how an AI system sounds. Live Avatar adds another layer by giving that system a face and visible expressions.

For consumers, the technology could eventually lead to more natural digital assistants and interactive AI characters. For businesses, it opens up applications in customer service, training, entertainment and digital communication.

Google’s latest Gemini rollout therefore focuses not only on what artificial intelligence can say, but also on how it speaks, how it sounds and how it appears while communicating.

With voice cloning, expressive speech and animated avatars now part of the Gemini ecosystem, Google is pushing conversational AI closer to a more human-style interaction while continuing to build safeguards around increasingly realistic synthetic media.

 

Leave a Reply

Your email address will not be published. Required fields are marked *