CRITICAL REGRESSION: Gemini Omni 1.1 Flash Update Destroyed Conversational Voice Continuity for UGC Creators

A recent change to Gemini Omni Flash has significantly disrupted my UGC video workflow in Flow.

Previously, I could give the Flow agent one reference image of my avatar and generate multiple clips for the same UGC-style video. Once the first clip established the character’s voice, subsequent clips would naturally maintain that same recognizable voice with very good consistency.

This was especially valuable for UGC because the naturally generated voice responded to the context of the prompt. I could create a character speaking inside a car, classroom, bedroom, outdoors, etc., and get believable environmental acoustics, microphone distance, personality, cadence, and natural vocal imperfections.

Now, when generating those same multi-clip videos conversationally, each clip frequently receives a completely different voice—as if every clip is an unrelated generation rather than another segment of the same video.

The current Voice/Character Ingredients do not fully replace the previous workflow for UGC. Preset voices can sound overly clean, corporate, or studio-recorded, and they may not naturally match the appearance, personality, environment, microphone characteristics, and acoustics of the generated character.

Please restore a way to preserve a naturally generated voice across multiple clips without requiring a preset Voice Ingredient.

A simple solution could be a “Keep voice consistent across clips” toggle. Once the user approves an anchor clip, Flow could preserve that character’s naturally generated voice identity for subsequent clips in the same video while still allowing the model to adapt the performance and recording acoustics to each scene.

Even better, allow us to explicitly select an existing generated clip and choose something like “Use this voice for subsequent clips.” This would let creators lock the voice that naturally fits their avatar instead of being forced to choose a preset voice that may not match.

The important distinction is that we want to lock the speaker’s identity, not lock the entire audio treatment. The same person should still be able to sound naturally recorded in a car, classroom, bedroom, outdoors, or on a phone microphone depending on the prompt.

For UGC creators, this would restore the authentic “shot on an iPhone” quality that made the previous conversational workflow so useful while still giving users explicit control over voice continuity.

Hello @eirenrocks ,

This community dedicatedly supports the Gemini API and Google AI Studio.If you have issues with Google Flow, please visit the Google Flow App, click on the three dots in the menu, and select Feedback to report your issue.

This is a problem we have been trying to work around as well for consistent voice across 8 or 10 second clips in that we are providing a starting frame of a talking dog from the last frame of the prior clip. Overall works fairly well except when it doesn’t. Recently stumbled across the @voice tag when I asked the AI agent in flow what prompt did it use. We are using the Omni API to generate the video shots. Having the @voice prompt embedded in the script I think is the better way to go related to making it easy. Ideally we would like to be able to pass in an audio file which appears to be an option but currently not supported in Flow or the API. Any documentation for the API? It looks like the path forward is to pick from a large selection of voices as a starting point with the ability to prompt them to what you want and then save that voice. Couldn’t get that to work in Flow related to saving the voice after testing and then not sure how that would extend to the API.