A recent change to Gemini Omni Flash has significantly disrupted my UGC video workflow in Flow.
Previously, I could give the Flow agent one reference image of my avatar and generate multiple clips for the same UGC-style video. Once the first clip established the character’s voice, subsequent clips would naturally maintain that same recognizable voice with very good consistency.
This was especially valuable for UGC because the naturally generated voice responded to the context of the prompt. I could create a character speaking inside a car, classroom, bedroom, outdoors, etc., and get believable environmental acoustics, microphone distance, personality, cadence, and natural vocal imperfections.
Now, when generating those same multi-clip videos conversationally, each clip frequently receives a completely different voice—as if every clip is an unrelated generation rather than another segment of the same video.
The current Voice/Character Ingredients do not fully replace the previous workflow for UGC. Preset voices can sound overly clean, corporate, or studio-recorded, and they may not naturally match the appearance, personality, environment, microphone characteristics, and acoustics of the generated character.
Please restore a way to preserve a naturally generated voice across multiple clips without requiring a preset Voice Ingredient.
A simple solution could be a “Keep voice consistent across clips” toggle. Once the user approves an anchor clip, Flow could preserve that character’s naturally generated voice identity for subsequent clips in the same video while still allowing the model to adapt the performance and recording acoustics to each scene.
Even better, allow us to explicitly select an existing generated clip and choose something like “Use this voice for subsequent clips.” This would let creators lock the voice that naturally fits their avatar instead of being forced to choose a preset voice that may not match.
The important distinction is that we want to lock the speaker’s identity, not lock the entire audio treatment. The same person should still be able to sound naturally recorded in a car, classroom, bedroom, outdoors, or on a phone microphone depending on the prompt.
For UGC creators, this would restore the authentic “shot on an iPhone” quality that made the previous conversational workflow so useful while still giving users explicit control over voice continuity.