Hi,
We use gemini-3.8-flash-tts (streamGenerateContent, responseModalities AUDIO, prebuilt voices, 24 kHz PCM16 mono) to voice museum audio-guide narrations of about 50 seconds, in 12 languages.
Some generations contain a low, dull, metallic hum/rumble underneath the voice for the whole clip. It sounds like room ambience or a resonance and is clearly audible on headphones. Most generations are clean. We saw the older threads about noise on the 2.5 TTS models: it looks like the 3.8 model is affected too.
What we observed
- Both clear cases came from the same style. The text is a deliberately pompous art-critic monologue, and the input starts with this style tag:
[haughty, snobbish art critic, slow condescending drawl, theatrical pauses] -
- It happened with two different voices: Despina and Algieba, both in French.
-
- The same pipeline gives clean audio with other styles and voices (for example Achird, French, plain narration with no style tag).
-
- Regenerating the same text sometimes produces a clean result, so the problem is intermittent.
-
- Each narration is sent in a single request (no chunking).
The forum only accepts images, so we could not attach the audio. We can share the three samples (two with the hum, one clean reference) and the exact request payloads with the team.
- Each narration is sent in a single request (no chunking).
Questions
(1) Is this a known artifact of the model, for example ambience generated from expressive style tags such as “theatrical”?
(2) Is there a recommended way to request a dry, clean studio sound, through a setting or a phrasing in the style tag?
(3) Would a different TTS model avoid it?
Thanks!