I am reporting a serious audio quality regression in Gemini 2.5 Pro TTS.
I have been using this model extensively for scripted narration and dialogue generation, and the current behavior is clearly different from how the same model performed before.
The problem is very specific: Gemini 2.5 Pro TTS now frequently produces short audio discontinuities inside words and phonemes. Parts of sounds are abruptly cut, interrupted, or corrupted, creating audible micro-skips in otherwise continuous speech.
This is not a subtle difference in voice quality, style, or pronunciation. It is an actual audio generation error.
Most importantly, this did not happen before.
Gemini 2.5 Pro TTS was previously capable of generating clean, smooth audio from the same type of scripts and prompts. Something appears to have changed recently, because the model now produces these artifacts frequently enough to make it unreliable for production use.
The issue is also very easy to reproduce.
Generate approximately one minute of normal narration or dialogue with Gemini 2.5 Pro TTS and listen carefully to the result. In my experience, the micro-cuts occur frequently enough that they are usually easy to hear. Repeating the same generation can cause the artifacts to appear in different positions, which strongly suggests that the problem is not tied to a specific word, sentence, or prompt.
I also tested Gemini 3.8 Flash TTS.
The specific micro-cut problem does not appear there in my testing. However, Gemini 3.8 Flash TTS is not an equivalent replacement for Gemini 2.5 Pro TTS.
When Gemini 2.5 Pro works correctly, I find its speech noticeably cleaner, softer, and more controlled. It also responds much better to voice-direction prompts.
Gemini 3.8 Flash TTS behaves differently. In particular, when I request slow and calm speech, the voice can sometimes become rough or hoarse, and metallic-sounding artifacts can appear. Its response to style prompts is also significantly less predictable in my testing.
So switching to Gemini 3.8 does not solve the underlying problem with Gemini 2.5 Pro. It simply means switching to another model with different voice characteristics and different limitations.
The key point is:
Gemini 2.5 Pro TTS was already able to produce clean audio without these micro-cuts. This is therefore a regression of existing functionality, not a request for a new feature or an improvement beyond the model’s previous capabilities.
Something has changed in the current Gemini 2.5 Pro TTS generation pipeline, and the model is now producing errors that were not present before.
Could the team please investigate whether there have been recent changes affecting any of the following?
- TTS inference or serving pipeline
- audio reconstruction
- streaming or chunk handling
- audio encoding
- post-processing
- internal generation parameters
I would specifically like to know:
- Has this regression been reproduced internally?
- Is the team currently investigating it?
- Were there any recent changes to the Gemini 2.5 Pro TTS audio pipeline?
- Is Gemini 2.5 Pro TTS expected to be restored to its previous audio quality?
I can provide example audio files if needed, although the issue should also be reproducible directly by generating approximately one minute of ordinary speech with the current version of Gemini 2.5 Pro TTS.
Before this regression, Gemini 2.5 Pro TTS was significantly better suited to my use case than Gemini 3.8 Flash TTS. At the moment, however, these repeated audio cuts make it unreliable for professional use.
I would appreciate it if the issue could be investigated as a regression in Gemini 2.5 Pro TTS itself, rather than treated simply as a reason to switch to Gemini 3.8.