When invoking the gemini-3.8-flash-tts speech model for audio generation, requests fail intermittently with an HTTP 429 Too Many Requests error. The accompanying error message reads:
“Google voiceover was rejected (HTTP 429). Google did not supply a recognized quota category. Check this speech model’s usage and limits in Google AI Studio. Completed recordings are saved.”
Observed Behavior:
-
The error suggests checking usage limits in Google AI Studio, but the Google AI Studio Rate Limits dashboard does not explicitly map a specific or recognized quota line item for
gemini-3.8-flash-ttsvoiceover generation. -
Standard metrics (RPM, TPM, RPD) for Tier 2 show peak usage well below project capacity ceilings, but the TTS endpoint still rejects requests with code 429.
-
The API payload returns standard HTTP 429 without referencing a clear quota category identifier (e.g.,
RESOURCE_EXHAUSTED/ specific limit type).
Environment & Setup:
-
Model:
gemini-3.8-flash-tts -
Account Tier: Tier 2 (Paid Google AI Studio Project)
-
API / Endpoint: Voiceover / Text-to-Speech Generation
Questions / Request:
-
Is
gemini-3.8-flash-ttssubject to an unlisted audio token rate limit or a separate regional concurrency pool limit? -
Why is no recognized quota category returned in the HTTP 429 response body?
-
How can we properly track and manage limits for speech/voiceover models in the AI Studio Dashboard to avoid unexpected request rejections?

