Hello,
I am currently planning a production application using the Gemini API. I am not requesting a rate-limit increase at this time; I only want to understand the limits and scalability before building my application.
I have the following questions:
-
What is the exact RPM limit for Gemini models on Tier 3?
-
What is the exact TPM limit for Tier 3?
-
When RPD is shown as “Unlimited”, does that mean there is no daily request limit?
-
If my application eventually receives 10,000+ requests at the same time, can Gemini API support this type of workload?
-
If the current RPM/TPM limits are not sufficient, can I request higher limits later when I move to production?
-
Is there a maximum RPM/TPM limit that Google can provide for a large-scale production application?
-
For an application that may eventually have 10,000+ concurrent users, what architecture does Google recommend for handling the requests reliably?
I am only looking for clarification at this stage. I will consider purchasing/upgrading the appropriate API tier later when the application goes into production.
Thank you.