Accidental change? Intentional? Or maybe someone didn't think through this change in Gemma's configuration?

They should reconsider Gemma’s TPM limit. Additionally, Google should raise the TPM ceiling to 200 or 250K tokens. RPD is a completely fine safeguard on its own. A higher TPM ceiling would restore the ability to send a single request with real context.

To reduce costs, they could also add Context Caching for the static portion of the prompt so that it doesn’t burn through compute like crazy. Both would be better solutions. 16K TPM is the same amount of memory we used to get with the GPT 3.5 16K context variant that was only available through the API.

16K Tokens are barely enough to get anything done, even for light users.

It’s affecting both casual users and developers alike according to this post: Here