Accidental change? Intentional? Or maybe someone didn't think through this change in Gemma's configuration?


Hello. A strange change has occurred regarding Gemma, making this model unusable in a way.

It was: 15 / unlimited / 1500. Changing to the new ones shown in the screenshot is totally against usability from the USER’s perspective.

Including context data, MCP servers, and agentic behavior causes it to very quickly reach 50-100k per minute. Restricting this to the new limits totally kills this model. What’s the point of changing from 1500 to 14.4K when you can’t do absolutely anything with this model.

The system prompt, tool definitions, skills, MCP servers, etc., after all, take up 5k-10k tokens right at the start.

Could I know if this is an intentional change, or perhaps someone’s mistake? Because in the current situation, what is the point of a free 14.4K if it cannot be used in any way, as it has killed agentic behavior and tool creation for these Gemma models. What is the point of a 256k context window if you cannot even use it.

The RPM was increased, but it makes no sense because it simultaneously increases the amount of context per minute, while TPM has been completely restricted, causing us to constantly hit a brick wall. What is the point of 14.4k RPD when that implies an intention for agentic behavior, but it is impossible because we immediately hit the TPM limit again? Can anyone explain this kind of change?

Because I have the impression that a model which could be used in an application. Which I did and it worked very well. At this point it is totally killed, because as I said at the start, the model needs definitions of tools, skills, MCP, etc. Agentic behavior is completely impossible. Using playwright-cli is completely unrealistic, and the Gemma model performed quite well in it.

Was this change intentional, or did someone accidentally change TPM from unlimited to 16k?

They can’t stop making bad decisions.
I used both of the models every day and my TPM was never under 20k, totally killed my projects.

But it’s impossible to have less than 20k. The model is fast, you can connect tools to it, and it’s supposed to operate agentically. It’s normal for TPM to skyrocket. Someone must have messed something up. What is the point of 14.4k if you can’t even use the model for a regular task right now? Maybe someone can provide an explanation, because the situation is absurd.

I’m always skeptical with google, it’s impressive how they manage to … up on every little aspect of their products at some point, be it on purpose or not.
I was also having tons of API erros before this change.

I also had a ton of errors during certain hours, but I just assumed there was high traffic.

Someone screwed up. This screenshot is from TIER3, (tier2, tier1 and free) look exactly the same. Someone accidentally killed the model. Because as you can see, TPM is an important matter on paid tiers, and even there, there is a limit that completely kills the model. To my eye, it’s 95% some human error/mistake.

Wait for a day, and we will know if its a config mistake , or they just “killed” the model.

The situation is crazy because, on the one hand, Google shows solutions like the ones in the screenshot, but it is only natural that most users will use their API for this instead of buying some monster from NVIDIA.

As you can see on their profile, the recent posts were often about tok/s, and at the same time, the current situation is a total disabling of TPM. This makes absolutely no sense, seriously, what is going on?

It is like in the subscription plans of Gemini, GPT, and Claude, someone did: ‘We are increasing your weekly limits 5x! Thanks to new chips, now you can have more for the same price.’ And at the same time, making a hard limit of 1M input in a 5-hour window. This would, after all, make such tools completely unusable in those subscriptions.

reported this on issuetracker. hope it wont get ignored.

Could all this confusion be related to the fact that Google released new versions? My project has been suspended since yesterday. From what I can see, I can only say hello to the model, as the very second message already exceeds the limit.

Google Gemma on X: “We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is a breakdown of what’s being fixed and updated in this release: :thread::backhand_index_pointing_down: https://t.co/SMIGbJaUZg” / X

Got a twitter reply from someone from AI Studio team saying “We’re looking into it”. Which might suggest it’s a mistake, but still not sure looking at Google’s “bad decisions” speedrun.

At least someone replied to you. Because this decision makes absolutely no sense from an operational perspective.

Why increase RPD when the goal would be to kill the model via TPM? I don’t know, someone must have just reflexively typed in 16k or their slider went too far.

if you want, you can upvote this issue on issuetracker (by pressing a [+1] button) so it gets more priority and we could get a fix or response from the devs faster.

adding a comment about facing the same issue (and maybe about how it affects your projects) would also help a lot

https://issuetracker.google.com/issues/534781949

Done, I cast my vote. It is important that someone who understands this handles the matter. Because I remember once getting someone for my issue who completely didn’t understand the problem.

Normal usage of Gemma + tools + sub-agents is about 1-2 million tokens per minute. At the moment of creating files, opening files, searching for information. I’m worried that someone will come along and change it from 16K to 24K ;/.

Simply put, if they want to save, let them change FREE to RPD 500, but TPM must be unlimited.

thanks. i hope that they do realize how ridiculous these limits are. google should care about their developers staying loyal, not greed (even if this happened because of their greed - it is still stupid having the same limits on both free and all paid tiers). with this level of service, sooner or later, many of developers may switch to other providers.

They never treated AI Studio seriously, they never treat anything seriously.
They always find a way to make things worse.

I reported interface issues, but unfortunately, they were ignored. I remember the days of 2023 and 2024 when the interface ran fast in AIStudio. And then, in 2025, they changed something, and if system instructions had more than 20k tokens, you couldn’t even use the model anymore because the interface lagged so badly, and it is still like that today, whereas before, you could drop in 100-200k and there was no issue, it was blazing fast.

As for the current RPM (30) and TPM (16K) situation, it is so absurd that, doing some quick math, we will exceed the 16k TPM if we write quickly with the model. Simply put, the conversation context will grow, and we hit the limit.

On Google Docs pages and X, Google boasts that Gemma understands 140 languages, but so what? Since the 16k limit makes translating any documents impossible, unless they want the user to translate them a dozen sentences at a time.

They fixed the quality issue at least :joy::face_with_symbols_on_mouth::rofl::crying_cat:

“We have folks looking to get this fixed based on you flagging the other day”

One of the people from Google commented to me this way, which means they are aware of the problem.

Additionally, if they really must limit traffic (although by increasing RPD I think they wanted to promote Gemma), please do it through RPD and god forbid through TPM.

I simply prefer to have RPD: 200 but be able to fully test my application, because if TPM is cut down to a value like 100K, I cannot even assign a task with a browser, because for example playwright-cli exceptionally strains the context, and sub-agents do it completely.

my ticket on issuetracker finally got assigned