It’s cool that something has started to happen, I would like to resume my project over the weekend and develop it. Unfortunately, I have some concerns. I use the FREE tier myself, because 1500 RPD was completely enough for my testing + usage. I’m worried that someone will simply drop TPM to 100k on FREE and in reality, it won’t change anything for me compared to 16K, and I’ll have to abandon Gemma.
30 / 100K / 14.4K - flawed from a user perspective.
If someone from Google is reading this topic, even 15 / unlimited / 200 would be much better for the user.
They should reconsider Gemma’s TPM limit. Additionally, Google should raise the TPM ceiling to 200 or 250K tokens. RPD is a completely fine safeguard on its own. A higher TPM ceiling would restore the ability to send a single request with real context.
To reduce costs, they could also add Context Caching for the static portion of the prompt so that it doesn’t burn through compute like crazy. Both would be better solutions. 16K TPM is the same amount of memory we used to get with the GPT 3.5 16K context variant that was only available through the API.
16K Tokens are barely enough to get anything done, even for light users.
It’s affecting both casual users and developers alike according to this post: Here
TPM was previously set to Gemma ‘unlimited’, because if someone uses sub-agents + browser + mcp, it jumps straight to a million even. RPD as a parameter is enough to manage consumption.
The main context with tools on my end, shared for all agents, is about 10k tokens. When we take skills for playwright-cli, notebooklm, and other tasks, it ramps up to 1 million per minute after a short while.
Without ‘unlimited’, Gemma is merely a chatbot; with unlimited, it can be a good tool.
They replied on issuetracker.
Status: Won’t Fix (Infeasible).
Hi @da…@gmail.com,
Thank you for reaching out to us.
We have reviewed the reports regarding Google AI Studio / Gemini API (Gemma 4 Models). While we appreciate the feedback, we would like to clarify that Google AI Studio is currently outside the scope of the Issue Tracker, which is for reporting bugs and requesting features on Google Cloud products. You can learn more here.
To ensure your report and feedback reach the correct product team, we recommend to use the Google AI Studio interface, go to Settings > Send feedback. This is the direct channel to the product team for bug reports and feature requests.
For technical questions visit the Google AI Developers Community or consult the official documentation for technical discussions and troubleshooting.
We will be closing this case. We recommend all interested users follow the feedback path mentioned above to ensure visibility with the appropriate support.
Thank you for your understanding.
Google Cloud Platform
Unfortunately. The only thing left is that two people from Google posted something on X, maybe it will work, otherwise we will have to wait for Gemma-5 in 2027, maybe then there will be a new TPM.
I’m afraid the matter might be dead. I don’t know, doesn’t Google care about Gemma? Gemma 31 is, unfortunately, a model that I stand no chance of hosting locally, especially with sub-agents and fast agentic performance, it’s unreal for me, and probably for many others as well.
On top of that, damn, I even dreamed that the TPM was changed from 16K to… to what? To 32! But not 32K, just straight up 32 tokens, to the point where I woke up pissed off hahahha.
It really hurts that I can’t continue my project. What do I need this 14.4K for? 500 would be enough for me, as long as the model can reach its full potential thanks to unlimited TPM.
Ugh!
This change in Gemma’s configuration is a disaster. Going from 1500 to 14,400, but making it unusable, what’s the point? Tokens run out incredibly fast. It’s frustrating!
Is it a bug or intentional? 
GITHUB APP NAME: @ANDRICK-IA-GITHUB-ENTERPRISES ®
PUBLIC LINK:
https://github.com/apps/andrick-ia-github-enterprises
Maybe someone didn’t think things through. Because doing such a thing intentionally would be mean or trolling. So I don’t even know what to think. Since then, I’ve used Gemma barely 4 times. Well, because what am I supposed to do with it, when even the basic functions won’t work.
A small flame of hope is still smoldering somewhere out there. That 14.4K still amuses me, because I can’t even write a second message to the model with a 16K limit ;p. Let alone a situation where I need to do a specific task with playwright-cli. I have a feeling that for me on the FREE TIER they’ll throw 100k at me, and still 95% of my tasks won’t go through. Because the main operating mode in my application is browser + google search and agentic action to create information sets. And when several agents use the browser, tokens fly like crazy.
This issue happened before I think, cause I saw a reddit post about it mark as 3 months ago so lets hope its just infrastructure instability and maybe them misconfig as a cover up…
I don’t think that this is a mistake. It is a policy implementation.
They want you to use the more recent, more advanced and more expensive models (it bring in more money using the same Google resources) .
Why would I think so? Because I was using gemini-2.5-flash and pro for AI assisted parsing and data extraction in Apps scripts, and it suddenly stopped working. It appeared that they deprecated the models, even though it is supposed to work at least until October, according to official promises.
After it got reported and complained about, they have reinstated the models, so they work again. This is yet another attemtp to test the waters, and see how the users react. If no big fuss follows, they quietly force us to use more expensive models. Simple, but sneaky.
But Gemma-4-31b-it is their flagship model when it comes to the Gemma family. They are constantly advertising it on X. First it’s some improvements, then someone presenting an app built on this model. For me, the Gemma model has something unique compared to Gemini: it doesn’t mess up my text formatting. Yeah, I know how that sounds, but Gemini since 1.5 totally ignores my formatting preferences. Meanwhile, the Gemma model gets it right always . In weeks of use, it only failed me once in that regard, and I probably broke something myself back then. The problem with Gemini-2.5 is that it’s an over-a-year-old model, while Gemma 4 is their fresh release.
No accident let’s think like this if it’s a whole lot of geniuses building and we have incomplete projects. Unless they private then we are literally feeding the module ideas for someone 
can easily access and capitalize on
It’s been over a week, and considering they can aggregate hundreds of tweets on X about a new model release and drop it into the AiStudio API in a second, changing from 16k to unlimited is a 3-minute task for whoever manages it, in a break between other duties.

Not letting this thread die, they need to see this.