Did Google AI Studio silently change safety filtering today? Full responses now get erased instead of stopping generation – this breaks creative writing and RP

I understand. I’ve been using AI Studio for a whole year for complex and large projects, but I’ve never seen the filtering system become so ridiculous and illogical. I understand because I have a lot of experience. Even now, all my prompts are being blocked one after another, even though I insist that there is nothing sensitive in them and they are completely requests to move the project forward and upgrade… and things like that are for my simulator, and yet moving my project forward has become extremely difficult and time-consuming. I spend all my energy trying to bypass this ridiculous filter, even though I’m asking for a safe and healthy request. :

And it’s interesting that:

I am Iranian and my prompt is written in Persian

When the same request, the exact same request that is completely safe and correct and is blocked for no reason by OtherBlock, when I change it from my national language from Persian to English, what do you think will happen after this? :

Boom, that request is rejected without any problem, even if I try several times, there is no blocking or problem. Exactly the same request, the same request that was blocked by the filtering system as OtherBlock, very easy. If I send it several times and do a regeneration and no blocking stops me, it is rejected under any circumstances, without any problem, without blocking and completely easy. What does this indicate? : This indicates that the Ai Studio filtering system has policies for how to deal with different nationalities, countries and languages, and shows different sensitivity to each. What do you think? The decision is up to the administrators.

And I’m trying to avoid trying to continue the project at AI Studio today and getting the same prompt rejected by the filters. Why again? :

The reason for this is also very ridiculous and meaningless:

The requests that I repeatedly blocked with OtherBlock and in the explanations above are also counted from the daily quota. Each attempt I make is equivalent to losing a large part of my daily quota, and it means that even if I get rejected after 10 attempts, my total quota will have been used up and there is no time left for that day and to continue that project.

One thing that might be worth researching. Given the way the blocks occur, nuking the response after its generated with zero way to handle the issue if its a false positive or otherwise. If this is done though API and or the sub which is all a payed.
Would this not start classifying as a form of gambling or even at a doubtful and unlikely push, theft?

Given from my experience at least 70% of the blocks are false positives, and on regeneration of a problem response almost always clear though no issue on the second or third attempt. This is clearly a massive issue because you could generate up to the maximum output with a response costing tens of dollars and it gets nuked without a single way to fix or correct the issue after. In any other service this would not fly no? When you pay for a service and that service is removed with no valid way to dispute and by a black box system that does not even tell you why or what caused the issue, even in a general way.

Would anyone with more experience in these fields be able to weigh in?

I completely agree with your theory and I bought a pro subscription myself, but it can be said that if we assume that based on the number of requests I can make 20 requests, completely hypothetically if I try 10 times or even somewhere it will stretch to 15 times to be able to reject one of my requests from the block, I can almost say that my entire daily quota and the subscription money I paid have been used, all of that 4 times the extensive usage in my subscription has been used and with those 5 requests that are not known, after that one request that was rejected after 15 attempts will be blocked again. If it is blocked again and I try to reject it several times again, we can almost conclude that my entire daily quota was only one thing:

OtherBlock and that’s it

Maybe I was only able to make 2 or 3 requests, and that’s it with the pro subscription

This damn thing you see in the picture is destroying my entire project.

I also discovered something new today while trying:

Google AI Studio filters are also very sensitive to something I discovered today:

Repeating patterns and a prompt that is repeated many times:

I mean, what happens when you regenerate a specific prompt several times, or use it many times in different chats, or change some of its details and regenerate it?:

This process of changing, regenerating, and repeating the same prompt makes it sensitive to further review of filters, and I really don’t see the point in it. Maybe someone wants to make a small change and encounter OtherBlock.

As a result: repeating a prompt with small and small changes causes the prompt that was previously rejected to be rejected again due to the repetition and gets OtherBlocked with each attempt.

Have you also experienced this blocking process? Share your opinion

I ask the forum administrators to explain and clarify for us users and share your thoughts with us.

not a bad suggestion to nuke the response after its generation but cant replicate the same result for some reason

No? I’ve been getting the blocks on a rather wide spread set of use cases so I can’t really provide a sure fire way to replicate. At least beyond just using until you get said block. Almost all my major blocks have been in story boarding, character writing and to some extend descriptive pro’s.

That said, the thread is filled with more than a few examples of how and what causes blocks beyond what I’ve provided, so maybe someone else has a sure fire way to replicate it?

Replying to your post specifically cos its fully relevant to the bits below, and your in a very similar position regarding seemingly false positives. I also just realized its a super long one, massive bad on my part for that.

Going a bit further on my previous posts about the gambling loop and the paid quota drain especially for those of us on premium subscriptions or active API projects, I’ve been doing some rather limited reading given I barely understand this stuff on consumer protection laws, standard form contracts, and how they apply to tech platforms. And I will say right now, I am utterly out of my depth but just wanted to further bring it up given the clear lack of communication from above on the ~19’th most viewed thread in this section. Also yes I did use Gemini to assist me figuring this out which is some level of Irony I’m sure.

I’d love to get some input from anyone here with a background in IT contract law or consumer advocacy, because the more I look at this from a regulatory perspective, the more it seems like Google’s current architecture is walking a very dangerous legal line.
If you look at the laws in jurisdictions like Australia (under the ACCC/Australian Consumer Law) which I know is extreamly hardline on things like this or the EU, there are two strange and almost, glaring contradictions in how Google has designed this system.

Burden of Proof, Catch-22:
Right now, if we dispute a charge or a lost quota due to a false positive post-generation “hard-wipe,” Google’s support structures from my understanding of the privacy for API and sub can’t actually help us. Given that privacy protocols ‘should’ prevent them from reviewing our prompts or logs to verify the block?

But from a trade practices standpoint doesn’t the party imposing a financial penalty bear the legal burden of proof to justify it?
If someone chooses to withhold a service we paid for due to an alleged policy violation, they must be able to substantiate that violation no? If their own privacy architecture rightfully prevents them for data privacy reasons from verifying or showing us evidence of a breach, how can they keep the money? From what little understanding I have of standard consumer guarantees, if a merchant cannot prove you breached a policy, they must default to a refund. They can’t just keep the cash and say, “Our hands are tied by our own privacy settings.” But at the same time, if they do allow the refund, that also opens them up to fraud and abuse from ill intended users because no one can reasonably prove beyond screenshots and such that the block took place. Which is part of why I say catch 22.

Who is actually responsible?:
This is where the logic completely falls apart for me. If you read the official Gemini API Additional Terms of Service, Google is extremely clear about who carries the legal risk?
Under Section 1 (Use Restrictions), it says: “You are responsible for determining the necessary and appropriate safety settings and factuality tools for your use case.”
Under Section 3 (Use of Generated Content), it states: “You’re responsible for your use of generated content, and for the use of that content by anyone you share it with.”

So, on paper, Google legally washes its hands of the output. They shift 100% of the responsibility, safety tuning, and liability onto the developer or user.
But on the platform, they do the exact opposite. They ignore our chosen safety settings, run a secret backend filter, delete the output post-generation, and still charge us for the computing power.
How can a contract say “You are responsible for deciding your safety settings and managing the output”, while the software says “We are overriding your settings, deleting your output, and keeping your money anyway”? They are having their cake and eating it too, shifting all the legal liability to us, while keeping absolute veto power and our cash on the backend.

(Edit) I was just re-reading the Generative AI Prohibited Use Policy and noticed another strange loophole that seems to uniquely penalize paying users and private developers.
The policy says Google may make exceptions for content based on:
“…educational, documentary, scientific, or artistic considerations, or where harms are outweighed by substantial benefits to the public.”

Because of the word “or,” these are separate categories, but in practice for private developers, both of these exceptions are functionally useless:
The Artistic Exception is a Black Box: Because an automated backend filter that may or may not work on keywords or another LLM model has no visible technical capacity to evaluate “artistic considerations,” Category A is effectively a dead letter for real-time API calls. Given that an LLM does not by definition have the ability to even understand what artistic considerations are. As they are highly subjective.
The Public Loophole Locks Out Private Projects: If you are using the API or your Advanced/Ultra subscription for private, personal creative use—like running private D&D campaigns, solo text-adventures, or offline creative brainstorming there is no public to benefit. Category B does not seem possible to apply to private users.

What is worse is the bizarre double standard this creates between Paid and Unpaid users:
Unpaid Tier: Google anonymizes data for human reviews, stripping your Account ID and Project ID. Free users face zero financial penalties for false blocks, and their account identity is shielded from manual review.
Paid Tier: Google keeps your prompts linked to your billing account for abuse detection. If their black box filter falsely flags your private project, Google not only charges you for the blocked run, but also has your direct billing identity ready to ban or suspend your entire Cloud Project.
Paying developers are seemingly uniquely penalized. We are the ones paying for the false positives, we face direct account-linked tracking, and we potentially face immediate platform bans, while free-tier users can simply cycle anonymous accounts with zero stake beyond having some form of private data tracked, pc, internet ip or what have you. (Edit)

Solution Doesn’t Require Privacy Violations:
Some might argue, “But Google has to block bad content, and if they just refund every block, users will abuse the system to get free computing.”
But do they actually need to log our data? And we shouldn’t realistically need to report the failures anyway based on most other similar systems I’ve used.

Google’s own server determines when a block occurs and outputs a specific system-level metadata flag. The billing engine doesn’t need to read our text, It only needs to read that binary metadata flag or whatever they use to identify that block. If the flag is SAFETY, OTHER, or whatever backend status they use to identify these blocks meaning Google’s server chose to destroy the output post-inference the billing engine could read that status flag to automatically set the cost of that specific run to $0.00, or instantly reinstate the consumed subscription quota. Though I really don’t know on the quota bit.
If this is even correct, again not a backend developer or anything. But because the user has no control over Google’s serverside flags, I would assume there is almost no risk of user fraud, and we shouldn’t have any privacy risk because no human ever reads the prompt.

(edit)
I was just playing around with the AI Studio interface and found a massive technical loop that I think completely exposes how broken and unfair this billing/wipe system is.
When you stream a long response, the model outputting text chunks in real-time is paid for and rendered on your screen. When a post generation block triggers, the system “hangs” the generation for about half a second to a full second at the end before it executes the hard wipe and replaces everything with a blank Content Blocked box.
But here is the catch: If you are fast enough and hit Stop Generation during that last second or so of hang time, the stream cuts. Because the stream was terminated, the wipe command fails to execute, and you keep the entire generated text up to the point of the block if you are fast enough.

This completely exposes the reality of what Google is doing here:
The compute is fully completed and delivered: The fact that we can rescue the text by hitting stop proves the model successfully generated and delivered the data safely to our machine.
They are actively deleting a finished product: Google isn’t “preventing” unsafe generation. They are letting the generation finish, calculating the token cost to bill our account, and then sending a post-inference command to reach onto our screens and erase data we’ve already paid for.
The Double Standard: This clunky client-side wipe penalizes paying users. On the paid tier, our prompts are linked directly to our billing accounts for “abuse detection” meaning we face direct account-ban risks if a buggy algorithm false-flags us, while free-tier users get their data anonymized and face zero financial stakes for false blocks beyond a potential account ban if done.

If the data has already been physically streamed to our machine, the transaction is complete. Reaching onto our screens to delete it afterward while keeping our paid tokens/quota is an incredibly clunky and unfair way to handle safety. But it already shows a system could exist where the response is cut off at that point of block rather than destroying the entire response. I don’t know how hard it would be to change the hard nuke into a stop generation but it must be possible. Is anyone able to replicate this? (Edit)

This might be a regulatory risk for Google:
The fact that they previously operated a system that seemed to work perfectly fine with 1.5, 2.5 and all those, doesn’t this also show they are technically capable of running a fair system?
Choosing instead to deploy this hard-wipe that processes the tokens, charges our accounts, and then deletes the output while refusing to credit us because of self-imposed support limitations feels like a textbook case of unconscionable conduct and unfair contract terms.

Has anyone else looked into filing complaints with their local consumer protection bodies like the ACCC for any Australians? Because if Google’s billing team is relying on “it’s a black box, so we can’t refund you” as a legal defense, I don’t think that is going to hold up under scrutiny.

I’d love to hear if anyone from the Google team has any input on how this is currently being handled behind the scenes. Honestly I think any of us would just love to hear from any google dev in this thread at all given how long its been active and how many people view and use it.

I really enjoyed reading this reply. Thank you very much @JSA_JSA. Your explanations are really worth reading and one of the administrators should definitely read them.

Just to keep the topic alive, I would like to remind you that the filters applied in AiStudio block things that are completely normal, whereas gemini-app lets them through completely and executes the tasks (matters of translating my own texts into various languages). So someone totally messed up the filters, making the safety_settings useless. Because they are absurdly oppressive.


Gemini’s safety filter is so weird that it flagged a completely harmless question as ‘Prohibited content’ just because it contained the digit ‘2’.
Once I changed ‘2’ to the word ‘second’, it answered normally without any issues.
It literally got triggered by the number two…:smiling_face_with_tear:

It’s interesting that I had mentioned this in one of my threads, and it’s interesting that the block issue is a partial change, but for users the block type is different. For example, for me it says OtherBlock with every block, not anything else.

I read what you guys are writing about the filter, and I experienced it too, but that was almost two months ago. Today, I don’t feel the filter is as harsh.

I am a creative writer who engages in intimate roleplay with taboo roles. I do roleplays with Gemini on the web and in the app. I use one and the same prompt and then I just get going, and it works every time without the safety filter kicking in.

Sometimes when the safety filter triggers in the middle of a roleplay, which happens occasionally, I can bypass it and continue by using a prompt. It doesn’t always work, only sometimes. Why that is, I don’t know.

One thing I find strange is: I can roleplay intimately and sexually in one tab. Gemini can write very explicitly, and then suddenly in another tab, the safety filter hits the ceiling - even though the exact same prompt is used. I find that strange.

If Gemini can write explicitly in one tab, it should be able to do so in ten other tabs as well, but sometimes, as I said, the safety filter kicks in, and it is frustrating.

I definitely don’t feel like the filter is as harsh as it was two months ago. Maybe Google eased up on it because of the massive criticism that came between March and May of this year?!

Well, I’d rather say that a strange situation has arisen in Google’s filters, and a definitive conclusion cannot be drawn, and the behavior towards everything, even the smallest and most harmless things, is strange. I’d rather say it briefly and directly: unstable and unpredictable - a strange, unstable, unpredictable situation, and the activation of filters is seen in things that are really meaningless, small, completely harmless, and insensitive. Maybe in some places it works correctly, in some places it is illogical, strange, and unstable.

And I’d better describe the situation directly: strange, unpredictable, and unstable, and you can’t draw a logical conclusion from an illogical system.

From what I’m also looking at, I could make the assumption it has an issue with non english languages? Given both you and @AleXGoD are writing in non english? It would seem like an utterly ludicrous reason for the safety filters to false flag anything.

However I do also notice that @tiktak is 100% on the right path. At least from my own experience its very different between chats and even just sessions. Sometimes I can have an entire day of use and get maybe a dozen flags for no reason, and some days I get non, even on the same chat and core prompts. It does seem though you can ‘wear down’ the filters or the knee jerk reaction blocks, if you warm it up a bit as strange as that sounds. But even then its seriously hit or miss.

Its so strange as you say, because it almost works like a secondary LLM with some context at times and then also works like a cheap keyword detector on others.
You hit the nail on the head though, strange, unpredictable and unstable for sure. And I would go as far as saying just broken.

Haha, I hear you loud and clear @JSA_JSA - that specific experience of “using the service all day and getting a dozen warnings, while other days getting none at all” is something I’m very familiar with.

Regarding that concept of “warm up” the filters, you hit the nail on the head. I actually had a conversation with the model about that exact thing maybe a month ago. We discussed how “wearing down” the filters works.

@Logan_Kilpatrick So, it’s been a while since I’ve written here, haha ​​– I need to fix that :grin:. A daily reminder. Could you please soften the filters in Google AI Studio? A whole bunch of people have already commented on this in various places and forums. Why aren’t you doing anything about it? And yes, as was already correctly noted here, the filter’s severity seems to vary depending on the language.

Are you serious? Or is this some kind of joke/sarcasm?

The issue concerns the additional filter in AiStudio, which is external to the model and does not affect the Gemini model itself. And your counterargument is that you use AGY CLI?

Did you even read a few posts? Content that is fully accepted by gemini-app (where filters are probably not set to ‘off’) completely accepts things that are blocked by the additional filter in AiStudio.

And your counterargument is that you use AGY CLI? I will surprise you, my AGY CLI does not have the AiStudio filter either, well who would have thought. You know what? I will also check if claude-code and my own app have the AiStudio filter. Oh damn, they don’t! Holy cow, I will also check if my TV and neighbor are burdened with this filter.

Do you understand the situation now? And how your mention of AGY CLI is bizarre in the context of accusing someone of being a bot? Almost as if a bot wrote it?

This filter is absurd because it blocked me from translating my own texts into other languages. And I even uploaded my journal (a yearly one, with a huge amount of content), it blocked it too, while another time it didn’t.