Hey everyone,
I’m opening this thread to raise visibility on a critical regression that’s been heavily impacting production and development workflows over the past 48 hours (September 8–9, 2026).
When handling complex, multi-step queries—especially deep code refactoring, mathematical proofs, and large context evaluations—Gemini Deep Think gets trapped in an infinite reasoning loop. The interface remains stuck on "Thinking…" or "Ready in a few minutes" indefinitely, never reaching an exit condition or emitting the final response.
Observed Symptoms & Impact:
-
Perpetual Reasoning / Stream Deadlock: The SSE stream remains connected, but the model endlessly iterates on internal self-reflection cycles (“Wait”, “Let me double-check”) without ever reaching a stop delimiter. We have threads running for hours without output.
-
Aggressive Token / Compute Depletion: Under the current test-time compute accounting, these runaway loops chew through massive amounts of tokens behind the scenes.
-
Premature 24-Hour Lockouts: Just two or three stuck queries are enough to trip the daily compute ceiling, slapping users with an unexpected 24-hour lockout and grinding our work to a dead halt.
-
Status Dashboard Disconnect: Both the Google Cloud Service Health and Workspace Status dashboards remain completely green. Because the API endpoint holds an open HTTP 200 streaming state, infrastructure monitors fail to catch what is essentially an algorithmic deadlock.
Workarounds Tested:
-
Model Toggling (Handshake Reset): Switching the active session over to Gemini 3.8 Flash, sending a lightweight 1-token prompt (e.g., “ping”), and then switching back to Deep Think occasionally unsticks the session state. However, sending another complex prompt immediately re-triggers the loop.
-
Cache & Profile Sanitation: Tested in clean Incognito sessions, performed hard refreshes (Ctrl + Shift + R), cleared all local site data for gemini.google.com, and disabled Custom Instructions / Saved Info. Brand-new chat threads still encounter the exact same stall.
-
Constrained Prompting: Explicitly prompting with “Do not perform iterative re-verification; limit internal reasoning to 3 steps and output immediately” occasionally bypasses the trap, but it strips away the analytical depth we rely on Deep Think for.
Questions for the Community & Dev Team:
-
Is anyone else seeing this sudden spike in infinite loops over Sept 8–9?
-
Has there been an undocumented backend rollout affecting the PRM verifier, attention thresholds, or subagent tool orchestration?
-
Feature Request: Could the team implement an explicit client-side timeout or an automatic reasoning circuit-breaker (e.g., a hard cap on thinking tokens per turn) so runaway processes don’t torch an entire day’s compute quota?
Would really appreciate acknowledgment or confirmation from anyone on the Google AI team.
Environment:
-
Tier: Google AI Ultra / Gemini API (Enterprise)
-
Primary Target: Gemini Deep Think Mode
-
Fallback Tested: Gemini 3.8 Flash
-
Platform: Web UI & AI Studio
Hello @chunghiadaidong ,
Thank you for providing the detailed environment context and workarounds.
To help our engineering team reproduce the reasoning stall, could you share a sample or minimal version of a prompt (and any system instructions) that reliably triggers this loop?
Once we have a reproducible test case, we will investigate the behavior with the team.
Hi @Sai_Deepika_K,
Thank you for following up and looping in the team.
Status Update (Sept 15):
The issue appears to have resolved on our end today. Complex queries that previously triggered the reasoning deadlock are now terminating cleanly and returning complete outputs within expected thinking windows (roughly 45–90 seconds).
For the engineering team’s regression test suite and post-mortem tracking, here is the workload profile and a sanitized, reproducible prompt skeleton capturing the exact multi-constraint mechanics:
-
Environment: Default settings (no custom system instructions or active memory anchors).
-
Workload Archetype: Multi-constraint text compression balancing quantitative bounds (strict speaking duration and word/character count) against qualitative constraints (preserving argumentative rigor and rhetorical structure).
-
Synthetic Prompt Skeleton:
“Compress a 3,000-word formal presentation to fit a strict 15-minute speaking limit (assuming a measured academic delivery pace of 180–200 words/minute). The output must strictly preserve all core logical proofs, refutations, and critical evidence without diluting academic rigor or authoritative tone.”
-
Observed Failure (Sept 8–9): The session held an open HTTP 200 SSE connection, spun endlessly in “Thinking…” for 9 to 10 consecutive hours—likely trapped in an internal self-correction loop evaluating quantitative cuts against semantic completeness—before abruptly terminating with a generic "Something went wrong" error.
Could you confirm whether an undocumented backend mitigation, PRM verifier adjustment, or rollback was deployed over the weekend?
Lastly, this incident strongly reinforces the value of an automated reasoning circuit-breaker (such as an execution timeout or a hard cap on internal thinking tokens). Allowing an unhandled process to spin silently for hours risks needlessly tying up backend resources and potentially exhausting daily compute quotas.
Thanks again for looking into this!
My deep think function has failed so thoroughly, that I fear using it out of anxiety that I will lose more long context threads and split and save threads like I did playing RPGs in the 90s. It initially resulted in multiple threads Hitting Infinite loops around June 28th, shortly after Flash rolled out. I was calculating geometric tensor mechanics for a QED theory and framework I have been working on when Gemini locked up. I tried the to run similar prompts in other models Pro 3.1 and deep think that were trained and switched to deep think for the prompt, and they all locked up as well. Tokens roasted day after day, when reluctantly I had to delete a year worth of chats spanning more than a dozen threads. This occurred on all of my devices. I have lost more threads trying Deep Thinks, since this major loss. It took me months to catch up to where I was, so now I print more and back more up. If the Split tool worked, I would use it, but it does not work. I have discovered work arounds that involve having the model generate a prompt to train a new thread using flash 3.8 for tensor calculations and complex dimensional and temporospatial calculations with high success rates. My PC’s now get Deep Think to run, and it spits out a response… but only once does it provide the broken and often hostile text while remaining in an Infinite Loop. I now delete any Infinite loop threat immediately and spread my work thin. I miss Deep Think. It’s why I downgraded from $200/mo ultra to $100 plan and am now considering new platforms… or building my own LLM.