Repeated gemini-3.8-flash generateContent HTTP 503 at low reported usage; TCP/TLS trace completes

We are investigating repeated HTTP 503 responses from gemini-3.8-flash through the Gemini Developer API v1beta generateContent endpoint on an isolated Linux test host. Our requests use synthetic data only.

Our latest controlled probe on 7 October 2026 around 03:09 UTC made exactly one HTTP attempt, with no retries or model fallback. It returned google.genai.errors.ServerError with HTTP 503 after 2311.236 ms. Transport tracing recorded successful TCP connection, TLS negotiation, request headers/body transmission, response headers/body receipt, and response closure. The wait for response headers was 2249.360 ms. We did not retain the error body or a request ID in this probe.

Environment: Python 3.12.3; google-genai 2.22.0; httpx 0.28.1; httpcore 1.0.9. Configuration: gemini-3.8-flash, structured JSON response schema, max_output_tokens=4096, thinking_budget=0, automatic function calling disabled, tools=None; outer deadline 8 seconds and transport timeout 12 seconds. This latest probe used a short synthetic Vietnamese text prompt, with no audio input, TTS, WebSocket, or tool execution.

An earlier isolated audio E2E pilot on 6 October 2026 returned HTTP 200 for input TTS and HTTP 503 for its single Brain request. A separate text probe later hit its 8-second outer deadline without recording an HTTP status; we cannot attribute that timeout to a transport stage or equate it with a server-side 504.

The AI Studio project we inspected shows Free tier and peak usage below the displayed limits: RPM 1/5, TPM 1.75K/250K, RPD 2/20. Its aggregate usage chart includes 503 and 504 errors. We have not independently verified that the key used on the test host belongs to the displayed project, and we are not treating aggregate dashboard bars as per-request correlation.

Could Google clarify:

  1. Whether there is a known capacity/access issue affecting this model and endpoint on Free tier, and which non-secret diagnostic identifiers would help investigate it?
  2. Whether thinking_budget=0 with structured outputs is supported for gemini-3.8-flash on generateContent; if not, which parameter and expected error should apply?
  3. Whether the prepay migration notice applies to projects still showing Free tier and Set up billing, and where to confirm project-specific eligibility?

We do not assume that upgrading billing will resolve the failures. We have paused further live probes pending clarification. No API keys or user data are included in this report.