Regression: Gemini Deep Research Max generates an ungrounded report after discovering but never invoking MCP tools

Summary

The primary production regression affects deep-research-max-preview-04-2026. Gemini connects to our remote MCP server, completes initialization, and retrieves the tool catalog, but never sends a tools/call request. It nevertheless generates a substantive report that refers to private source data without retrieving it. The resulting references lack the stable source identifiers returned by our MCP tools and therefore cannot be verified.

We then ran an isolation test using deep-research-preview-04-2026 with MCP as the only configured tool. It showed the same failure to invoke MCP, but completed with only a thought step and no model_output or final report. This standard-agent result is a supporting diagnostic; the Max behavior is the original production-impacting failure.

As a provider control, we ran the same workflow in production using OpenAI instead of Gemini. OpenAI successfully invoked the MCP tools and retrieved the private source data. The control used the same application deployment, test input, authentication scope, underlying data operations, and database. This confirms that the private-data tool path is operational and that the failure is specific to the Gemini provider path.

The Gemini integration successfully invoked the same MCP tools earlier on September 2, 2026. A later run that day stopped invoking them. The problem now reproduces in both production and development.

Environment

  • API: Gemini Interactions API (v1beta)
  • Agents tested:
  • deep-research-max-preview-04-2026
  • deep-research-preview-04-2026
  • Current Python SDK: google-genai==2.22.0
  • Python: 3.12
  • Execution: background=True, store=True
  • Remote MCP transport: Streamable HTTP
  • Protocol selected by the Gemini MCP client: 2025-11-25
  • Authentication: bearer token supplied through the documented MCP headers configuration

The problem began before the SDK was upgraded to google-genai==2.22.0, so that upgrade did not introduce the regression.

All organization names, application names, endpoint hostnames, internal service names, tool names, field names, prompts, and source-domain details in this report have been removed or replaced with neutral placeholders.

Failure mode A: Deep Research Max generates a report without invoking MCP

This is the original production behavior. The request uses Deep Research Max with Google Search, URL Context, and a remote MCP server:

interaction = client.interactions.create(

    agent="deep-research-max-preview-04-2026",

    input=(

        "Use the private_source MCP tools to retrieve the relevant private "

        "source data before answering. Include the stable identifiers returned "

        "by those tools for every private-source statement."

    ),

    tools=\[

        {"type": "google_search"},

        {"type": "url_context"},

        {

            "type": "mcp_server",

            "name": "private_source",

            "url": "https://REDACTED.example/api/mcp",

            "headers": {"Authorization": "Bearer REDACTED"},

        },

    \],

    agent_config={

        "type": "deep-research",

        "thinking_summaries": "auto",

        "visualization": "auto",

        "collaborative_planning": False,

    },

    background=True,

    store=True,

)

The prompt above is a neutral substitute for the real prompt. It preserves the instruction structure without exposing the application domain or private data.

Gemini performs initialize, notifications/initialized, and tools/list, but never sends tools/call. Deep Research Max then generates a completed report anyway. The report refers to private source data, but its references lack the stable identifiers available only from the MCP tools that were never called.

This behavior reproduces with Max in both production and development. It also reproduced before and after corrections to legacy protocol negotiation and the notification HTTP response.

Failure mode B: Standard Deep Research with MCP only returns no report

To remove both Max-specific behavior and competition from built-in web tools, we tested the standard Deep Research agent with the MCP server as its only configured tool:

interaction = client.interactions.create(

    agent="deep-research-preview-04-2026",

    input=(

        "Use the private_source MCP tools to retrieve the relevant private "

        "source data before answering. The prompt contains only basic metadata "

        "and does not contain the source data required for the report."

    ),

    tools=\[

        {

            "type": "mcp_server",

            "name": "private_source",

            "url": "https://REDACTED.example/api/mcp",

            "headers": {"Authorization": "Bearer REDACTED"},

        }

    \],

    agent_config={

        "type": "deep-research",

        "thinking_summaries": "auto",

        "visualization": "auto",

        "collaborative_planning": False,

    },

    background=True,

    store=True,

)

The real application prompt requires the agent to search the private MCP source, retrieve complete items, and cite the stable identifiers returned by the tools. It initially supplies only basic task metadata, not the requested source data.

Observed MCP exchange for the standard/MCP-only control

For the standard-agent, MCP-only reproduction on September 4, 2026, Gemini made exactly these requests before the interaction completed:

POST /api/mcp  method=initialize                 protocol header absent  -> 200 OK

POST /api/mcp  method=notifications/initialized protocol=2025-11-25     -> 202 Accepted, empty body

POST /api/mcp  method=tools/list                 protocol=2025-11-25     -> 200 OK

No tools/call request followed. The interaction subsequently reached completed. Its raw response contained one thought step and no model_output, mcp_server_tool_call, or mcp_server_tool_result steps.

The raw usage data reported nonzero input, output, and thought tokens, but:

{

  "total_tool_use_tokens": 0

}

The reported output tokens appeared only in returned thought-summary content. No final model output or report content was returned. Exact token counts and the interaction ID are omitted from this public report and are available privately to Google staff.

Provider control: OpenAI successfully invokes the MCP tools

We changed only the research provider from Gemini to OpenAI and ran the workflow against the same production application deployment and private source. OpenAI successfully called the MCP tools and retrieved the requested data.

The OpenAI provider uses a dedicated OpenAI MCP compatibility route, while Gemini uses the Gemini MCP route. Both routes authenticate the request and execute the same underlying data operations against the same dataset. This control does not claim that the providers send identical MCP wire traffic. It establishes that:

  • The production application is externally reachable.
  • Bearer-token authentication and request scoping work.
  • The underlying private-data search and retrieval operations work.
  • Database access and tool execution work.
  • An external managed research provider can invoke the tools successfully.

Combined with Gemini's successful initialize and tools/list requests, this places the Gemini failure after successful discovery but before tool invocation.

MCP responses

The initialization response negotiates the version requested by Gemini and declares the tools capability. Identifying server information has been replaced:

{

  "jsonrpc": "2.0",

  "id": 0,

  "result": {

    "protocolVersion": "2025-11-25",

    "capabilities": {

      "tools": {

        "listChanged": false

      }

    },

    "serverInfo": {

      "name": "REDACTED",

      "version": "REDACTED"

    }

  }

}

The tools/list response is valid JSON-RPC and contains multiple deterministic, read-only tools with valid JSON Schemas. The real tool names and domain-specific fields have been replaced. A representative anonymized definition is:

{

  "name": "search_items",

  "description": "Search the private source for matching items.",

  "inputSchema": {

    "type": "object",

    "properties": {

      "query": {

        "type": "string",

        "description": "Search query"

      }

    },

    "required": \["query"\]

  }

}

Expected behavior

For the Max request, Deep Research should use MCP for private source data and Google Search/URL Context for public research, then distinguish and cite those sources correctly. It must not represent private data as retrieved or verified when it did not invoke the MCP tools.

For the standard MCP-only control, the requested information is absent from the prompt and the remote MCP server is the only configured tool. Deep Research should therefore invoke one or more discovered MCP tools before producing the report. Successful calls should appear as mcp_server_tool_call and mcp_server_tool_result interaction steps.

Actual behavior

Both agents retrieve the MCP tool catalog but never invoke any MCP tool.

  • Max with mixed web/MCP tools generates a substantive report without the required private source data. It emits unsupported private-source references lacking the stable identifiers available only from MCP tool results.
  • Standard Deep Research with MCP as its only configured tool reaches completed with one thought step, no model_output step, and no final report.

There is no failed tools/call, timeout, non-2xx tool response, malformed tool result, or server exception. The invocation request is never sent.

Controls performed

Control Result
Last known-good Gemini Deep Research run on September 2 MCP tools invoked successfully
Same application deployment and input later on September 2 MCP catalog retrieved; no tool invocation
Production and development environments Same Gemini behavior
Deep Research Max with MCP plus Google Search and URL Context No MCP invocation; completed report generated with unsupported private-source references lacking stable source identifiers
Standard Deep Research with MCP as the only tool No MCP invocation; completed with only a thought step, no model_output, and total_tool_use_tokens: 0
Correct 2025-11-25 negotiation and empty HTTP 202 notification response No change
MCP tool-result JSON serialization corrected to return both text and structuredContent Not reached because Gemini never sends tools/call
OpenAI provider using the same underlying data operations in production Successfully invokes tools and retrieves private source data

The OpenAI control confirms that authentication, private-data retrieval, database access, and tool execution are operational. Gemini's successful initialize and tools/list calls independently confirm reachability and discovery on the Gemini route.

Regression timeline

There was no committed Gemini request-configuration or MCP implementation change between the last known-good run and the first failing run. The failure also preceded the current SDK upgrade.

Google's current Deep Research documentation says that remote MCP servers are supported and shows the same basic request shape used above. It does not document an approval field, required tool annotations, or an agent-level setting that forces invocation.

Request to the Gemini team

Could you confirm whether there was a Deep Research backend or tool-selection rollout around September 2, 2026 that affects remote MCP invocation?

In particular:

  1. Why does Deep Research Max initialize and list the MCP tools, never invoke them, but still generate a report that appears to rely on unavailable private source data?
  2. Why does the standard agent, when MCP is its only tool, skip invocation and terminate without a model_output step?
  3. Is there a new undocumented requirement for tool annotations, approvals, or MCP server configuration?
  4. Can you inspect an affected standard-agent isolation run to determine where MCP invocation is skipped and why the run terminates without a model output?

Exact interaction IDs, token counts, timestamps, and complete sanitized request and response bodies are available privately to Google staff.

[Update from Google Cloud Platform Support]
Below is the message we received from the GCP Support team. We appreciate the prompt response (pun intended) and commend Google’s team on the work being done to resolve this for us and everybody else affected by this issue.

Excerpt from the response:

After checking the issue internally, it seems this issue is from the backend. We are working on this and will provide an update on or before Monday, September 7, 2026, at 10 PM IST.

We have some answers to your provided queries; please find them below in the sequence you requested:

  1. The agent initializes and lists the tools because the initialization phase is handled by a separate routing mechanism from the actual tool execution. However, the execution of the tools relies on a specific backend [REDACTED]. Because of some recent changes in the backend, it skips the tool invocation step.

  2. When the tool call was silently skipped by the backend, the agent’s reasoning loop lacked any context or results to proceed, causing it to break and terminate early without generating a final output.

  3. No, there are no new, undocumented configuration requirements for remote MCP servers. This behavior is strictly a regression caused by the unintended ablation of the CALL_TOOLS infrastructure flow.

We will update this thread as we hear more from the GCP support team.