Title: Context caching or cached-input pricing for the Gemini Live API (native audio)?

Hello. We run a voice companion app for older people on the Gemini Live API (gemini-3.8 Live, native audio), in 17 languages. Each call sends the same system instructions and tool definitions (about 7,000 text tokens), and they are re-billed as input on every turn. That is roughly two-thirds of our cost per minute.

Does the Live API support context caching, or implicit caching of the system instruction, so the repeated instructions are billed at a cached-input rate? If not now, is it planned?
If there is no caching, is there volume or committed-use pricing for the Live API, through AI Studio or Vertex AI?
On Vertex AI, does the Live API pricing differ, or does it offer caching that AI Studio does not?
We are pre-launch and estimate tens of thousands of call minutes a month within a year. Thank you.

Hello @Deepak_T_Patel,

  • Context caching (explicit and implicit): Neither explicit nor implicit context caching is supported for gemini-3.8-live. Active session context—including system_instruction and tool declarations—are billed at the standard text input rate.
  • Managing multi-turn session token growth: You can enable sliding-window context window compression (ContextWindowCompressionConfig with SlidingWindow) to automatically trim older conversation turns so historical dialogue does not accumulate across a long call; however, system_instruction is always preserved at the start of the window and remains billed on each turn.
  • Volume and committed-use pricing (Gemini Developer API): The Enterprise plan on the Gemini Developer API Pricing page offers volume-based discounts and provisioned throughput. You can reach out via Contact Sales to discuss high-volume commitments.
  • Vertex AI inquiries: For questions specific to the Vertex AI API (including Vertex AI pricing and feature support), please contact your Google Cloud sales representative if you have one, or post on the Google Cloud Forum.