Hello. We run a voice companion app for older people on the Gemini Live API (gemini-3.8 Live, native audio), in 17 languages. Each call sends the same system instructions and tool definitions (about 7,000 text tokens), and they are re-billed as input on every turn. That is roughly two-thirds of our cost per minute.
Does the Live API support context caching, or implicit caching of the system instruction, so the repeated instructions are billed at a cached-input rate? If not now, is it planned?
If there is no caching, is there volume or committed-use pricing for the Live API, through AI Studio or Vertex AI?
On Vertex AI, does the Live API pricing differ, or does it offer caching that AI Studio does not?
We are pre-launch and estimate tens of thousands of call minutes a month within a year. Thank you.