Hello everyone!
I’m experiencing an issue with large-context conversations in the Google AI Studio Playground.
I mainly use AI Studio for a long-form roleplay. I have one continuous chat that has grown to around 400,000 input tokens.
Recently, all the newer Flash versions I tried started returning this error:
“Input token count exceeds the maximum number of tokens allowed for this model. Please adjust your prompt and try again.”
This happens with Flash 3.8, 3.7, 3.6, and 3.5.
My understanding is that these models support context windows of up to 1 million tokens, so I don’t understand why a ~400k-token conversation is being rejected.
Flash 3 was working somewhat better, but today I started getting another error:
“Model isn’t available right now. Please wait a minute and try again.”
This can happen in two different ways. Sometimes the model generates the thinking/reasoning and then returns this error instead of the actual response. Other times it starts generating the response normally, but the same error appears partway through the post.
I also have an issue with the daily usage limit in this particular conversation. I have a Google AI Pro subscription, but the limit is exhausted very quickly when I use this long chat. At the same time, new short chats work normally.
I understand that a 400k-token conversation is large, but it is still well below the advertised 1M-token context window. I would therefore like to understand whether there is another limitation on long conversations in the Playground, such as a lower effective context limit, a per-conversation limit, or some other restriction.
Has anyone else experienced this with a large long-running conversation in AI Studio Playground, especially for roleplay or other long-form writing?