Audio transcription regression: Gemini 3.5 Flash-Lite has 3.55× higher WER than 3.1 Flash-Lite on a french radiology test dataset

We observed a substantial audio-transcription quality regression when migrating from gemini-3.1-flash-lite to the officially recommended gemini-3.5-flash-lite replacement.

On the same 50 French radiology dictations (634 seconds total), using the same prompt, temperature 0, minimal thinking, and sequential execution:

  • Gemini 3.1 Flash-Lite: WER 6.19%, CER 2.94%, median latency 1.320 s
  • Gemini 3.5 Flash-Lite: WER 21.97%, CER 13.66%, median latency 1.243 s
  • 3.5 Flash-Lite had worse WER on 40/50 cases, equal WER on 7, and better WER on only 3.

Is this audio-transcription regression expected? With Gemini 3.1 Flash-Lite scheduled for end of support in May 2027, we are concerned about migrating to the recommended replacement because Gemini 3.5 Flash-Lite performs substantially worse on our ASR benchmark.