Feedback on gemini-3.5-transcribe and Comparison with General Gemini Models on STT Accuracy: Improvement Proposals and Recommendations for AI Project Managers on Model Training and Key Focus Areas

Feedback on gemini-3.5-transcribe and Comparison with General Gemini Models on STT Accuracy in languages Slovak and Czech: Improvement Proposals and Recommendations for AI Project Managers on Model Training and Key Focus Areas

Introduction and Review of the Gemini Model for Speech-to-Text

Dear Google employees and developers working on LLMs, Gemini AI, and specifically… my words and this review are aimed at speech-to-text, specifically the Gemini 3.5 Transcribe model.

Further in this piece, I will also share unique know-how and guidelines on how to train artificial intelligence when it comes to speech-to-text. Because you have a lot of catching up to do in this area; it is currently neither good nor ideal. So, on the one hand, I will certainly criticize you quite heavily, but on the other hand, I won’t leave you stranded. I am a very good person; I will help you as a benefactor and tell you what could be improved.

The Value of Customer Feedback and Reviews

Customer reviews like this carry immense value—worth millions of euros—because receiving this caliber of feedback is simply priceless. Specifically, I tested this speech-to-text service in Slovak and Czech. I don’t know how it performs in English; I assume you get better feedback in English, as it is logically the primary and default language.

Because I am exceptionally talented in management and am a very creative individual, solid solutions on how to improve things naturally occurred to me. On the one hand, it is very depressing and sad that this product of yours is truly garbage. On the other hand, once refined, it could be genuinely perfect and outstanding. Not to sound entirely negative, it is commendable in itself that you developed a dedicated model focused purely on speech-to-text in the first place. That deserves praise; that is very good. Up until now, that didn’t exist.

Comparison of the General Gemini Model with Google Cloud Console and Chirp 3

As for the Google Cloud Console, that is complete garbage. Chirp 3 is absolute garbage, even worse than this. But this model is at least slightly better. Let me explain how it works and what my experiences have been.

To be honest, the general Gemini model works best for me. The general Gemini model delivers substantially better and far higher-quality speech-to-text. When I explicitly instruct it in the prompt to also make stylistic revisions, it handles the task relatively well.

Two Extremes of Transcription: Verbatim with Slips vs. Unwanted Summarization

Even with the general model, I had quite a hard time fine-tuning it, because the AI constantly swung between two very unpleasant and ugly extremes. I dislike both extremes; I prefer a golden mean. It either transcribes every single slip of the tongue and filler sound completely verbatim, or—when I instruct it to polish the text slightly—it turns it into an aggressive summary. It simply omits very important words, leaves out crucial and essential details, and strips away the genuine original energy that I want to convey.

When I record a podcast, I want it to be as authentic as possible, retaining the emotion and style, rather than turning into an executive summary where every third word is dropped. Because when it skips every third, fourth, or fifth word, it becomes a problem, and critical details are lost, which bothers me greatly.

Critical Bug in Gemini 3.5 Transcribe: Occurrence of Foreign Characters and Cyrillic

Regarding Gemini 3.5 Transcribe: while on one hand it’s a disaster, on the other hand, I appreciate that you are specializing in this capability. I understand that you likely wanted to create a solution to offload Google’s servers with lower hardware demands. That in itself is good and correct. This is the right direction to pursue, but it must be executed properly, sensibly, and well.

Then again, it is obviously better than nothing. Beginnings are always difficult, and even though this truly feels like an alpha build, in Slovak it simply does not work as it should. Since I personally write entire articles this way—first making a voice recording and then drafting emails or articles from it—I like using this functionality. Unfortunately, this Gemini 3.5 Transcribe is practically worthless.

Let me explain specifically what bothers me the most. Even when I explicitly configure only three allowed languages—inputting English, Slovak, and Czech—it still spontaneously outputs certain words in Cyrillic or something similar. As a Slovak citizen, I do not use Cyrillic; we use the Latin alphabet, so I cannot identify with certainty what those characters were. They resembled Cyrillic, or perhaps Greek script—I don’t know, but I dislike it and it is unacceptable. Even with explicit instructions provided, this is a very serious bug and a critical flaw. Unless such issues are fixed, the service will remain complete garbage and unusable.

How often do these errors occur? I would say roughly five to eight times within an approximately six-minute recording. That is how frequently these errors happen, which is understandably a major issue.

Doubts Regarding the LLM Architecture and Suspicion of a Primitive Neural Network

Furthermore, this Gemini 3.5 Transcribe model really does not strike me as an LLM. To me, it feels more like a legacy neural network—the predecessor to LLMs. If errors are found in a five-minute recording with phonetic, overly literal transcriptions… Of course, in Slovak, certain words are pronounced differently from how they are spelled. Here, it produced a completely literal phonetic transcription, as if it didn’t even grasp or know basic Slovak grammar.

This indicates that it certainly wasn’t powered by an LLM, or at the very least, the LLM must have been so aggressively miniaturized that it de facto behaved like an outdated neural network. It is an utter disaster and poorly executed; this needs improvement. You have to start improving it, otherwise it will remain unusable. I would be very pleased if you upgrade it, release a new version, and implement specialized training for this type of artificial intelligence so that it handles these tasks properly.

Proposal for Stylistic Scalability via API and Parameter Settings

Next, I propose completely unique know-how. It would be beneficial if users could—if not in the GUI, then at least via the API (I mostly use the API and rarely the GUI, just as a side note)—configure the degree of stylistic editing applied to the text. This would prevent the two extremes I mentioned and allow for a scalable adjustment. Perhaps a parameter from 1 to 5, where you simply select the intensity of text revision. This is a crucial feature.

Text Formatting: Paragraph Configurations, Subheadings, and Proper Interpretation of Speech Errors

Furthermore, it should allow paragraph formatting options: whether one wants raw unformatted text, paragraphs grouped by topic, or even subheadings for individual sections. It doesn’t have to be the default setting, but having it as an available option would be very valuable.

Another very important point: it often fails to correctly interpret verbal slip-ups. This sometimes happens even with the general model. When I misspeak and subsequently correct myself, an LLM—striving to comprehend semantic meaning—should naturally evaluate such context automatically.

For example, if I make a mistake regarding a number and say: “I traveled 6 km, but in reality I traveled 9 km,” and immediately correct myself: “Sorry, I didn’t travel 6 km, I traveled 9 km,” it should simply omit the mistake so I don’t have to edit it manually later. This is another essential capability: accurately interpreting speech errors and self-corrections. When I correct myself during a podcast, the output should directly reflect the corrected phrasing rather than the initial blunder.

Artificial intelligence and its new models must be trained with this in mind. The training should be specialized in this manner to handle these nuances properly, avoiding the two extremes I pointed out, while perfecting the handling of self-corrected speech errors.

Processing Speech Errors in Audio Domain Prior to Text Conversion

AI can only correct various verbal stumbles accurately if it processes them as sound rather than pure text. Vocal timbre and intonational cues are vital for proper comprehension—not just for humans, but for AI as well. I have had very poor results—it absolutely never works well—when asking a model to stylistically polish a raw verbatim transcript containing every blunder. It simply falls short and never does it well.

It requires an enormous amount of effort to post-process that text, whereas it will always be noticeably better and higher quality if those stumbles are resolved directly within the audio domain rather than after the fact in text. This is critical, as pitch, inflection, and tone convey the true intended meaning in ways text cannot always capture in detail.

Offer of Collaboration, Managerial Know-How, and Financial Compensation

Thus, I have shared completely unique know-how and insights with you. While what I am saying might seem obvious on the surface, it represents distinct know-how, and without my guidance, you likely would not have arrived at these essential solutions. I would be very glad if you incorporated these proposals.

Writing to you in such detail takes a tremendous amount of my time, dear developers. If you are from senior management, I would certainly welcome your contact, and I would gladly accept financial compensation for refining and significantly contributing to the strategic development and data training of Google’s artificial intelligence. Do not hesitate to reach out; you have my email address, so simply send an email, and I will be more than happy to discuss financial terms.

Speaker Diarization, Smart Transcription, and Confusion Between Slovak and Czech

Another crucial enhancement for Gemini 3.5 Transcribe is support for speaker diarization (speaker separation), alongside retaining the Smart Transcription capability in diarization mode—meaning at least minimal stylistic cleanup of the text. Remarkably, even when I enabled Smart Transcription, it still inserted Cyrillic characters. Even under these conditions—and I exclusively use Smart Transcription—it performs poorly, and I am unsatisfied. If I unnecessarily repeat the same word multiple times in a sentence, it simply fails to prune the redundancy. I dislike this, and it works poorly. Once again, the general Gemini AI model works reasonably well here. That model performs genuinely well, and it deserves credit for that.

It has also happened that it transcribed a Slovak word into Czech, which is another huge mistake. Distinguishing between Slovak and Czech is fundamental… When I speak Slovak the entire time and haven’t uttered a single Czech word, why is it outputting in Czech? That is a glaring flaw. It behaves so primitively, as if it were an old-generation neural network or Chirp 3. In fact, even Chirp 3 handled that specific distinction slightly better.

This is a serious problem that must be resolved. I am providing you with managerial advice: put substantial effort into training this new model so that you iron out all these bugs and flaws, fix them once and for all, and eliminate these problems. I don’t know about English, but in Slovak, this Gemini 3.5 Transcribe model is an absolute disaster.