Use of the Google AI Studio API and Criticism of the Chirp 3 Model Quality
Another very important point is what specific services I used. I exclusively used an API key from Google AI Studio. This is very important for you to know regarding the types of Speech-to-Text services I utilized. Regarding the Chirp 3 service for speech-to-text, for instance, that is complete garbage. It is absolutely unusable and has extreme inaccuracies, at least in Slovak. I primarily use Slovak, but I also used Czech, and it is a total disaster, unusable, garbage, and worthless. That is why I actually use Gemini in AI Studio. If you have contact with senior managers, you can ask them to discontinue the Chirp 3 model for Slovak because it is garbage and makes a mistake in every other sentence. Conversely, when speech-to-text is performed via Gemini in AI Studio, it is quite tolerable, and the training data in this area is much more up-to-date, making it far superior.
Issues with File Length, Free API Overload, and the Value of the Opus Format
This time, I used speech-to-text services mostly for Czech, and these were MP3 files. However, there were major issues. I had to create a script that split two-hour files into four separate half-hour files. That is how it worked, and it was quite inconvenient because I didn’t have time to deal with it. Naturally, it cut off either mid-sentence or in the middle of an unfinished word, and the resulting transcription was not as high quality, especially at the beginning and the end. This was very frustrating because the system simply cannot handle larger files. There is a huge problem that it clearly cannot handle larger files in Slovak and Czech, and it definitely fails.
I also try to use the free API whenever possible. However, there is a huge problem there as well, because the free version constantly tells me that the servers are overloaded and cannot be used. These servers are overloaded about 95% of the time, so I can rarely use it. This is a great pity because training data is worth far more to Google. When a service is free, you acquire training data that can be utilized, and that data is truly worth its weight in gold. When I upload such files in Slovak, I always use Opus format files. When you upload training data in an Opus file, it is worth gold to Google’s developers and team. This way, you are missing out on truly valuable training data in Opus format, which has multi-million-euro value because it is not just an ordinary file. Average amateurs use all sorts of MP3s and nonsense, but I used Opus—that is the absolute best for training, the top tier. I am an IT professional and elite, and this way you are pointlessly losing training data because the free version is constantly overloaded.
Failures of the Paid Gemini Version, “Thinking Level Low” Setting, and Timestamp Errors
Even in the paid version, it did not work as it should have. There was also a massive problem because even this paid version could not keep up. I had to set Thinking level to Low, which reduces the model’s reasoning ability, and it still experienced massive issues, even though it was a paid version. I cannot pay for things and services that were completely substandard and garbage. I explicitly demand an 85% refund for the speech-to-text services back to my account to settle this dispute. Alternatively, you can refund the entire amount, which would also be fair, reasonable, and just, as server overload was noticeable there as well. Perhaps there is a bigger issue with Slovak and Czech than with English, as these languages are more complex and put more load on the servers.
To protect Google financially, so that I do not have to constantly dispute these services and Google does not lose money, you should logically fix these services. They should work without constantly showing service overloads, even in the paid version. Even then, it made various, often unbelievably primitive errors; I did not like it at all, and there were many mistakes in the transcript. Although I was actually running this on the Gemini 3.7 version, it felt as though I were running it on the 2.5 Lite version. The text instructions and basic formatting guidelines I provided for the transcript were not followed. My prompt requested timestamps approximately every 15 to 30 seconds. Sometimes it followed these instructions reasonably well, but other times it placed timestamps every 5 seconds, every 3 seconds, or even every 2 seconds. This is a major issue, and that was with Thinking level set to Low; otherwise, the server completely froze, hung, and failed to complete the task. Furthermore, there is a timeout—I believe of 10 minutes—which is another major issue: if a timeout occurs, the task fails to finish and you are just wasting tokens. I spent money needlessly because Google has technical problems and overloaded servers. This is a serious issue that needs to be addressed because I cannot pay for poorly functioning services.
Failures in Processing Entire Audio Segments and Proposal to Enable System Logs
It also happened, particularly in the case of Czech where I processed exceptionally long recordings totaling up to 24 hours split into half-hour segments, that in one instance an entire half-hour segment was not processed at all. It was unable to generate any output from the audio file. I checked whether the audio file was fine, and it was; all the others went through without issues, but this one half-hour segment came back completely empty. There was unequivocally an error.
It is very important that system logs are kept. If Google records logs, you will certainly find the reason why such a case occurred. If Google does not keep logs, that is no longer my problem, and you will have to deal with it yourselves. Alternatively, you can enable logging specifically for my account, but not at my expense—you should cover the cost yourselves since it is an additional service. If it is covered at Google’s own expense, I fully agree to it. It will also facilitate dispute resolution so you know that such an outage occurred. I was at my wit’s end trying to solve this because it would not work via the API at all. In Slovak, there was even a segment of about 12 to 13 minutes that did not work through the API at all, and I had to insert it via the graphical interface (UI). While all other files worked fine, this one glitched and failed. This is another bug report that needs to be resolved.
Strengthening Server Hardware Capacity and Management Know-How Proposals
It is absolutely necessary to strengthen the servers and increase hardware capacity so that I do not have to set Thinking level to Low and throttle the AI’s performance, ensuring the system works properly. This is your task. Please contact the highest-level managers possible at Google and forward them my proposal and know-how on how these issues could be resolved. I am a very good, merciful, and kind person, which is why I do not just want to criticize, but also provide solutions. I possess top-tier management skills and know-how, and I have come up with a brilliant idea regarding the managerial actions managers must take to solve this problem. As an experienced IT professional, I am able to provide you with very high-quality advice.
Diagnostics, Server Load Monitoring, and Cluster Optimization
First and foremost, high-quality diagnostic tools must be implemented, and an audit and inspection of diagnostic tools regarding server load must be conducted. An analysis should be done on whether this was running, or normally runs, in the United States on their servers that handle speech-to-text or others. So, very high-quality tools must be created. If these tools are not of sufficient quality, the engineering team must naturally create an application, code an application that will properly monitor the load. And of course, you first need diagnostic data so that you can then find a solution.
Furthermore, from this diagnostic data, a solution must be devised to distribute the server load evenly, preventing a situation where one server cluster runs at 50% capacity while another cluster runs at 100% or 110% capacity. The system should be able to balance the load evenly and optimize utilization so that the potential of these servers is maximized to the fullest extent. Another option is to purchase additional new servers for AI workloads.
Strategic Solution via Custom ASIC and NPU Chips
And then another solution, which is truly a strategic one, is to create highly specific custom chips that will help minimize energy consumption. Minimizing energy consumption will in turn reduce cooling costs for Google’s data centers. This can be accomplished by creating a highly specific custom chip… Of course, experienced engineers capable of designing such chips from scratch and then ordering their fabrication are very scarce. Only the largest corporations employ such engineers. Google has such engineers—specifically, they develop chips for smartphones, for Google Pixel. These same individuals are technically qualified to create specific custom chips. I believe they are called ASIC chips, specialized custom-made chips to keep power consumption as low as possible, and this way Google could save billions.
In the short term, there are initial capital expenditures on both sides, but in the long term, it is highly profitable and the investment will pay off very quickly. However, such an endeavor takes perhaps one and a half to two years to design and subsequently start fabrication of these chips, which will be dedicated strictly as a subcategory of NPU chips for speech-to-text, minimizing server overload and excessive load.
Naturally, a proper feasibility analysis must first be conducted to determine whether it would make sense to manufacture such highly specific custom chips, a subcategory of NPU chips, exclusively for speech-to-text purposes. When these servers are overloaded, one must think strategically about what will happen five years down the road. You must plan five years into the future and start designing such things now. Top Google executives should think strategically not just about tomorrow, but about five years from now. These are strategic plans to optimize profits while simultaneously reducing costs, and so forth.
Dedicated Monitoring Personnel and the Value of Training Data
This is also a very important matter to prevent Google from running at a loss on these operations when I am unfortunately forced to dispute these charges because it does not work at all. That is an issue. It works best with recordings or audio files under 14 minutes, but when files are larger, the servers cannot keep up, fail, and produce nothing but confusion.
Furthermore, it is very important that Google dedicate at least two employees whose sole responsibility will be to monitor server loads and manage this issue, preventing occurrences where valuable training data is lost because the free version is constantly overloaded 95% of the time and completely unusable. I reject this; this is an improper procedure. Such issues cause severe damage and massive financial losses to Google when it loses data worth its weight in gold. The Opus files and training data that I uploaded to Google’s servers are worth millions of euros. Such data should not be disregarded. Secondly, when it comes to the paid version, it must truly function without issues and errors.
Importance of Slovak and Czech and Bug Fixes
I do not think it is acceptable that the maximum file length must be around 13 minutes. I consider that problematic. It is worth considering adding a geographic server targeting option to the API. If that would help, I would have no issue routing requests to US servers, but the API lacks this option—and I assume US servers are the most robust and have the highest hardware performance compared to European servers.
It is also very important to mention that while collecting training data in English might not be as critical, gathering training data in Slovak is immensely valuable. In English—understandably, as the most widely spoken language—there is an abundance of training data, but for Slovak, Czech, etc., training data is worth gold. Do not disregard it; appreciate this training data. There must be a free version, and with the free version, data is automatically used for training.
It is also a great mystery that I truly do not understand: why with certain files that are identical to the others, everything works smoothly and is technically sound, while with others you get a completely blank output with no text output for speech-to-text. That is a mystery and clearly a bug, but that is not my problem; Google’s engineers must resolve it. Please forward this issue to them; the problem has been reported. For Google not to lose money, it is vital to fix these issues so that a substandard service becomes a high-quality one, which is fully achievable.
Slovak Transcription Error Rate and Time Burden on Podcasts
Furthermore, what is very important to add is that the transcription in Slovak still produces quite a few errors. It is not completely terrible—perhaps one error every 10 minutes of audio—but it is still very inconvenient when I have to correct it. I certainly want to save time, and it is burdensome since I often produce podcasts lasting several hours, so having to correct mistakes every 10 minutes is very annoying when the AI fails at certain things.
Significance of Training Data for Google AI and the Server Overload Issue
I do not want to see this:
503 UNAVAILABLE. {'error': {'code': 503, 'message': 'This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.', 'status': 'UNAVAILABLE'}}
This can be solved precisely by having sufficient training data. With enough training data, these issues will gradually be resolved in newer versions. However, what I previously mentioned must not happen: Google must not, through such foolish errors, lose what is worth gold—training data. It immensely harms Google—costing millions of euros in damages—when it disregards training data like this, preventing me from submitting training data even if I wanted to, simply because it displays: “error, server overloaded, unable to process.” I do not accept this.