Hi,
I’m using 2.0-flash and 2.0-flash-lite to OCR PDF’s and images. Am I the only one who noticed that 2.0 is much better on performing those kind of tasks? Both accuracy and following instructions is much better on previous model.
Hi,
I’m using 2.0-flash and 2.0-flash-lite to OCR PDF’s and images. Am I the only one who noticed that 2.0 is much better on performing those kind of tasks? Both accuracy and following instructions is much better on previous model.
Hello,
We recommend using our 2.5 generation on models, these are significantly improved than previous generations.
Also could please share some information on this, like which models are you comparing and possible share your observation in little more detail?
TBH, I see a regression in OCR capabilities between the 2.5 and 2.0 series. The 2.5 series has problems recognizing characters; for example, ‘e’ often becomes ‘ë’ or ‘è’. There are other quirks, like instruction following. In my case, most documents have a tabular format. With an appropriate prompt and structured output, the 2.0 family works great. If I have a prompt where I specify that if, for example, there are two rows that are identical, treat them as two separate rows, 2.0 follows my instruction, but very often, 2.5 ignores me and returns one row. I know that this is a general example, but I hope it helps. If you need more details, feel free to ask. And one last question from me: are you guys working on a dedicated OCR Gemini model? I see that more and more of them are coming to the markets from other providers.
Could you please share some of the pictures where you observed this problems, so that we can try to reproduce the issue at our end?