You take a photo, run it through a text reader and get back something that looks almost right, except that a price has the wrong digit, a word is split in two, or a whole line has vanished. It is frustrating, and it can look random. It is not. Text recognition mistakes follow patterns, and once you know the patterns you can usually predict them, prevent them and fix them.
This article is a practical troubleshooting guide. We go through the most common reasons OCR misreads text, show you how to recognize each one from the result and give you a specific fix. If you want the basics first, start with what OCR is and how it works.
First, read the result like a detective
Before you change anything, look at what kind of mistake you got. The type of error points to the cause.
| What you see | Most likely cause | Go to |
|---|---|---|
| Random symbols and nonsense words | Wrong language or a very blurry image | Causes 1 and 2 |
| One or two wrong characters in otherwise good text | Look-alike characters or small print | Causes 3 and 4 |
| Missing lines or words | Low contrast, glare or cropping | Causes 5 and 6 |
| Lines in the wrong order | Columns, tables or tilted text | Causes 7 and 8 |
| Letters merged or split by spaces | Tight or wide letter spacing, decorative fonts | Cause 9 |
| Handwriting turns into garbage | Handwriting is hard for any reader | Cause 10 |
Cause 1: The wrong language is selected
This is the most common cause of complete nonsense, and the easiest to fix. A reader uses language knowledge to choose between similar shapes and to correct likely mistakes. If you tell it the text is German when it is Spanish, it may “correct” real Spanish words into German looking ones, and it will not expect accents in the right places.
How to spot it: the output looks like words from another language, or is full of unusual symbols.
Fix: change the source language and run it again. If you do not know the language, look for clues such as special letters, a different alphabet or characters instead of letters.
Cause 2: The image is blurry
Blur happens when the camera is not focused, when your hand moves, or when the picture has been enlarged from a small version. Edges of letters melt into the background, so the reader cannot tell an e from a c or an o.
How to spot it: many small errors across the whole text, not just a few.
Fix: tap the text on your phone to focus, hold the phone with both hands or rest it on something, and move closer instead of zooming digitally. If you are working with an existing image, find the original rather than a shrunken copy.
Cause 3: Look-alike characters
Some characters have nearly identical shapes. The reader picks the more likely one based on context, and sometimes guesses wrong.
| Often swapped | Where it hurts |
|---|---|
| 0 (zero) and O (capital letter) | Codes, serial numbers, passwords |
| 1 (one), l (lowercase L) and I (capital i) | Prices, IDs, names |
| 5 and S | Amounts, product codes |
| 8 and B | Measurements, codes |
| rn and m | Ordinary words |
| cl and d | Ordinary words |
Fix: there is no way to remove this risk completely, so check numbers and codes against the picture by eye. Wherever a wrong character would matter, such as a price, a date, a medicine dose or a reference number, confirm it manually.
Cause 4: The print is too small
If each letter is only a few pixels tall, there is not enough detail to separate similar shapes. Small print on labels, receipts and contracts is the classic example.
Fix: get closer so the text fills the frame. If you cannot, take several photos of smaller sections. For screenshots, zoom in before capturing.
Cause 5: Low contrast
Light gray on white, dark blue on black, or text over a patterned background all give the reader very little difference between letter and background. Faded thermal receipts are a well known example because the ink fades over time.
How to spot it: missing words or lines, mostly in the faint areas.
Fix: add light, raise the brightness, or increase the contrast in a photo editor before uploading. For screenshots, try a different theme or increase the app’s text contrast.
Cause 6: Glare, shadows and reflections
A bright reflection on glossy paper, laminated menus or glass can wipe out a whole region of text. Shadows from your own hand have a similar effect from the other direction.
Fix: change your position or tilt the object until the reflection moves. Turn off the flash, which usually makes glare worse. Use soft, even light from the side or above.
Cause 7: Columns and tables
Readers decide the order of text by looking at its layout. On a page with several columns, or a table with many cells, the engine may read straight across the page instead of down each column, mixing the lines together.
How to spot it: the sentences seem to jump between topics, or table values are separated from their labels.
Fix: crop and process one column at a time. For tables, try one row or one block at a time and rebuild the table yourself in a document.
Cause 8: Tilt and perspective
If the page is rotated or photographed from an angle, letters change shape and lines slope. The reader may break lines in the wrong place or misjudge the size of letters.
Fix: hold the camera parallel to the page. If you already have a tilted image, rotate it until lines are level before you upload it.
Cause 9: Fancy fonts and tight spacing
Decorative, condensed, handwritten style or very bold fonts look different from the standard shapes an engine expects. Letters that touch each other may be read as one shape, and widely spaced letters may be read as separate words.
How to spot it: logos and headings fail while the plain body text is fine.
Fix: there is no complete fix, because the font itself is the problem. Read headings yourself, and rely on the tool for plain body text.
Cause 10: Handwriting
Handwriting differs from person to person, and cursive connects letters in ways that are hard to separate. Even strong systems make more mistakes here than with print.
Fix: if you are the writer, use clear block letters, dark ink and plain paper. If you are reading someone else’s writing, treat the result as a rough draft and ask them to confirm important words.
A fast checklist before you upload
- Is the text sharp when you view it at full size?
- Is the light even, with no glare or shadow?
- Is the camera straight on, so lines are level?
- Is the picture cropped to the text?
- Did you pick the right source language?
- For numbers and codes, have you planned to double check them?
When the image cannot be improved
Sometimes the original is simply poor, such as an old faxed page or a photo of a distant sign. In those cases, set your expectations. Use the tool to get the gist, then rely on context. For example, you can often work out a damaged word from the words around it. If the content matters, ask for a better copy or have a person read it.
A note on translation errors
Not every strange result is an OCR problem. If the left box matches the picture but the translation sounds odd, the problem is in the translation step. Idioms, slang, abbreviations and missing context are the usual causes. In that case, search the original phrase, or ask someone who knows the language. Our guide to translating text in an image covers how the two stages fit together.
A ten minute experiment you can try
The best way to understand these causes is to cause them on purpose. Print or open a short paragraph with a few numbers in it. Then photograph it five times: once sharp and straight, once slightly blurry, once at a steep angle, once with a shadow across it and once with a hand covering the lower half. Run each picture through the translator with the correct language selected and compare the left boxes. You will see exactly how each problem changes the result, which words disappear first and which characters are swapped. After this exercise, you will be able to look at a bad result and guess the cause within seconds, which is the most useful skill in working with any text reader.
Frequently asked questions
Why does OCR confuse zero and the letter O?
They look almost identical in many fonts. The engine uses context to guess, and context is not always enough.
Can I improve a blurry image after the fact?
Sharpening tools can help a little, but they cannot recover detail that was never captured. A new photo is better.
Why is my receipt missing lines?
Thermal receipts fade and have low contrast. Photograph them on a dark surface under bright, even light and increase the contrast.
Does a higher resolution photo always help?
Usually, yes, as long as the text is in focus. Very large images can be slow, so crop to the text.
Should I trust numbers from OCR?
Verify them. A single wrong digit can change a price, a date or a dose.
Put it into practice
The quickest way to learn these patterns is to try them. Take a photo, translate it with the translator and compare the left box with the original. For situations where fast, reliable reading matters, see our guide on translating street signs while traveling and the tips for Japanese images.
