This site contains affiliate links. We may earn a commission at no extra cost to you.
GuideWhy OCR Fails on Manga: Sound Effects, Handwriting and What Works
August 8, 2026
Translation apps have got genuinely good. Point a phone at a Japanese menu and you get a usable translation in real time. Point the same phone at a page of manga and the results are wildly inconsistent — clean on some panels, blank or nonsense on others, and often worst exactly where the story is.
This is not the app being bad. It is a mismatch between what these tools are built to read and what a manga page actually contains. Understanding the mismatch tells you which works will work, which won’t, and what to do differently.
What OCR is actually doing
Optical character recognition finds shapes in an image and decides which characters they are. It was built for documents — printed text, dark on light, in horizontal rows, in a consistent typeface.
Every one of those assumptions is a design decision, and every one of them is violated somewhere on a manga page.
Consistent typeface. OCR works by matching shapes against learned forms. Printed text varies little; hand-drawn text varies enormously.
Dark on light. Contrast is what separates a character from its background. Manga panels frequently place text over artwork, screentone, or black.
Horizontal rows. Japanese is often set vertically, and manga mixes both on the same page. Better tools handle this; many do not.
Isolated from imagery. A document has text and margins. A manga page has text sitting inside a drawing, sometimes overlapping it.
Speech bubbles happen to satisfy all four assumptions almost perfectly — which is why bubble dialogue works, and why everything else is a lottery.
The four places text lives on a page
Sorting a page by where its text sits predicts your results better than any app comparison.
Inside speech bubbles. Printed typeface, white background, clean edges, isolated from artwork. This is the best case and it usually works. If a work keeps its text here, you can read it with an app.
Thought text and small asides. Often handwritten, often smaller, often placed outside a bubble against artwork. Recognition drops sharply. These lines frequently carry the internal narration that gives a scene its meaning.
Sound effects drawn into the art. Stylised, distorted, sometimes stretched across a whole panel or integrated into an object. OCR essentially never handles these, and no app is close to solving it. They are drawings that happen to be letters.
Text on the artwork itself — signage, phone screens, handwritten notes as props. Variable, and usually unimportant. Occasionally a plot point.
The rough hierarchy: bubbles work, asides are unreliable, sound effects don’t work at all.
Why sound effects matter more than they seem
English-language comics carry sound effects too, so it is easy to assume they are decorative. In Japanese comics they frequently are not.
Japanese has an unusually large inventory of words describing states, textures, and manners of action — not just noises. These appear on the page as sound effects, and they do work that English comics would do with a caption or a line of dialogue: conveying atmosphere, physical sensation, the quality of a silence.
A page can carry its emotional information almost entirely in drawn text, with the bubbles handling only the literal conversation. Translate the bubbles and you get the plot while missing the register.
This is also why official translations feel different from app-assisted reading in a way that goes beyond convenience. A translator handles these; an app cannot see them. It is the single largest gap between the two experiences, and it is invisible if you have only ever read one way.
What actually improves your results
Practical, in order of how much difference they make.
Choose works with light text. By a wide margin the most effective change, and the one that requires no tooling at all. A work with four lines of dialogue has no OCR problem to solve.
Read on the largest screen available. Recognition scales with resolution. A page viewed full-size on a monitor gives an app far more to work with than the same page on a phone, and you can see the composition while you’re at it.
Zoom before capturing. Filling the frame with one bubble rather than the whole page improves accuracy substantially. Slower, and worth it on anything you care about.
Read the whole page before translating anything. Manga is composed as a page, and the layout carries a large share of the meaning — who is where, what is happening in the background, what the panel sizes are doing to the pacing. If you translate bubble by bubble from the top, you assemble the page out of fragments and never see it whole. One pass for the picture, then one pass for the words, is both faster and better.
Try more than one app. They fail differently. If one produces nonsense on a panel, another sometimes handles it — particularly for vertical text, where support varies.
Accept the sound effects. Trying to force them is the biggest time sink available. Read them as visual texture, which is a large part of what they are, and move on.
What does not help: better lighting, a newer phone, or any tool that operates on text rather than images. The Japanese in a manga page is not text in any sense a computer recognises — it is pixels arranged to look like characters, exactly as the artwork is pixels arranged to look like people.
The things people try that don’t work
Worth naming, because each one costs an afternoon to discover independently.
Uploading a page to a general-purpose AI assistant. Newer, and it fails in a more confusing way than OCR does. These tools will produce fluent, plausible output for a page they have only partly read — filling gaps with what the scene looks like it should say. The result reads better than a machine translation and is less reliable, because you cannot tell which lines were read and which were inferred. For a story built on a specific revelation, that is the worst possible failure mode.
Whole-page translation tools. These attempt to detect every text region and overlay translations in place. When they work, they are genuinely pleasant. They work best on exactly the material that was already easiest — clean printed bubbles — and they fail on the same handwriting and sound effects as everything else, except now the failures are hidden under a confident-looking overlay.
Turning up the contrast or “cleaning” the page first. Marginal at best. The failure is not that the characters are unclear; it is that they are stylised beyond the shapes the model learned. Sharpening a hand-drawn sound effect gives you a sharper drawing.
Asking in a forum what a specific panel says. Works, occasionally, and does not scale past a page or two.
Vertical text is the one variable worth checking
Japanese is set vertically in most comics — columns read top to bottom, right to left. Some recognition tools handle this natively; some assume horizontal rows and return garbled output that looks like a bad translation rather than a layout failure.
If an app produces consistent nonsense on clean, legible bubbles, this is usually why, and the fix is a different app rather than a different page. It is worth testing once on a work you can verify, so that you know which of your tools is doing what.
Judging a work before you buy it
Everything above is downstream of one decision: which work you bought. That decision is where the leverage is, and you can make it from the product page.
Open the previews and look at where the text sits. Not what it says — you cannot read it — but where it is. Text neatly contained in bubbles means an app will cope. Text scattered across artwork in a loose hand means it will not.
Check whether the bubbles are printed or hand-lettered. Printed lettering is the friendly case and is what most commercial-scale circles use. Hand-lettering inside bubbles is common in smaller releases and sits somewhere in between — legible to a person, unreliable to a tool.
Count the bubbles per page. Two or three is comfortable. Eight is a page you will spend several minutes on.
Look for text in the margins. Handwritten commentary crammed around panels is a strong signal of a wordy work, and it is the exact category OCR handles worst.
Check the format. Illustration sets carry less text than manga as a rule, and novels are pure text.
Look at the panel density too. Pages with many small panels tend to carry more dialogue than pages with a few large ones, simply because there are more moments to narrate. A work laid out in big, open panels is usually a work that trusts its art, and works that trust their art need fewer words.
This takes about fifteen seconds and decides whether an app-assisted read will be pleasant or exhausting.
What good results actually look like
Worth calibrating, because “it works” means different things to different people.
A good session looks like: you read a page for the art, point the app at three or four bubbles, get lines that clearly make sense in context, and move on. Perhaps one line in ten is ambiguous and you infer it from the panel. The reading takes maybe twice as long as it would in your own language.
A bad session looks like: you point the app at a bubble and get a fragment. You point it at the next and get something that contradicts the first. You start zooming, rotating, re-capturing. Twenty minutes in you have read four pages and lost the thread of the story.
The difference between those two sessions is almost entirely the work, not the app, not the phone, and not your technique. The same tool that produced the first session produces the second on a different book.
Which is worth knowing, because the natural response to a bad session is to go looking for a better app — and the effective response is to go looking for a different work.
The honest summary
For clean bubble dialogue, translation apps work well enough that the language barrier is a speed bump rather than a wall. For handwritten asides they are unreliable. For sound effects they do not function, and no forthcoming version will change that, because the problem is not recognition quality — it is that the text has been absorbed into the drawing.
Which means the real decision is not which app but which works. A well-chosen work needs almost none of this. A poorly chosen one cannot be rescued by any tool currently available.
That is the judgement recorded against every work on this site: how much Japanese you actually need, and why — including whether the text sits somewhere a tool can reach. It is the one thing no storefront records, and it is the thing that decides whether the evening goes well.
Related reading: Doujinshi With Almost No Text · What You Actually Get When You Buy a Doujinshi · Silent Manga and Doujinshi
When OCR is not the answer
If a work defeats OCR, the fix is usually to pick a different work rather than a different app:
Also worth reading: how to read untranslated doujinshi