The Readability Dataset

Every work on this site carries a judgement of how much Japanese you need to enjoy it. This page explains how that judgement is made, publishes the underlying figures, and states plainly what the method cannot tell you.

Why this exists

DLsite records a great deal about a work — circle, genre tags, price, release date, file format. It does not record how much text is on the page, because for the audience the store was built for, that is not a variable worth recording. A Japanese reader does not need to be told whether a manga has dialogue.

For a reader who cannot read Japanese fluently, it is the single most decisive fact about a purchase. A work you cannot read costs you one hundred per cent of what you paid, whatever you paid. So we measure it ourselves.

How it is measured

Every work on DLsite has free sample pages. We run optical character recognition over those pages and record three numbers, and only three:

  • Characters per page — how much text the OCR extracted
  • Text regions per page — how many separate areas the OCR identified as containing writing
  • Confidence — how certain the OCR was about what it read

We do not read the contents. The measurement is deliberately blind to what the work is about; it only asks how much writing sits on the page and how machine-readable that writing is.

The four bands

Band What it means for you
English An official or community English edition exists. No Japanese needed.
Minimal text Very little writing. The artwork carries the story.
OCR OK There is text, and it is typeset cleanly enough that a translation app genuinely cooperates.
Text-heavy Dialogue carries the work, or the lettering is hand-drawn. You want real reading ability.

The correction that mattered

Our first attempt used character count alone. It put more than half the catalogue into Minimal text, which is useless — when half the shelf carries the same label, the label has stopped telling you anything.

The fix came from the region count. Low-confidence works turn out to be two different animals:

  • About one character per page and half a text region per page — genuinely no text
  • Six or seven characters per page across two or three regions — text that the OCR could not read

When the OCR finds regions but cannot turn them into characters, it is looking at handwriting or drawn lettering. That is harder for a human learner too, not easier. Works matching that pattern now fall to Text-heavy.

Where a work sits near a boundary, we push it toward Text-heavy. Telling someone a work is easy when it is not costs them money and trust. The opposite error costs us a sale. Those are not equivalent.

What this does not cover

  • It describes this catalogue only. Everything here is drawn from the NTR catalogue we cover, not from DLsite as a whole.
  • The English band is not OCR-measured. It records that a translated edition exists — a fact about the store, not about the page.
  • Sample pages are not the whole work. A work can open quietly and turn talkative.
  • Unreleased works are absent. Announced titles have no sample pages to measure.
  • OCR is not comprehension. “A translation app cooperates” is not the same as “you will follow the story”.

The data

The full classification is published as a CSV file. You are welcome to use it, including commercially, with attribution — CC BY 4.0.

Download readability.csv

Column Contents
rj The work’s DLsite product number
readability One of the four bands
reason The measured figures behind the judgement, in plain English
work_type Manga, illustration collection, novel, and so on
pages Page count where the circle stated one
released Release date
has_english Whether an English edition exists
url The work’s page on this site

The reason column is the point. Every judgement shows its own working, so you can disagree with where we drew the line rather than having to take our word for it.

If you use this in your own work and something looks wrong, tell us. A measurement nobody can check is not a measurement.