Distant Reading
Franco Moretti's provocation: nobody has read the nineteenth-century novel. People have read a canon — a few hundred books out of tens of thousands published, selected by a filtering process that ran for a century and had opinions. Any claim about "the Victorian novel" is really a claim about the survivors of that filter.
Distant reading is the response: don't read the books, count them. Chart genre lifespans, title lengths, character networks, word frequencies across thousands of texts at once. You lose everything that makes a novel worth reading and gain the only thing close reading structurally cannot see — the shape of the whole.
The two are not rivals. Close reading answers what a text means; distant reading answers what was typical, when a pattern began, and what the canon left out. Asking either to do the other's job produces nonsense.
What Computation Is Good At
| Task | Why the machine wins |
|---|---|
| Counting | Word frequencies across 50,000 texts; no human can |
| Linking | Record linkage across censuses, registers, shipping lists |
| Mapping | Plotting thousands of events; patterns emerge spatially |
| Networks | Who corresponded with whom, at a scale no reader can hold |
| Finding the odd | Anomalies in a mass no one could ever eyeball |
Real results follow: correspondence networks of the Republic of Letters, trans-Atlantic slave-voyage databases that turned scattered manifests into a countable system, name-linked censuses that made nineteenth-century mobility measurable for the first time.
What Computation Cannot See
Digitization is a new filter stacked on the old ones. The archive already filtered by recorded/kept/survived; digitization now adds scanned, and scanning follows money, copyright, institutional prestige, and OCR-friendliness. So the digital corpus is not the archive — it is the well-funded, out-of-copyright, cleanly-printed corner of it, and searching it feels exactly like searching everything.
The specific traps:
- The OCR floor. Optical character recognition fails on damaged print, gothic type, and handwriting. Your "complete search" of a newspaper run silently skips the smudged pages — and smudged is not random.
- Searchability bias. You find what you can search for. Concepts without a keyword — and phenomena people had no word for — are invisible to full-text search, which quietly reshapes research toward whatever happens to be named.
- Metadata is an argument. Every category in a database was a decision. Fields for gender, race, occupation, and status impose a classification the sources may not share, and once entered it hardens into a fact and gets counted.
- False precision. "1,847 mentions" looks like a measurement. It is a measurement of a corpus, and the corpus is a sample nobody designed.
The rule that keeps this honest: a digital method changes the scale, not the epistemology. Every question you would ask a document — who made it, why, what is missing, what is it for — you must ask the corpus, the OCR, and the schema. A number produced from an unexamined corpus is not evidence; it is a rumour with a decimal point.
The productive loop is neither pure: distant reading finds the anomaly, close reading explains it, and the explanation suggests the next thing to count. The machine finds where to look. It does not know what it found.