Distant Reading and Computation
Nobody has read the nineteenth-century novel. Any claim about it is based on a canon, a few hundred survivor books selected by a century of filtering. Distant reading is the response: don't read the books, count them. You lose the nuance of individual texts to see the shape of the whole.
| Task | Why the machine wins |
|---|---|
| Counting | Word frequencies across 50,000 texts; no human can |
| Linking | Record linkage across censuses, registers, shipping lists |
| Mapping | Plotting thousands of events; patterns emerge spatially |
| Networks | Who corresponded with whom, at scale |
Distant reading and close reading are not rivals. Close reading answers what a text means; distant reading answers what was typical, when a pattern began, and what the canon left out.
What Computation Cannot See
Digitization is a new filter stacked on the old ones. The digital corpus is just the well-funded, out-of-copyright, cleanly-printed corner of history. Watch out for specific traps:
- The OCR floor: Optical character recognition fails on damaged print, gothic type, and handwriting.
- Searchability bias: Concepts without a keyword are invisible to full-text search.
- Metadata is an argument: Database categories impose a classification the sources may not share.
- False precision: A number from an unexamined corpus is not evidence; it is a rumor with a decimal point.
The productive loop: distant reading finds the anomaly, close reading explains it, and the explanation suggests the next thing to count. The machine finds where to look, but it does not know what it found.