Courses / History I
Historical Sources

Digital Humanities and Big Data

History I 554 words Free to read

Distant Reading

Franco Moretti's provocation: nobody has read the nineteenth-century novel. People have read a canon — a few hundred books out of tens of thousands published, selected by a filtering process that ran for a century and had opinions. Any claim about "the Victorian novel" is really a claim about the survivors of that filter.

Distant reading is the response: don't read the books, count them. Chart genre lifespans, title lengths, character networks, word frequencies across thousands of texts at once. You lose everything that makes a novel worth reading and gain the only thing close reading structurally cannot see — the shape of the whole.

The two are not rivals. Close reading answers what a text means; distant reading answers what was typical, when a pattern began, and what the canon left out. Asking either to do the other's job produces nonsense.

What Computation Is Good At

TaskWhy the machine wins
CountingWord frequencies across 50,000 texts; no human can
LinkingRecord linkage across censuses, registers, shipping lists
MappingPlotting thousands of events; patterns emerge spatially
NetworksWho corresponded with whom, at a scale no reader can hold
Finding the oddAnomalies in a mass no one could ever eyeball

Real results follow: correspondence networks of the Republic of Letters, trans-Atlantic slave-voyage databases that turned scattered manifests into a countable system, name-linked censuses that made nineteenth-century mobility measurable for the first time.

What Computation Cannot See

Digitization is a new filter stacked on the old ones. The archive already filtered by recorded/kept/survived; digitization now adds scanned, and scanning follows money, copyright, institutional prestige, and OCR-friendliness. So the digital corpus is not the archive — it is the well-funded, out-of-copyright, cleanly-printed corner of it, and searching it feels exactly like searching everything.

The specific traps:

The rule that keeps this honest: a digital method changes the scale, not the epistemology. Every question you would ask a document — who made it, why, what is missing, what is it for — you must ask the corpus, the OCR, and the schema. A number produced from an unexamined corpus is not evidence; it is a rumour with a decimal point.

The productive loop is neither pure: distant reading finds the anomaly, close reading explains it, and the explanation suggests the next thing to count. The machine finds where to look. It does not know what it found.

Practise this lesson

The explanation above is free to read. The graded practice for this lesson lives in the Tryals app.

13practice questions
2interactive scenes
Start History I free

Historical Sources