AI NewsProductivityReported
An engineer searched 4 million archive pages and found a lost meteorite
Software engineer Jesse Waites built a pipeline that reads millions of handwritten archive pages with a small decision model as the first filter, and reported it found a meteorite fall and three lost rhinos.

Image: Jesse Waites
Why it mattersPutting a cheap decision model in front of a large one, as the first reader of noisy data, is a pattern that turns a workload too big for a reasoning model into one that fits a hobbyist's budget.
Reading one of the world's largest handwritten archives at a cost a hobbyist can pay turns out to need two models, not one.
Jesse Waites, a software engineer, published an account of a research pipeline he built with Claude Code to search four hundred years of Dutch East India Company papers and digitised newspapers for forgotten events. The pipeline found a letter about a meteorite that fell in Maharashtra in 1812, three live rhinoceroses that the Company shipped from Batavia and never delivered to the King of Kandy, and references to volcanic eruptions missing from the standard catalogues. Each finding is tied back to a scan of the original page.
The scale, and the trick that made it affordable
The Dutch East India Company archive runs to 4.35 million pages, which the GLOBALISE project has turned into searchable text. Waites says reading those pages himself at two minutes a page, five days a week, would take him about 70 years. A home GPU in his office turned every passage into a semantic fingerprint in a single overnight run, 5.7 million passages from the Company archive alone, so he could search for what a passage was about rather than its exact words. That mattered because 17th-century Dutch spells "rhinoceros" about 15 different ways.
Semantic search still returns tens of thousands of hits per query, which is more than a large model can read affordably. So the first reader of each hit was TypeSafe's Jev, a small System One decision model, which Waites says answers one narrow question at a time at a few cents per million words. Having Jev read 59,000 mentions of elephants cost him about three dollars. Only the passages Jev flagged were passed to Claude Haiku for a close read and translation, and then to the Claude Code agent to open the original handwritten scan and check the transcription against it.
What he found
The meteorite report comes from the Java Government Gazette, reprinting a letter from the Bombay Gazette: a British officer at Pandharpur described a heavy iron-rich stone with a thin black crust, buried a foot deep in open ground after a sound "like a rustling fire of Musquetry". Waites says the fall is not in any of six catalogues he checked, from Chladni's 1819 list through the British Museum catalogues and the 1933 List of Indian Meteorites to today's Meteoritical Bulletin, and the stone is not in the Natural History Museum's collection. Once confirmed, he says it would be the earliest recorded meteorite fall in Maharashtra.
The rhinos sit in the Company's own ledgers. One live Javan rhinoceros, caught near Batavia in 1738, reached Colombo and died in the governor's horse stable the following year. Two more sailed on the Loverendaal three days later and died at sea, according to a sworn statement by the officers before the ship's bookkeeper.
The control that stopped the pipeline fooling itself
The method step worth copying is what Waites did before trusting any new find. The pipeline had to re-find Benjamin Breen's 1615 dodo sighting, the Laki eruption, Tambora, and a dozen other known events before the search for new ones began. A pipeline that cannot find a known answer in known data has nothing to say when it finds nothing new.
Waites also stayed involved throughout, decided which leads to pursue, and recognised when a search was running out. The models read at a scale he cannot. The judgement of where to go next and when to stop is still his.
Source
I Pointed AI at 400 Years of Historical Archives, Jesse Waites.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


