The University of Oxford allowed OpenAI to ingest historical texts from the Bodleian Library — one of the oldest research libraries in the world — into its model training pipeline. The arrangement was announced in March 2025 as a digitisation partnership that would make rare texts more accessible. Internal documents obtained via freedom of information request reveal the scanned material was used to "populate the OpenAI training set." The training-data element was not mentioned in the public announcement. By June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI, including PhD theses from European and American universities written in the 19th and 20th centuries, a collection of 10,000 16th-century broadside ballads, and discussions were underway about digitising 18th-century Irish state papers, letters of novelist Marie Edgeworth, and Dorothy Hodgkin's penicillin notebooks. The contract also raises the prospect of mass digitisation across the Bodleian's 23 million items. Oxford is the sole UK member of OpenAI's NextGenAI project, which has struck similar agreements with Boston Public Library, Caltech, MIT, and the University of Michigan. The pattern is clear: as scraped websites become saturated with AI-generated slop, physical book and manuscript collections represent the last reservoir of clean, high-quality, human-authored training data. The Bodleian's centuries of accumulated knowledge is exactly what model developers need. The university insists the digitised material is "modest in scale," covers only out-of-copyright works, and that the Bodleian retains rights to the scans, which will be published openly online within months. Staff were said to have been "open" that the project would contribute training data, though internal meeting minutes record concerns from Bodleian governance committee members about reputational risk and the environmental cost of energy-intensive AI infrastructure. The broader context makes the deal look less benign. Anthropic has spent tens of millions of dollars buying secondhand books, slicing off their spines for scanning, then pulping them. A tracking device placed in a secondhand book order by 404 Media traced it to an Amazon facility where books were dismantled and scanned. Booksellers report mysterious bulk orders for obscure titles unlikely to exist in digitised form — agricultural implement guides, 1950s motor-racing biographies — that appear designed to harvest fresh training data. Oxford's deal is gentler than spine-slicing, but it operates on the same logic: physical collections as untapped data mines. The fundamental asymmetry is structural. Oxford gets digitisation services it could have funded independently or through public grants. OpenAI gets irreplaceable training data from a collection built over 423 years by scholars, donors, and public investment — data that will be baked permanently into commercial models generating billions in revenue. The university spokesperson's framing — "make the material more accessible to a wider number of people" — describes a side effect, not the commercial core of the transaction. Whether this is a savvy partnership or a giveaway depends on what Oxford negotiated beyond free scanning. The FOI documents suggest internal discomfort but not a renegotiation. The deal's real significance is as a template: if Oxford trades its collection for digitisation services, every research library in the world will face the same pitch, and most will have less leverage to negotiate.