The University of Oxford entered a partnership with OpenAI in March 2025 that allows the AI firm to use digitised material from the Bodleian Library to train its large-language models. The agreement was presented as a digitisation project intended to make historic texts more accessible to students and researchers. Internal OpenAI documents, however, confirm that the scanned content has been added to the company’s training set.
An OpenAI spokesperson told the press that the company is ‘proud’ to ensure today’s AI models preserve the world’s historical knowledge for the future, adding that the technology’s billion-plus daily users require representation of diverse cultures, histories and perspectives. The spokesperson’s remarks were issued alongside the university’s announcement, linking the digitisation effort to broader goals of cultural inclusivity in AI outputs.
Minutes from a University of Oxford meeting, obtained through a freedom-of-information request, record staff worries about the partnership’s reputational risk and its clash with the university’s environmental commitments, given the energy-intensive nature of large-scale model training. Members of the Bodleian governance committee expressed particular concern that associating the library with a commercial AI developer could affect public perception of the institution.
Secondhand booksellers have reported a sudden rise in orders for obscure titles such as an 18th-century guide to African agricultural implements and biographies of 1950s race-car drivers. Owners of used-book shops speculate that these rare, non-digitised works are being sought as fresh data sources for the next generation of AI models, as web-scraped content becomes increasingly saturated with AI-generated text.
OpenAI has forged comparable deals with several U.S. research libraries, including the Boston Public Library, Caltech, MIT and the University of Michigan, under the NextGenAI programme. Oxford remains the sole UK participant in that initiative, extending the model-training network beyond American institutions and highlighting the growing demand for historical print collections in AI development.
By June 2025 the Bodleian had supplied OpenAI with 125,000 scanned images of historical dissertations, covering 19th- and 20th-century PhD theses from European and American universities. The digitisation effort also encompassed a rare set of 10,000 sixteenth-century broadside ballads that combine lyrics with musical notation. Staff discussions have mentioned extending the work to eighteenth-century Irish state papers, private letters of novelist Marie Edgeworth and Dorothy Hodgkin’s penicillin notebooks.
The contract opens the possibility of mass digitising the Bodleian’s 23 million-item collection, and meeting minutes reference plans for an ‘Ask the Bod’ chatbot that would query the library’s holdings. A university spokesperson emphasized that the material being digitised is modest in scale, out-of-copyright, and that the Bodleian retains rights to publish the scans openly online within months, with OpenAI’s use being non-exclusive.
Unlike the Bodleian’s approach, rival AI firm Anthropic has reportedly spent tens of millions acquiring books only to slice and pulp them after scanning, though the company claims it does not destroy rare volumes. A 404 Media investigation placed a tracking device in a secondhand book order and traced it to an Amazon facility in the United States where the books were dismantled and digitised, underscoring divergent industry practices.