The University of Oxford announced in March 2025 a collaboration with the creator of ChatGPT to apply OpenAI software to digitise holdings from the Bodleian Library. The university said the effort would broaden access for scholars and students by creating digital copies of historic works. The public statement did not mention that the resulting files would also be fed into OpenAI’s machine-learning pipelines.
Internal records indicate that the scanned pages have been incorporated into OpenAI’s training corpus, a process described as ‘populating the OpenAI training set.’ An OpenAI spokesperson told the press the company is ‘proud’ to ensure ‘the AI models of today preserve the world’s historical knowledge for the future,’ adding that with over a billion daily users the technology must reflect diverse cultures, histories and viewpoints.
Minutes obtained through a freedom-of-information request reveal faculty worries about the partnership’s impact on the university’s reputation and its sustainability commitments, given the high energy consumption of large-scale AI training. The documents also note discussions about an “Ask the Bod” conversational agent and the possibility of extending digitisation to the Bodleian’s 23 million-item collection, raising further ethical and logistical questions.
By June 2025 the project had supplied OpenAI with roughly 125 000 scanned images of historical dissertations, including nineteenth- and twentieth-century PhD theses from European and American institutions. Additional material comprises a rare set of about 10 000 sixteenth-century broadside ballads featuring lyrics and musical notation, and staff have proposed adding eighteenth-century Irish state papers, private correspondence of novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin laboratory notes to the digitised pool.
University officials stress that the current digitisation effort is limited in size, focuses solely on works no longer under copyright, and that the Bodleian retains ownership of the scans. The library plans to release the digital files openly within months, and it rejects claims that the machine-learning aspect was concealed from students or the public. This approach contrasts with practices reported at other firms that acquire books, dismantle them and then pulp the originals.
Second-hand booksellers have reported a surge in orders for obscure titles such as an eighteenth-century African agricultural guide and biographies of 1950s motor-sport figures, items that are unlikely to exist in digital form elsewhere. Industry observers note that as web content becomes saturated with AI-generated text, developers are increasingly turning to physical, historic collections for fresh training data, a trend that has drawn both interest and criticism.