Full Breakdown
Oxford University Partners with OpenAI to Digitise Bodleian Library Collections
By Drooid · · How we work
Oxford–OpenAI Partnership Overview
The University of Oxford entered a partnership with OpenAI in March 2025 that allows the AI developer to digitise and use historical texts from the Bodleian Library for training its language models. The agreement is part of OpenAI’s “NextGenAI” initiative, which also includes collaborations with U.S. research libraries such as the Boston Public Library, Caltech, MIT, and the University of Michigan. Oxford is the sole UK participant in the programme.
Background and Scope
The Bodleian Library, housing roughly 23 million items, began scanning out-of-copyright material under the deal. By June 2025, about 125 000 images of 19th- and 20th-century dissertations and a collection of 10 000 16th-century broadside ballads had been shared with OpenAI. Staff meeting minutes obtained via a freedom-of-information request indicate that the university also discussed digitising 18th-century Irish state papers, private letters of novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin notebooks. The university emphasised that the digitisation effort is “modest in scale” and that the Bodleian will retain rights to the scans and publish them openly online in the coming months.
Official Statements & Responses
An OpenAI spokesperson stated that the company is “proud” to help preserve the world’s historical knowledge and that the models should reflect diverse cultures and perspectives. A university spokesperson characterised the amount of text being digitised as modest, confirmed that only out-of-copyright material is involved, and clarified that OpenAI’s use of the material is non-exclusive. The spokesperson also affirmed that the Bodleian will make the digitised content publicly accessible, mirroring the approach taken with other digitisation partnerships.
Staff Concerns and Criticism
Internal minutes reveal that members of the Bodleian governance committee voiced worries about reputational risk and the environmental impact of supporting an energy-intensive technology. Some staff rejected claims that the machine-learning component had been hidden from students, insisting that digitisation was the primary university interest while acknowledging that the project would also supply training data to OpenAI.
Data, Scale, and Potential Impact
The partnership could eventually enable the mass digitisation of the Bodleian’s 23 million-item collection, providing a new source of training data as developers turn away from web-scraped content that is increasingly saturated with AI-generated text. By retaining rights and publishing the scans, Oxford aims to increase public access to rare historical works while contributing to the development of large-scale language models.
