The University of Oxford has allowed OpenAI to use historical texts from its Bodleian Library to train AI models. Internal documents obtained via freedom of information requests show digitised material has been used to “populate the OpenAI training set.” The partnership was announced in March 2025 as a digitisation project that would make rare public-domain texts more searchable for students and researchers. That announcement did not state the texts would also train OpenAI models. By June 2025 about 125,000 images of historical dissertations, including 19th- and 20th-century PhD theses, had been shared. Oxford says the work is modest in scale and covers only out-of-copyright material. The Bodleian retains rights to the scans and plans to publish them openly. Meeting minutes recorded staff concerns about reputational risk and the environmental impact of the energy-intensive technology. Oxford is the only UK participant in OpenAI’s NextGenAI programme.
The University of Oxford has allowed OpenAI to use historical texts from its Bodleian Library to train AI models. Internal documents obtained via freedom of information requests show digitised material has been used to “populate the OpenAI training set.” The partnership was announced in March 2025 as a digitisation project that would make rare public-domain texts more searchable for students and researchers. That announcement did not state the texts would also train OpenAI models. By June 2025 about 125,000 images of historical dissertations, including 19th- and 20th-century PhD theses, had been shared. Oxford says the work is modest in scale and covers only out-of-copyright material. The Bodleian retains rights to the scans and plans to publish them openly. Meeting minutes recorded staff concerns about reputational risk and the environmental impact of the energy-intensive technology. Oxford is the only UK participant in OpenAI’s NextGenAI programme.