Wow, this wasn't on my radar at all.
"OpenAI has struck similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI. Oxford is the only UK member of the project."
From the article:
The OpenAI contract with Oxford also raises the prospect of the mass digitisation of the Bodleian’s collection consisting of 23m items. The minutes also discussed the creation of an “Ask the Bod” chatbot.
A spokesperson for the University of Oxford said the amount of text being digitised was “modest in scale” and covered only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said.
From me: but what does the contract say? I can't tell from reading this if "the contract raises the prospect" means that the contract gives OpenAI the obligation to digitize all the Bodleian's out-of-copyright works and the right to use it all for training models or if that is merely an unfounded concern raised in the staff meeting whose minutes were got via a freedom of information request.
The second paragraph doesn't settle the issue in my mind - "modest in scale" is not specific at all, and the continuous aspect could mean that whatever "modest in scale" means, it applied right then at the time of the interview, and it might stop applying later on.