Like other sectors that publish rich content online, libraries have been confronted with questions about how large language models (LLMs) are extracting content from databases and the web to train their models. Many libraries will have seen the impact of ‘bots’ scraping content from their websites in their stats or in reduced performance of their sites.
This rapidly evolving and highly complex landscape warrants in-depth study. CENL’s AI in Libraries Network Group therefore invites one coordinated response from each European national library. The results will inform a CENL White Paper on the creation, publication, governance and reuse of text datasets, especially their use in AI.
We are particularly interested in the relationship between datasets of texts at national libraries and commercial or non-commercial AI models. In this context, a text dataset could mean a corpus, collection export, package, API, bulk download or controlled data-access arrangement that enables computational reuse beyond ordinary item-level browsing. It may include digitised or born-digital text, automatic text transcription output, associated images, metadata, annotations and/or derived features. A text dataset could include several similar datasets based on different collections (e.g., newspapers and periodicals).
Answering the survey may take approximately 20-30 minutes, depending on how many dataset profiles are to be added. It might require coordination with colleagues to respond from an institutional perspective.
The deadline is 15 August 2026.
https://docs.google.com/forms/d/e/1FAIpQLSejv3_6tB0GxR4tCx3QQ1JWcVvsMEg4RCJw29tdy6j_fxMlMw/viewform