Journalism as Data: GDPR Implications of Licensing Journalistic Content to Large Language Models
Abstract
This article explores the data protection implications under the GDPR of integrating journalistic content into Large Language Models (LLMs). The number of commercial partnerships between AI companies and news publishers for the licensing of daily news and archival content has rapidly increased. We contend that while LLMs and their top-layer applications do offer innovative solutions for news dissemination, publishers should carefully evaluate their position under the GDPR. Comparing different technical solutions available, in particular pre-training, fine-tuning and Retrieval Augmented Generation (RAG), we analyse the relevant regulatory barriers and opportunities, focusing in particular on the distribution of processing roles, lawfulness and transparency of these deals, and the application of the special regime for journalistic processing under art 85(2) GDPR.
GDPR, Journalism, language models, Privacy