Journalism as Data: GDPR Implications of Licensing Journalistic Content to Large Language Models external link

Bouchè, G. & Steketee, M.
Technology and Regulation, pp: 102-121, 2026

Abstract

This article explores the data protection implications under the GDPR of integrating journalistic content into Large Language Models (LLMs). The number of commercial partnerships between AI companies and news publishers for the licensing of daily news and archival content has rapidly increased. We contend that while LLMs and their top-layer applications do offer innovative solutions for news dissemination, publishers should carefully evaluate their position under the GDPR. Comparing different technical solutions available, in particular pre-training, fine-tuning and Retrieval Augmented Generation (RAG), we analyse the relevant regulatory barriers and opportunities, focusing in particular on the distribution of processing roles, lawfulness and transparency of these deals, and the application of the special regime for journalistic processing under art 85(2) GDPR.

GDPR, Journalism, language models, Privacy

RIS

Save .RIS

Bibtex

Save .bib

Dangerous Criminals and Beautiful Prostitutes? Investigating Harmful Representations in Dutch Language Models external link

Lin, Z., Trogrlić, G., Vreese, C.H. de & Helberger, N.
FAccT '25: Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, pp: 11005 - 1014, 2025

Abstract

While language-based AI is becoming increasingly popular, ensuring that these systems are socially responsible is essential. Despite their growing impact, large language models (LLMs), the engines of many language-driven applications, remain largely in the black box. Concerns about LLMs reinforcing harmful representations are shared by academia, industries, and the public. In professional contexts, researchers rely on LLMs for computational tasks such as text classification and contextual prediction, during which the risk of perpetuating biases cannot be overlooked. In a broader society where LLM-powered tools are widely accessible, interacting with biased models can shape public perceptions and behaviors, potentially reinforcing problematic social issues over time. This study investigates harmful representations in LLMs, focusing on ethnicity and gender in the Dutch context. Through template-based sentence construction and model probing, we identified potentially harmful representations using both automated and manual content analysis at the lexical and sentence levels, combining quantitative measurements with qualitative insights. Our findings have important ethical, legal, and political implications, challenging the acceptability of such harmful representations and emphasizing the need for effective mitigation strategies. Warning: This paper contains examples of language that some people may find offensive or upsetting.

Artificial intelligence, language models

RIS

Save .RIS

Bibtex

Save .bib