PrivacySignal

Search & browse the archive

The full corpus — beyond today's front page.

Reset

8 results

News
WIRED — AI · · International

He Scraped All of Their Art for AI. Now He’s Collaborating on a Tool to Help Them

Cara, a portfolio platform built for artists who want to keep their work out of AI training datasets, has faced coordinated attacks by bad actors scraping and publishing its data. The person responsible for an earlier mass scraping of artists' work is now reportedly working with Cara on a protective tool.

Who should care: General readers · AI governance · Policy

Enforcement
NPR — Tech · · International

Why tech companies are buying up tons of rare old books to train their AI models

An investigation by 404 Media found that rare books have been routed to an Amazon warehouse used for AI training data. The reporting suggests tech companies are acquiring physical archival materials to expand the text datasets used to build AI models.

Who should care: Lawyers · Privacy officers · Compliance · General readers · AI governance · Policy

#enforcement#ai Read original →
News
BBC — Tech · · International

Secondhand book sales are booming. Is it because of AI?

Secondhand book sellers are reporting unusual bulk purchases, with speculation that buyers are acquiring physical books to use as AI training data before destroying them. The pattern has drawn attention to how AI companies may be sourcing text at scale through unconventional channels.

Who should care: General readers · AI governance · Policy

News
EFF — Deeplinks · · International

“Stealth Crawlers” Are Not a Threat to the Open Web. Bills Targeting Them Would Be.

Legislation is being proposed that would require automated web crawlers to disclose their identity when collecting publicly available data. Critics argue such bills, backed by publishers wanting to control AI training data, would harm legitimate uses like investigative journalism, academic research, and security work.

Who should care: General readers · AI governance · Policy

GDPR / Intl
IAPP · · International

Thought for the week: Web scraping for generative AI is subject to the GDPR

A commentary from the IAPP argues that web scraping used to build generative AI training datasets falls within the scope of GDPR, meaning the collection and processing of personal data found online is not exempt simply because it occurs at scale or at the infrastructure level.

Who should care: Lawyers · Privacy officers · AI governance · General readers · Policy

News
BBC — Tech · · International

Meta halts worker tracking for AI training due to privacy fears

Meta has stopped monitoring employee computer activity for AI training purposes, a program it launched only two months ago. The decision follows concerns about privacy implications of using worker behavioral data to develop AI systems.

Who should care: Privacy officers · Cybersecurity · General readers · AI governance · Policy

#surveillance#ai#privacy Read original →
GDPR / Intl
OECD AI Policy Observatory · · International

Rethinking AI data: From scraping to sustainable and ethical data sharing

The OECD's VIADUCT initiative examines growing tensions in AI training data acquisition, questioning the sustainability of web scraping and exploring alternatives centered on ethical data-sharing frameworks that account for copyright, GDPR compliance, and equitable access.

Who should care: Lawyers · Privacy officers · AI governance · Administrators · General readers · Policy

#gdpr#ai-governance#ai Read original →