Open Editor
Privacy
7 min read

EU's New Web Scraping Rules for AI: Why Publishing Publicly Isn't Consent

The EDPB's July 2026 draft guidelines confirm GDPR applies in full to personal data scraped from the web to train AI models — with no AI exemption, and a site's own robots.txt file now carrying real legal weight.

DetailFact
Adopted (draft, version 1.0)July 2026 EDPB plenary
Public consultation deadlineOctober 30, 2026
Consent as legal basis for scraping at scaleGenerally not viable
robots.txt / ai.txt / CAPTCHAsNow treated as legally relevant signals
AI-specific exemption from GDPRNone

The European Data Protection Board adopted draft Guidelines 03/2026 on web scraping in the context of generative AI at its July 2026 plenary — the first pan-EU framework specifically addressing whether GDPR applies to personal data pulled from the open web to train AI models. It does, in full, with no carve-out for AI, and the draft is open for public comment until October 30, 2026.

The Central Clarification: Public Doesn't Mean Consented

The guidelines directly address a common assumption: that data published on a public webpage is fair game because it's already visible to anyone. The EDPB states plainly that publishing personal data publicly does not constitute consent to have it scraped and used for AI training — the two are legally distinct acts, and the second doesn't follow automatically from the first. For anyone who has ever assumed that "if it's on the internet, it's fair use," this guidance draws a clear, explicit line against that assumption.

Consent Is Effectively Off the Table at Scale

Because obtaining individual consent from every person whose data might be swept up in a large-scale scrape is practically impossible, the guidelines confirm consent generally isn't a viable legal basis for this kind of collection. That leaves legitimate interest as the primary path forward for companies scraping data to train AI — but it comes with a rigorous three-part test covering the interest itself, whether scraping is actually necessary to serve it, and a balancing test weighing that interest against the individual's rights.

A framework built around legitimate interest rather than consent shifts the burden from asking permission to justifying the collection after the fact — a company has to be able to show its reasoning holds up, not just that no one explicitly objected. That's a meaningfully higher bar than simply scraping what's publicly reachable and assuming that's sufficient.

Technical Signals Now Carry Legal Weight

The guidelines treat robots.txt files, the emerging ai.txt standard, CAPTCHAs, and login walls as relevant indicators of a data subject's or website operator's reasonable expectations under GDPR — meaning a site that technically blocks or restricts automated access has stronger legal standing than one that doesn't, even though robots.txt itself has always been advisory rather than legally binding by design. This effectively gives a website's own technical configuration a role in the legal assessment of whether scraping that site was appropriate, not just a courtesy signal crawlers may or may not respect.

What the Guidelines Openly Admit They Can't Fully Solve

Legal analysts reviewing the draft have noted it's unusually candid about its own limitations — acknowledging that a company doing large-scale scraping may not always know exactly what personal data it collected, and that what a trained model has already learned from scraped data can't currently be easily removed or "unlearned" after the fact. That second point is significant: the guidelines describe an ideal of upstream governance — getting the collection right before training happens — precisely because undoing it afterward isn't a reliable technical option.

Special Category Data Gets the Strictest Treatment

Sensitive personal data — health information, political opinions, and similar special categories under GDPR — is treated as near-prohibited for scraping purposes, and any such data collected incidentally during a broader scrape must be minimized and deleted once discovered rather than retained. This applies even when the sensitive data wasn't specifically targeted, placing responsibility on the scraping party to actively filter it out once found rather than treating incidental collection as an acceptable byproduct.

What This Means While It's Still a Draft

Since the guidelines remain open for public consultation until October 30, 2026, the specific requirements described here reflect the current draft rather than a settled final rule — the version that eventually takes effect could shift based on feedback received during that window. What's already clear, though, is the direction: GDPR applies to AI training data scraped from the web with no special exemption, and the specific mechanics of demonstrating compliance are the open question the consultation period exists to work out.


The EDPB's draft guidelines close a specific gap that's existed since generative AI models started training on web-scraped data at scale — confirming GDPR applies fully, that public visibility isn't consent, and that a site's own technical restrictions now factor into the legal analysis. Whether the final version changes these specifics after consultation remains open, but the underlying position — no AI exemption from data protection law — is already the clear signal from this draft.

For questions or inquiries contact us at info@cleartexteditor.com