Skip to content
All posts

13 July 2026

5 min read

Written by

Clément Lacaille

Clément Lacaille

Founder, Tech-Bharat

About the author
Business & compliance

What actually counts as “anonymous” data: the EDPB’s new rules and what they mean for your AI projects

On July 8, 2026, the European Data Protection Board adopted guidelines on anonymisation and the first pan-EU guidelines on web scraping for generative AI, open for public consultation until October 30, 2026. If your SMB feeds customer data to an AI tool or runs an agent that reads the public web, both texts change what counts as compliant.

On July 8, 2026, at its 122nd plenary session, the European Data Protection Board — the EU body that groups national regulators such as France’s CNIL — adopted two texts that concern any SMB handing customer data to an AI tool: guidelines on anonymisation, and Guidelines 03/2026 on web scraping in the context of generative AI, the first pan-EU framework addressing how AI training data gets collected. Both are open for public consultation until October 30, 2026, meaning the final wording can still move, but the direction is already clear enough to act on.

A three-criteria test for what counts as anonymous

The anonymisation guidelines, informed by the Court of Justice’s September 4, 2025 ruling in case C-413/23 P, set out three cumulative criteria a dataset must meet to be considered genuinely anonymous: no possibility to single out an individual record, no possibility to link records to the same person across sources, and no possibility to infer information about a person. The Board allows two ways to apply the test — a “contextual” approach that factors in the different re-identification capabilities of whoever might hold the data, or a simplified approach that ignores those differences for the sake of a clearer, more conservative answer. If any one of the three criteria fails, the data stays pseudonymous, not anonymous — and pseudonymous data remains personal data fully subject to GDPR.

Web scraping now has a documented legal basis to satisfy

The scraping guidelines confirm that GDPR applies the moment scraped data includes personal information, and that legitimate interest — the legal basis most scraping relies on — must pass a three-step test: a genuine, specific interest; a real necessity for scraping to serve it; and a balancing exercise showing that interest does not override the rights of the people whose data gets collected. The Board recommends scraping only from reliable sources, recording when each item was collected, and validating data before using it, to satisfy the accuracy principle. Special categories of data — health, opinions, and similar — stay off-limits by default: scraping them requires both a lawful basis under Article 6 and a specific exception under Article 9(2) of GDPR.

What it means for your SMB

Two situations are directly affected. First, any SMB feeding customer data to an AI tool — a support agent, a WhatsApp sales assistant, a client portal — on the assumption that stripping out names and emails makes it “anonymous” and therefore outside GDPR’s scope: under the new three-criteria test, that is very often still pseudonymisation, meaning the tool, its hosting, and its access logs remain fully regulated. Second, any SMB running or considering a regulatory-watch or competitive-intelligence agent that reads public websites: it now needs a documented legitimate-interest assessment, not an assumption that public data is free to use, before that agent goes into production.

  • Audit what your AI tools actually receive: run the three-criteria test (no singling out, no linkage, no inference) on any dataset you call “anonymised” — if it fails on one criterion, treat it as personal data.
  • If you run or plan a regulatory-watch or competitive-intelligence agent that scrapes public sources, document the legitimate-interest test before launch: the specific interest, the necessity, and the balance against the rights of the people concerned.
  • Never scrape or process health data, opinions, or other special categories through an AI agent without both a lawful basis and an explicit exception under Article 9(2) — that combination is not optional.
  • Bring these two texts to your next review with your DPO, alongside the EU AI Act transparency obligations due August 2, 2026 — the two deadlines now sit on the same compliance calendar.

The consultation stays open until October 30, 2026, so the final text may still shift on details. What will not shift is the underlying logic: “anonymised” is a claim you now have to be able to demonstrate against a defined test, not a label you attach because a spreadsheet no longer shows a name. For an SMB running — or building — an AI agent on customer or public data, that is the difference between a compliant deployment and a five-minute conversation with a regulator you did not want to have.

Free resource

The self-assessment grid: 20 tasks AI can automate

Sales, admin, support, operations: the 20 tasks AI agents already handle in SMEs — with, for each one, the tell-tale sign that your team is concerned.

Read next