Eva Jarbekk
Partner
Oslo
Norway, Sweden, Denmark, UK
by Eva Jarbekk and Sofie Axelsson
Published:
Think your data is anonymous? The EDPB is not so sure — new guidelines explained
On 7 July 2026, the EDPB published its new draft Guidelines 02/2026 on Anonymisation. At first glance, it may look like a technical update; a refresh of the 2014 Article 29 Working Party Opinion, with new terminology and updated assessment criteria. Look more closely, and a considerably more significant shift emerges. Since the 2014 Opinion, there have been major changes in the legal, privacy, engineering and technological landscapes, including the development of CJEU case law, the establishment of EU-wide data spaces, and developments in AI and technology in general. All of which raised the need to update the previous guidance. The Guidelines are open for public consultation until 30 October 2026. They are worth reading carefully.
Anonymity is relative — and that changes everything
The most headline-grabbing concept in the Guidelines is what privacy practitioners are already calling "relative identifiability." The idea flows directly from the CJEU's 2025 judgment in EDPS v SRB, and the EDPB now builds it into a full framework. Information may be anonymous for some entities but not for others. The same dataset can simultaneously be personal data for the organisation that holds the re-identification key, and genuinely anonymous for a recipient who has no realistic way of linking it to any individual.
For example, information may unambiguously relate to an individual, but only one organisation may be able to actually identify that individual. In this case, the information would be considered personal data for that organisation, while it could be considered anonymous for everyone else. The practical implication is significant: if you share data with a recipient who genuinely cannot re-identify it, the GDPR may simply not apply to what they do with it. No legal basis required. No data subject rights. No deletion obligations. They could, in principle, use it to train AI models.
But before you reach for that as a solution, read on. Because the Guidelines come with exceptions that change the picture.
The processor exception nobody saw coming
Take a dataset and share it with a recipient who processes it on your instructions. This could typically be a vendor, a SaaS provider or an AI partner operating under a data processing agreement. That recipient is a processor. Under the Guidelines, where a processor acts on behalf of the controller, it inherits the controller's perspective on identifiability. If the controller can re-identify individuals in the dataset, the data is personal data for your processor too.
In other words: relative identifiability applies when data flows to an independent controller. It does not apply to processors. (Here one may speculate if this is a direct consequence of the SRB-verdict, where it was assumed by many that the recipient was a separate controller, not a processor – even though this was not clear in the verdict.)
A processor is bound by the controller's perspective, regardless of whether the processor itself has any realistic means of re-identifying anyone.
Why does this matter? First – because the existing practise of identifying processors with controllers will continue. Further - almost every SaaS contract in existence contains a clause along the lines of: "the service provider may use aggregated or anonymised customer data to improve its own products." That clause has long been treated as a convenient shortcut, a way for vendors to extract value from data without triggering full GDPR obligations. Under the logic of these Guidelines, it is evident that this may not be used in this manner if the controller can still re-identify the individuals.
At the same time – it is challenging to follow EDPB's reasoning. The same dataset, shared to an independent controller, is anonymous and free from GDPR constraints. Shared to a processor, it is personal data, fully regulated, SCCs and all. If this rule remains in the final Guidelines, it will be even more important to scrutinize the structure of set-ups in data sharing. Independent controllers have more freedom than processors.
Two approaches, one very high bar
In the Guidelines, the framework for assessing anonymity can be applied in two ways: the contextual approach, which considers the differences in capabilities between those who might identify the data subject, and the simplified approach, which does not.
The contextual approach is the actual legal standard. Relevant entities who might identify individuals may include unauthorised and malicious actors, including rogue employees, investigative journalists, domestic and foreign intelligence agencies, unethical companies, and cybercriminals. You are expected to assess whether any of them, through means reasonably likely to be used, could re-identify individuals from your dataset. That is a demanding exercise in practice.
In theory, the simplified approach offers a shortcut: assume the worst case and skip the individual assessment of each entity's capabilities. This simplified approach is not an alternative legal standard. On the contrary, it is best understood as a subcategory of the contextual approach. The test is straightforward: if the data cannot be re-identified even under that worst-case assumption, it qualifies as anonymous and no further analysis is needed. But if the data fails, which will be the case for most real-world datasets, the controller must fall back on the full contextual approach regardless.
A new approach to the classical identification analyzis
The familiar trio, singling out, linkability, inference, has been replaced. The framework presents three criteria for testing whether data is anonymous: No Record Isolation, No Linkage and No Inference. The renaming is more than cosmetic. The new formulations better reflect how modern re-identification works.
The No Inference criterion is the one to watch. In an era where AI inference attacks can determine whether a specific person's data was included in a model's training set, this criterion has teeth. This is so important that I choose to quote para 82 from the EDPB text:
[..]inference could also be drawn from data that has been aggregated to represent correlations between elements of the original information, rather than that information itself. This would, in particular, be the case for AI models or synthetic data. Specific inferences can be made by querying or prompting the given data with additional information to elicit new information about a particular individual. Where such an inference is also meaningful, this would be a violation of the No Inference criterion.
This means an organisation that has trained a model on personal data and published it cannot simply declare the output anonymous and move on. The inference risk follows the data.
The consequences of this may seem very significant, and I believe this will generate considerable engagement.
Anonymisation requires a legal basis — and always did
The principles of data protection apply whenever personal data is processed, which includes processing operations carried out to anonymise information. This means, amongst other things, that you need a valid legal basis just to run the anonymisation algorithm. If your dataset contains health data, employment data, or any other special category information, you also need a qualifying exemption under Article 9. Neither of these requirements is new. However, they are frequently overlooked in practice, particularly in data science and AI development contexts where the focus is on the output rather than the process. Controllers should clearly state that personal data will be processed to produce anonymous data which then falls outside of the scope of the GDPR, and should avoid any ambiguous or misleading statements. In particular, controllers should not use descriptions like "anonymous", "de-identified" or "de-personalised" if individuals are still identifiable.
Anonymisation is never truly final
Perhaps the most sobering point in the entire document is one that organisations have been slow to internalise: anonymisation is not a one-time event.
The likelihood of re-identification typically increases over time due to advances in the technology and techniques used for re-identification, as well as the increased availability of additional information. If the likelihood of identification increases to a level that is no longer insignificant, the previously anonymous data should again be considered personal, and the entity will be accountable for any processing of that personal data.
What does this mean in practice? A security incident may lead to a reassessment of anonymity if the assessment of anonymity relied on certain information being kept confidential, but the incident means that individuals can now be identified.
In other words: a data breach at a third party, completely unrelated to your organisation, could instantly transform a dataset you considered anonymous into personal data, and start a 72-hour breach notification clock. That is a risk many organisations have not yet factored into their incident response planning.
What this means for your organisation
The EDPB's new Guidelines shift anonymisation from a technical deliverable into something closer to an ongoing operational obligation. The question is no longer only whether your data has been anonymised. It is whether you can continuously demonstrate that it remains so, in light of changing technology, new auxiliary information, and evolving capabilities across the full chain of entities who might access it. That requires documentation, periodic reassessment, and a clear-eyed view of what your processors and third-party recipients are doing with the data you share with them.
A few practical starting points worth considering:
You can find the EDPB's Guidelines here.
Speaking of anonymisation, here is a case that illustrates just how relative the concept can be in practice.
An Austrian data subject posted about their ADHD diagnosis on a publicly accessible online forum. They used a pseudonym. From most readers' perspective, the post was entirely anonymous. But one reader knew exactly who was behind it: a follower of the data subject who recognised the pseudonym and promptly forwarded the post to a mutual contact via WhatsApp, explicitly linking the pseudonym to the data subject's real identity.
The data subject complained to the Austrian DPA, arguing that their health data had been disclosed without authorisation. The DPA dismissed the complaint. Its reasoning rested on Article 9(2)(e) GDPR, which permits the processing of special category data where the data subject has manifestly made that data public. Actively posting a diagnosis on a publicly accessible forum, the DPA concluded, constituted precisely such a conscious and unambiguous act of disclosure.
The outcome is legally sound, and the DPA's reasoning is difficult to fault. The data subject made a conscious choice to post sensitive health information on a publicly accessible forum. The fact that they used a pseudonym does not change the nature of that act, the information was genuinely made available to the public. The case highlights that pseudonymity and anonymity are not the same thing, and as a caution about the limits of pseudonymity as a privacy strategy. The protection afforded by a pseudonym is only ever as strong as the weakest link in the chain of people who might recognise you.
You can read more about the matter here.
We have just spent considerable time unpacking the EDPB's new anonymisation guidelines and the concept of relative identifiability. Here is a timely reminder of what happens when an organisation gets that assessment wrong.
On 26 May 2026, the CNIL fined IQVIA Operations France €5 million for a series of violations relating to two health data warehouses aggregating data from approximately 14,000 pharmacies and several thousand doctors, covering the health journeys of tens of millions of patients across France.
The anonymisation argument that did not hold
When the sanction procedure got underway, IQVIA invoked the CJEU's SRB judgment and argued that the data in its warehouses was anonymous, and that the GDPR therefore did not apply. The CNIL's restricted committee rejected this entirely. Each patient was assigned a unique identifier tracking their full care journey across time, and the depth of data collected was considerable: diagnoses, symptoms, prescriptions, allergies, weight, height, socioeconomic status, and more. Re-identification using reasonable means was possible, particularly by cross-referencing with publicly available data.
There is also a detail worth noting: prior to the SRB judgment, IQVIA had never disputed that it was processing personal data. It had actively sought and obtained CNIL authorisations for both warehouses on exactly that basis. Raising the anonymisation argument mid-enforcement, after years of operating under a personal data framework, was always going to be a difficult position to defend.
Beyond the anonymisation question, the inspections revealed a broader pattern of non-compliance: no systems to detect abnormal access, no multi-factor authentication, patients never informed that their data was being transferred, and, most strikingly, pharmacy software that transmitted customer data to IQVIA even where patients had refused.
For any organisation holding large health or behavioural datasets, the message is straightforward: pseudonymisation is a risk-reduction measure, not a legal reclassification. And an authorisation obtained on the basis that data is personal cannot easily be unwound by arguing, mid-enforcement, that it was anonymous all along.
You can read more about the matter here.
2 August 2026 has come and gone. The EU AI Act's transparency obligations are now formally in force, and the question has shifted from "are we ready?" to "are we compliant?" It is not the full weight of the regulation, the major obligations for high-risk AI systems are still some years away, but the rules that kicked in this month are real and immediate. And as with most EU legislation, the details matter considerably more than the headlines suggest.
Norway is not directly bound by the 2 August deadline, but the picture is worth following closely. The AI Act has been assessed as EEA-relevant, and the Code of Practice on transparency of AI-generated content will therefore also become relevant for Norway. However, the timing of formal EEA incorporation remains unclear. That said, organisations with EU-facing operations are exposed regardless of the EEA timeline. So when incorporation does come, the supporting documents, including the Code of Practice and the Commission's guidelines, will carry direct relevance.
For organisations operating in or doing business with the EU, the regulation applies regardless of where the organisation is based. The rules are live, and the exposure is real.
Who actually has to disclose what?
The centrepiece of the transparency rules is Article 50(1) of the AI Act. The obligation is straightforward in principle: AI systems that interact directly with people must be designed so that users are clearly informed they are talking to a machine, unless this is already obvious. Chatbots, virtual assistants, automated customer service tools all fall squarely within scope.
But here is where it gets interesting. Article 50(1) binds providers, not deployers. A provider is the entity that develops an AI system and places it on the market or puts it into service under its own name or trademark. A deployer is an organisation that uses a ready-made system for its own purposes. If an organisation has embedded a third-party chatbot under that third party's branding, it is generally a deployer, and the disclosure obligation rests primarily with the provider who built the system.
The critical word is "generally." Consider a scenario that is far more common than it might appear: a company commissions a chatbot from an external developer, gives it a name, and deploys it under its own brand and visual identity. Under Article 3(3) of the AI Act, a provider includes anyone who has an AI system developed and puts it into service under their own name or trademark. Article 3(11) makes clear that "putting into service" includes supplying a system for a company's own use. The result is that the company may well have crossed the line from deployer to provider, and inherits the Article 50(1) disclosure obligation itself. A small attribution tag from the underlying developer is unlikely to change that conclusion if the predominant branding is the company's own.
The question to ask is therefore not "did we build this?" but rather quite often: "whose name is on it?"
The Code of Practice and the Commission guidelines
Two supporting documents have arrived alongside the regulation. A Code of Practice on transparency of AI-generated content was published on 10 June, developed by six independent experts following input from over 180 stakeholders. Signing up is voluntary, but organisations that adhere to it can point to that adherence as evidence of compliance with Article 50. Those that do not sign must demonstrate compliance through their own means.
The Code has two main parts: one for providers, focusing on machine-readable marking and detection of AI-generated content, and one for deployers, addressing labelling of deepfakes and AI-manipulated text, including guidance on placement and internal compliance processes. On 20 July, the European Commission published its own guidelines on the implementation of Article 50, clarifying the scope of the obligations and covering aspects the Code does not address.
You can find the Code of Practice here and the Commission Guidelines here.
Deployers are not off the hook either
While Article 50(1) targets providers, deployers carry their own obligations under Article 50. Where a deployer uses an AI system to generate or manipulate content constituting a deepfake, disclosure is required. The same applies to AI-generated or manipulated text on matters of public interest that has not been subject to human review. If you look to the Commission Guidelines, human review entails at least fact-checking the accuracy of the content. The review should include a deliberate examination of the substance of the content by someone with relevant knowledge on the subject; solely formal checks will not do. That is a provision with obvious implications for anyone using AI in communications, public affairs, or information contexts. Deployers using emotion-recognition or biometric categorisation systems must also notify the individuals concerned – this is something that may seem quite uncommon – but in many countries it is not unusual to have emotion-recognition in call centers in order to direct particular angry customers to employees that are trained to handle them. To my knowledge, it is (yet?) not common in Scandinavia.
Machine-readable marking: the invisible layer
Alongside the human-facing disclosure obligations, Article 50(2) of the AI Act requires providers of AI systems that generate synthetic audio, image, video or text to embed machine-readable markings in their output. These are technical signatures that automated detection systems can identify, even where no visible label is present. The regulation does not prescribe a single technique; permitted approaches include watermarks, metadata tagging, cryptographic methods and fingerprinting. The purpose is to ensure that AI-generated content remains traceable after it has been shared, cropped or stripped of its original context, which explains why the obligation sits with providers rather than deployers.
Key things to have in place now
Here is a practical summary of what the transparency rules require as of 2 August:
The Digital Omnibus: a revised timeline
One other development that deserves attention in its own right: on 27 July 2026, the EU Digital Omnibus on AI entered into force, amending the AI Act in several important respects. The key dates for the broader regulation now look like this:
2 August 2028: High-risk AI systems embedded in regulated products.
The Digital Omnibus also clarifies the relationship between the AI Act and the GDPR, explicitly stating that the AI Act does not displace the GDPR, the e-Privacy Directive or the Law Enforcement Directive, each framework continues to apply on its own terms. This is directly relevant given the overlap between AI transparency obligations and data protection requirements around profiling and automated decision-making. Furthermore, a practical duplication is removed: where a deployer has already conducted a data protection impact assessment under Article 35 of the GDPR, that assessment can be cross-referenced directly in the AI Act's fundamental rights impact assessment, rather than repeated from scratch. It also introduces a new prohibition on AI systems designed to generate non-consensual intimate imagery.
You can find the Digital Omnibus here.
A note from Greece
Finally, a detail that is too striking to leave out. Greece is currently transposing the AI Act into national law, and a last-minute amendment to its draft legislation would make the removal of deepfake labels, including invisible watermarks, a criminal offence carrying a potential prison sentence. Germany's implementation includes no penalties for transparency rule violations at all. Ireland and Spain allow for fines.
The AI Act gives member states leeway in setting penalties, provided they are effective, proportionate, and dissuasive. One wonders whether the drafters of the AI Act had prison sentences in mind when they wrote "dissuasive". It is, however, a useful illustration of how differently the same regulation can land across legal cultures, and a reminder that as national implementations accumulate, the "EU AI Act" will not look quite the same in every country.
You can read more about the transparency rules here and Greece's penalties here.
7 July was a busy day at the EDPB. Alongside the anonymisation guidelines we covered above, the Board also published Guidelines 03/2026 on web scraping in the context of generative AI. If you are building or buying training data scraped from the web, this is the closest thing to an operating manual that EU data protection authorities have ever published for you. It is open for public consultation until 30 October 2026.
The guidelines are, by EDPB standards, surprisingly pragmatic. That is worth noting, because the Board has not historically given AI developers much ground. Its Opinion 28/2024 on AI models offered only case-by-case agnosticism, wrapped every concession in demanding evidential burdens, and left the problem of special category data entirely out of scope. These guidelines do not. They offer a worked example in which a compliant scraping operation is effectively pre-cleared, and they propose a solution to the special category problem that actually functions in practice. That is a meaningful shift.
Consent is out, legitimate interest is in
The legal basis of legitimate interest under Article 6(1)(f) GDPR is often relied on for scraping in the context of development of generative AI by private bodies. Consent is essentially ruled out: organisations scraping data from the internet do not have a direct relationship with the data subjects and are most probably not able to identify and obtain consent from each and every person before scraping. Importantly, the absence of a robots.txt file on a website does not amount to consent within the meaning of the GDPR.
In order for legitimate interest to apply lawfully, three cumulative conditions must be met: a legitimate interest must be pursued by the controller or by a third party; the processing of personal data must be necessary for that purpose; and the interests or fundamental rights of the data subjects must not override the controller's legitimate interest — a balancing test.
On the first condition, an interest may be regarded as legitimate if it is lawful, clearly and precisely articulated, real and present, and not speculative. "We might need training data at some point" will not do.
On necessity, narrowing the collection criteria to exclude unnecessary personal data, rather than scraping a wide part of the internet, may be crucial to ensure the necessity condition is met. Using pseudonymised or synthetic data may be another less intrusive way of pursuing the same purpose.
The balancing test is where most of the operational work lies. The restrictions imposed by the scraped website are a key consideration, for example through robots authentication, robots.txt (a file communicating that the text is not to be indexed by crawlers) or ai.txt files (a similar file specifically addressing access by AI systems), or CAPTCHA (a human-verification test, such as "select all images containing traffic lights", designed to block automated access). If data subjects are aware that a website implements these measures and the website is nevertheless scraped, it is less likely that they can expect the processing of their personal data by a scraping entity. robots.txt and ai.txt are not technicalities; they are evidence the EDPB explicitly weighs against you if you ignore them.
The guidelines include a worked example that will be of practical interest: an organisation that exclusively uses data from freely and publicly accessible online sources, where the data subjects have manifestly made the content public, excludes copyright-protected content, implements safeguards to limit data memorisation and regurgitation, restricts problematic content generation through technical or contractual measures, facilitates the exercise of data subjects' rights when re-identification is possible, and clearly indicates data sources in a publicly available privacy policy. In such a case, the balancing test may generally be considered as met. That is a meaningful concession from the Board.
The special category problem — and the search engine solution
This is the issue that has quietly rendered countless AI training projects legally indefensible, and the guidelines address it head-on this time around. Despite organisational and technical measures, it remains likely that a controller will residually process special categories of personal data that it did not intend to collect.
Rather than leaving controllers stranded, the EDPB reaches for the CJEU's 2019 judgment in GC & Others (C-136/17), originally decided in the context of search engines. The Court recognised that the specific features of the processing carried out by the operator of a search engine may have an effect on the extent of the operator's responsibility and obligations, and that the prohibition in Article 9(1) GDPR applies "within the framework of his responsibilities, powers and capabilities." The EDPB considers that this reasoning can be relevant for the incidental and residual collection of special categories of personal data in the context of web scraping for the training of an AI model, though it should not be seen as a general exemption from the requirements in Articles 9 and 10 GDPR.
For this to apply, a case-by-case analysis must show that: the processing activity has relevant similarities with the processing of a search engine; the processing includes only incidental and residual special category data, not intentional collection; it is difficult or impossible to assess in advance whether scraping will capture such data; and the controller implements measures within the framework of its responsibilities, powers and capabilities to prevent the dissemination of that data.
What does that mean in practice? The measures must span the full lifecycle: defining precise criteria and filters before collection to prevent special category data from being gathered; deleting any such data identified after collection as soon as possible; applying resistance to extraction attacks and output filters during model development; and continuously monitoring outputs after deployment, with model unlearning considered as a longer-term mitigation as technology develops.
This is a narrow, fact-specific route, not a blanket exemption. The controller must be able to demonstrate, in line with the accountability principle, that the conditions are relevant to its specific processing activity and that the measures adopted are relevant and effective.
What to have in place
A few practical points drawn directly from the guidelines:
The consultation on these guidelines closes on 30 October 2026, the same date as the anonymisation guidelines. If your organisation has a view on either document, now is the time to submit it.
You can read more about the matter here.
The only single matter included in this newsletter is an important one. We will revert to many other matters in the next issue.
Search engines have long enjoyed a degree of legal shelter that most content publishers do not. The traditional reasoning is straightforward: a search engine does not create content, it finds it. Google points you to a website; if that website says something defamatory, Google is not the one saying it. A German regional court has now decided that this reasoning does not extend to AI Overviews, and the implications reach well beyond Germany.
The Regional Court of Munich issued a temporary injunction against Google after its AI Overview feature falsely linked two Munich-based publishers to scams, subscription traps, and dubious business practices. The AI had mixed up information about genuinely problematic companies with the plaintiffs, drawing connections that did not appear in any of the linked sources. When the publishers sent Google a cease-and-desist letter, Google did not respond appropriately. The court then stepped in.
Google's own words
The core of the ruling is a distinction that sounds simple but carries significant weight. A traditional search result lists sources and quotes them directly. An AI Overview rewrites, evaluates, and synthesises. It speaks in its own words and according to its own structure. In this case, the overview opened with the confident assertion that one of the companies "is known for dubious business practices," then constructed its own narrative complete with a summary, red flags, and user tips. None of that language appeared in any of the underlying sources.
The court concluded that these were Google's own statements, not third-party content that Google had merely made findable. Because Google alone has influence over the AI's output and the algorithms driving it, Google alone bears responsibility for what it produces. The existing case law shielding traditional search engines from liability, developed by Germany's Federal Court of Justice precisely because search engines only point to outside content, simply does not apply when the AI is generating new and independent claims of its own.
Google argued that users could simply check the linked sources themselves and should know not to blindly trust AI-generated content. The court was unimpressed. A statement that is self-contained and independently understandable does not become legally harmless because a diligent reader could theoretically disprove it through further research. The court drew a parallel to press law: a publisher is liable for a teaser that stands on its own, even if the full article tells a more nuanced story. Google's argument, the court noted, would also rather undermine the purpose of the feature; if the overview is unreliable, what exactly is it for?
Why this matters beyond Munich
Google was ordered to cover eighty percent of the legal costs. But the broader significance of the ruling lies elsewhere. An analysis cited in the source material found that Google's AI Overviews answer correctly approximately ninety-one percent of the time. At the scale Google operates, that remaining nine percent still amounts to an enormous volume of incorrect answers generated every hour. If a meaningful portion of those errors concern identifiable individuals or organisations, the liability exposure across the industry is considerable.
The ruling is not yet final, and Google has indicated it is reviewing the decision. Whether the reasoning survives appeal, and whether courts in other jurisdictions follow the Munich court's logic, remains to be seen. But the direction of travel is clear: AI systems that generate their own claims, rather than merely surfacing those of others, are increasingly likely to be treated as the authors of those claims. For any organisation deploying AI features that summarise, synthesise, or comment on third-party information, that is a development worth following closely.
You can read more about the matter here.
Partner
Oslo
Associate
Stockholm
Managing Associate - Qualified as EEA lawyer
Oslo
Partner
Oslo
Partner
Oslo
Partner
Oslo
Partner
Oslo
Senior Lawyer
Oslo
Associate
Stockholm
Partner
Oslo
Senior Associate
Oslo