
The mismatch between the territorial nature of copyright and the global architecture of AI training pipelines creates a legal risk that many IP managers have not yet priced in. A dataset that is built from works that are in the public domain in one jurisdiction can still infringe copyright in a neighbouring Member State where the term of protection has not yet run. The Court of Justice of the European Union (CJEU) has now handed down a judgment that forces AI developers to confront this territorial trap directly.
In July 2026, the CJEU ruled on a reference concerning the online publication of the original manuscript of The Diary of Anne Frank. The case exposed the consequences of making a work available on a website without geo-blocking when the work is protected in some Member States but not in others. The judgment is not about AI, but its logic applies with full force to the datasets used to train large language models and multimodal systems. For Korean companies that train models on EU-origin data and deploy them in the European market, the ruling is a compliance signal that cannot be ignored.
The copyright in Anne Frank’s works expired in the Netherlands on 1 January 2016, 70 years after her death in 1945, as the Korean Copyright Commission notes (Source 7). The original manuscript, however, was made available online. The CJEU was asked to decide two questions, as reported by the IPKat (Source 1):
The Court held that making the work available on a website that is accessible from a Member State where the work is still protected constitutes a communication to the public in that Member State, even if the server is located in a country where the work is in the public domain. The person who decides to make the work available is the actor. The judgment also addressed the role of effective technological measures: the absence of geo-blocking to prevent access from the protected Member States points to liability. Conversely, the use of a VPN by an end user to circumvent geo-blocking does not automatically cure the uploader’s unauthorised communication, as the Korean Copyright Commission’s issue report on the case indicates (Source 7).
The CJEU’s reasoning is not confined to the right of communication to the public. It rests on the principle that copyright is territorial: the existence and scope of protection are determined by the law of each Member State for which protection is claimed. The same logic applies to the reproduction right. When a developer scrapes or downloads copyright-protected works from a website to build a training dataset, the act of reproduction occurs in each jurisdiction where the copying takes place or where the resulting copy is used. If the work is still protected in Spain, but the scraping is done from a server in the Netherlands where the work is in the public domain, the reproduction may still infringe Spanish copyright, unless the developer has ensured that the data is not accessed from Spain.
This territorial fragmentation is not a hypothetical edge case. For works whose authors died in the mid-20th century, the 70-year post mortem auctoris term means that the same text can be in the public domain in one EU country and protected in another for several years, depending on the exact date of death and any national wartime extensions that were preserved before full harmonisation. A dataset that includes 20th-century European literature, journalism, or personal archives will almost certainly contain works with a patchwork of copyright statuses across the EU.
The EU AI Act adds a further layer. Providers of general-purpose AI (GPAI) models must, under Article 53(1)(d), draw up and make publicly available a sufficiently detailed summary of the content used for training, using a template provided by the AI Office (Source 2). The template, published on 24 July 2025, requires disclosure of data sources, including whether publicly available datasets and crawled data were used, and what measures were taken to respect opt-outs under the text and data mining (TDM) exception of the DSM Directive. The CJEU’s Anne Frank judgment adds a spatial dimension to this transparency obligation: the summary must be understood against the backdrop of territoriality. A claim that training data was “lawfully accessible” for TDM purposes is not a single Boolean answer; it is a set of Member-State-specific answers.
Meanwhile, a parallel CJEU reference, Case C-250/25, asks whether the output of a large language model (Google’s Gemini) infringes press publishers’ rights under Article 15 of the DSM Directive and whether the training of the model falls within the TDM exception under Article 4 (Source 5). The Budapest District Court submitted those questions on 3 April 2025. When the CJEU answers, it will likely address the territorial scope of the TDM exception as well. Already, the Anne Frank judgment signals that the Court will not accept a “location of server” safe harbour to bypass territorial copyright protection.
Korean companies face a double exposure. The EU AI Act’s transparency and opt-out requirements apply to any GPAI model placed on the EU market, regardless of where the provider is established. The Korean AI Basic Act amendment proposals, as noted by the Korean Copyright Commission, would introduce a similar obligation for AI operators to establish a procedure for verifying whether a copyright holder’s works have been used as training data (Source 2). A Korean AI developer that cannot demonstrate, jurisdiction by jurisdiction, that it respected territorial copyright in its training data will face regulatory risk in both Europe and Korea.
The Anne Frank judgment gives IP managers a concrete reason to run a jurisdiction-by-jurisdiction copyright clearance on their training datasets. Do not rely on the public domain status of a work in the country where the data was collected. That is a single-point check that the CJEU has now rendered insufficient.
For every text source in the training corpus that was originally published or created in the EU, map the year of death of the author and the applicable copyright term in each Member State where the AI model will be deployed. Flag any work for which the term has expired in some Member States but not in others. For those works, implement one of two measures before training: (1) exclude the work from the dataset entirely for any use that will reach the protected Member States, or (2) deploy effective geo-blocking during the data ingestion and training phases to prevent the copying from taking place in a jurisdiction where the work is still protected. The second option requires documented technical controls that are not trivially circumvented—the CJEU’s reference to effective technological measures suggests that a negligible geo-blocking effort will not shield a defendant.
Prosecution and transactional attorneys drafting data supply agreements should insert a representation that the data provider has verified the copyright status of the supplied works in all EU Member States and has obtained licences where necessary. If the data is scraped from the open web, the IP owner must keep a record of the TDM opt-out protocols that were respected at the time of scraping, broken down by domain, and document the geo-blocking configuration in place. The record should be granular enough to be disclosed in the AI Office’s training data summary template without triggering a claim of misrepresentation.
Korean filers who are on the patent side but whose organisations also deploy AI products should treat this as a cross-functional risk. The IP function can run the copyright clearance audit while the data science team builds the geo-blocking architecture. In-house legal-ops teams should align the training data governance policy with the territorial logic of the Anne Frank judgment now, rather than waiting for the CJEU’s answer in Case C-250/25 to close the remaining gaps. The territorial principle is already clear. The only remaining question is how far the reproduction right will be read to reach the training stage itself—and the Anne Frank judgment suggests the CJEU will read it broadly.