Abstract: This white paper establishes a legal and technological governance framework for AI's stewardship of cultural assets and digital heritage during large-scale artificial intelligence model training. It reconciles the text and data mining exceptions under Articles 3 and 4 of Directive (EU) 2019/790 (DSM Directive) with the mandatory transparency and risk-mitigation obligations of Regulation (EU) 2024/1689 (EU AI Act). The analysis introduces institutional provenance verification protocols, safeguards for proper algorithmic style appropriation, and lifecycle audit mechanisms to preserve cultural integrity within enterprise synthetic data ingestion pipelines given their role in carrying on human civilisations' cultural heritage. While examining the lawful ingestion of digitized cultural heritage into artificial intelligence training sets under European Union law, this legal study analyzes the interaction between classical copyright exceptions – specifically Articles 3 and 4 of Directive (EU) 2019/790 (DSM Directive) and Directive 2012/28/EU (Orphan Works) – and the data governance, transparency, and copyright compliance requirements of Regulation (EU) 2024/1689. The paper defines institutional consent protocols for cultural institutions and identifies product-liability intersections where AI systems trained on cultural data trigger cybersecurity obligations under Regulation (EU) 2024/2847 (Cyber Resilience Act).
Intellectual Property, Curatorial Fidelity, and EU Compliance under the AI Act, TDM and DSM Directives
Growing up surrounded by museography, national and international cultural patrimony cataloguing and archival conservation, research and scientific presentations, and thousands of catalogued artifacts, one truth becomes clear: preserving culture is not about keeping heritage in a glass box – it is an active, rigorous science of transmission. Generative AI is the most powerful tool for cultural preservation since the invention of the printing press. Yet, just as physical restoration requires meticulous care to prevent damaging an ancient canvas, digital ingestion demands equal discipline. The objective is not to lock AI out, but to ensure that machine learning honors the integrity, provenance, and truth of human memory across centuries.
A digitised icon, an oral history recording, or an archaeological dataset is not just source material for a generative AI model, it is a legal object with its own rights, consents, and protocols attached. This article maps the EU rules that govern when, and on what terms, cultural heritage may lawfully enter an AI training pipeline, and what cultural institutions should demand from AI providers in return.
Digital cultural heritage is no longer a secondary concern. It has become a necessary condition for the preservation and continuity of what humanity has created across centuries. Physical objects, sites, and collections remain vulnerable to war, natural disaster, neglect, and the simple passage of time. The interrupted excavations at places such as Uruk remind us how easily knowledge can be fragmented or lost.
In this context, the careful digitisation of cultural materials and the responsible use of powerful AI instruments are not threats to culture; they are tools that, in the right hands, can help safeguard, reconstruct, and transmit it. Yet the power of these instruments requires precise legal and institutional governance. When digital representations of cultural heritage, museum collections, archival materials, ethnographic recordings, traditional knowledge, and digitised archaeological data, are used to train generative AI systems, fundamental questions of intellectual property, consent, control, and legitimacy arise under European Union law.
The Legal Question
Under what conditions may digital cultural heritage be lawfully used as training data for AI systems placed on the EU market or used within the Union?
The answer requires simultaneous attention to classical intellectual property rules, the data governance and fundamental rights obligations of the AI Act, and, where the resulting system qualifies as a “product with digital elements,” the cybersecurity framework of the Cyber Resilience Act. Who holds the relevant rights? What constitutes valid consent or a legitimate exception? How do the transparency and risk-management duties of the AI Act interact with the text and data mining exceptions of the Digital Single Market Directive? And when does an AI system trained on such material trigger the obligations of the Cyber Resilience Act?
Where cultural data originates from Earth-observation or other space-based systems, parallel questions of liability and data use also arise under the emerging EU Space Act framework.
Legal Instruments: Copyright, Consent, and the AI Act
Copyright and related rights remain the primary formal constraints. Directive (EU) 2019/790 on copyright and related rights in the Digital Single Market (the DSM Directive) addresses text and data mining (TDM) directly.
Article 3 establishes a mandatory exception for research organisations and cultural heritage institutions that have lawful access, allowing them to carry out TDM for scientific research. Article 4 provides a broader exception for the mining of lawfully accessible works, subject to the right holder’s express reservation of rights (the “opt-out”), including by machine-readable means.
The practical operation of that opt-out is no longer purely theoretical. In a significant German case, Kneschke v. LAION e.V. (Regional Court of Hamburg, case no. 310 O 227/23), a photographer challenged the scraping and processing of his image for an AI training dataset. The Regional Court of Hamburg (judgment of 27 September 2024) held that the use fell within the TDM research exception. On appeal, the Hanseatic Higher Regional Court of Hamburg (judgment of 10 December 2025, case no. 5 U 104/24) dismissed the photographer’s appeal and confirmed that the reproduction was covered by the TDM exception. Although a further appeal to the Federal Court of Justice remains possible, these decisions currently provide the clearest judicial guidance in the EU on how the Article 4 exception and its opt-out function in practice.
It is equally important to note Article 14 of the DSM Directive, which ensures that faithful digital reproductions of out-of-copyright visual artworks do not attract new copyright. This prevents institutions from claiming exclusive intellectual property rights over basic 2D scans of public domain art, pushing them instead toward contractual controls, access restrictions, and strict licensing terms.
Furthermore, where ethnographic archives, oral histories, or contemporary photo collections are concerned, the legal framework expands beyond intellectual property into the GDPR. Here, valid consent or established legitimate interest dictates whether personal data embedded in cultural heritage can lawfully enter a model’s training pipeline. Similarly, for traditional knowledge or sacred community materials, which Western IP frameworks typically leave in the public domain, institutional control must rely heavily on ethical governance, soft law (such as WIPO guidelines), and strict access agreements to prevent cultural distortion.
The AI Act adds a distinct, operational layer to these rules. Crucially, Article 53(1)(c) of the AI Act turns the copyright reservation of Article 4 DSM into an explicit compliance duty, requiring providers of General-Purpose AI (GPAI) models to demonstrate active policies for respecting right holders’ opt-outs where they correctly stand for curatorial guardrails & provenance integrity. GPAI providers must also publish a sufficiently detailed summary of the content used for training, under Article 53(1)(d). Systems that process sensitive cultural data may trigger enhanced fundamental-rights impact assessments. Romania’s designation of ANCOM as the national market surveillance authority and single contact point for the AI Act gives these obligations a clear national enforcement anchor (see our earlier analysis: Romania Designates ANCOM as National AI Market Surveillance Authority).
From Statute to Practice: The GPAI Code of Practice
The abstract duties of Article 53 now have an operational counterpart. On 10 July 2025, the European Commission published the General-Purpose AI Code of Practice, a voluntary instrument developed with independent experts and confirmed by the Commission and the AI Board as adequate to demonstrate compliance with Articles 53 and 55 of the AI Act. Its Copyright chapter translates the Article 53(1)(c) opt-out duty into concrete commitments and measures for curatorial guardrail and provenance integrity, while a companion template, published later that same month, structures what a “sufficiently detailed summary” of training content under Article 53(1)(d) should actually contain.
For cultural institutions, this matters practically rather than academically. A vendor that has adopted the Code, and can point to a completed training-data summary against the official template, offers a materially more auditable counterparty than one relying on Article 53 alone. Institutions negotiating AI procurement contracts should treat adherence to the Code, and the completeness of the published summary, as a due-diligence checkpoint alongside the contractual guarantees discussed below.
Filling the Gaps: Orphan Works and Out-of-Commerce Collections
Cultural heritage digitisation rarely deals with clean rights situations. A significant share of museum, library, and archival holdings feature unclear, dispersed, or untraceable rights holders – a problem the TDM exceptions alone do not solve, since Articles 3 and 4 DSM presuppose lawful access to an identifiable work, not resolution of an authorship gap.
Two further instruments exist precisely for this situation, and belong in any institutional AI-readiness assessment:
- Directive 2012/28/EU (the Orphan Works Directive), which permits certain uses, including digitisation and making available, of works whose rightsholders cannot be identified or located after a diligent, documented search, provided the use falls within the institution’s public-interest mission.
- Articles 8–11 of the DSM Directive, which establish an extended collective licensing mechanism, and a fallback exception where no representative collective management organisation exists, for out-of-commerce works held in the permanent collections of cultural heritage institutions.
Where an institution intends to make such works available for AI training or research use, documenting the diligent search (for orphan works) or the collective licence and its territorial scope (for out-of-commerce works) is a precondition, not a formality, and should be built into the institution’s internal governance before any AI producer is engaged (accounting for and amounting to vendor's algorithmic custodianship & data fidelity objectives).
The Relevance of the Cyber Resilience Act (CRA)
When a generative AI system, or a system incorporating generative AI components, is commercially distributed or placed on the EU market as a product with digital elements and has a direct or indirect data connection, it falls within the scope of the Cyber Resilience Act (Regulation (EU) 2024/2847). In that case, the manufacturer must comply with secure-by-design obligations, vulnerability handling, software bill of materials requirements, and the mandatory reporting of actively exploited vulnerabilities and severe incidents that begins on 11 September 2026.
For a cultural institution procuring such a system, the CRA serves as a critical vendor compliance standard. Residual risk can remain with the institution if the producer does not clearly assume CRA obligations, or if the system is substantially customised under the institution’s sole control. Effective risk allocation requires that the institution translate its own expertise into precise contractual requirements and verify the vendor’s capacity to meet them (Algorithmic Custodianship & Data Fidelity). This process requires coordinated input from curatorial, technical, and legal teams. Our Cyber Resilience Act Implementation Guide provides a practical roadmap for managing these concurrent timelines.
A Framework in Motion: The Digital Omnibus
None of the above should be read as settled once and for all. On 19 November 2025, the European Commission proposed a “Digital Omnibus,” a legislative package amending the AI Act and, in a parallel Data Omnibus strand, touching the GDPR itself. Political agreement on the AI Act elements was reached in May 2026, and the resulting simplification regulation entered into force on 27 July 2026, chiefly deferring high-risk AI system obligations. The GPAI transparency and copyright duties discussed above, already applicable since August 2025, are not among the provisions being delayed.
More consequential for this article is a strand still under negotiation: proposed clarifications to the GDPR that would treat AI training as capable of qualifying as a “legitimate interest,” alongside a narrowed definition of personal data. If adopted, this would directly affect the consent analysis for ethnographic archives, oral histories, and photographic collections containing identifiable individuals. Institutions should treat their current GDPR posture on cultural heritage data as provisional and revisit it once this strand of the Omnibus is finalised.
Practical Consequences and Institutional Control
Museums and cultural institutions already possess sophisticated systems for classifying cultural material, by period, typology, material, function, and religious or secular character. This existing organisational intelligence is a significant institutional strength. When an institution decides that a generative AI system is necessary for internal research, specialist access, or exhibition support, that same structured knowledge should shape the requirements imposed on the AI system and its producer.
Distortion of sensitive cultural material can occur in several ways: the generation of reconstructions that mix authentic elements with invented ones without clear labelling; stylistic alterations that change meaning; recombination of sacred or community-restricted material in ways that violate cultural protocols; and the loss of provenance information.
Prevention begins with clear internal governance regarding what may be digitised, with whose consent, and for what purposes. Only then should precise demands be placed on the AI producer. The distinction that should govern those demands is straightforward. Legal and protocol status – consented, reserved, orphaned, restricted, out of commerce – is to be flagged and bound to the pipeline. Meaning is not a fifth label. It is to be retrieved from the interpretive record, attributed to its source, and left open to later reading. Institutions are then in a position to contractually and technically require:
- Robust access controls that limit training or generation to authorised, carefully curated subsets of the collection.
- Technical guardrails that prevent the unauthorised recombination of designated sensitive, sacred, or personal works.
- Provenance retention, ensuring original source tracking is preserved and clearly displayed in any generated output.
- Effective audit and human-oversight tools that ensure meaningful control remains with the institution, not just the algorithm.
- Retrieval and attribution of the interpretive record – catalogue notes, conservation files, specialist commentary, art-historical monographs, community protocols, and later corrections of attribution – wherever generation would otherwise present a completed meaning as if it were the source.
The Romanian Anchor: The National Heritage Institute
In Romania, this framework does not operate in an institutional vacuum. The National Heritage Institute (Institutul Național al Patrimoniului - INP), operating under the Ministry of Culture, is the country’s designated national aggregator for the digitisation of cultural resources, feeding Romania’s Digital Library into Europeana. INP also maintains DOCPAT, the national documentation system for movable cultural heritage collections.
INP’s activity already implements the two EU soft-law instruments most relevant to this article: Commission Recommendation 2011/711/EU on the digitisation and online accessibility of cultural material, and Commission Recommendation (EU) 2021/1970 on a common European data space for cultural heritage. For Romanian museums, archives, and libraries, INP is therefore the natural first point of institutional coordination when structuring a digitisation or AI-training project, both to align with the national inventories it maintains (the List of Historic Monuments, the National Archaeological Repertory, and the National Inventory of Movable Cultural Heritage) and to ensure any AI-training use of digitised holdings is consistent with the aggregation and provenance standards INP already applies to Romania’s contribution to Europeana.
Institutions intending to license out-of-commerce or orphan works for AI use, as discussed above, should likewise coordinate the diligent-search and collective-licensing documentation with INP’s existing inventories, which in many cases already hold the provenance data needed to satisfy that documentation requirement.
A Balanced Pathway
Digital cultural heritage is necessary for preservation and continuity. Powerful AI instruments can serve that purpose when they are governed with transparency, accountability, and good faith. The legal frameworks already in place, intellectual property rules, data privacy laws, the AI Act, and the Cyber Resilience Act, provide the necessary structure, and are actively being refined through instruments such as the GPAI Code of Practice and the Digital Omnibus. What remains is the institutional capacity to translate those frameworks into clear internal policies, careful data governance, and well-designed contractual requirements.
Cultural institutions that approach the procurement of generative AI systems with this level of intentionality, and AI developers prepared to meet rigorous provenance, control, and compliance standards, are better positioned to turn regulatory obligations into sources of trust and long-term value.
Organisations developing or deploying AI systems that draw on cultural heritage materials, and cultural institutions seeking to protect or strategically govern their digital collections, are invited to discuss tailored legal strategies and integrated compliance pathways with us.
Coda – From Regulation to Legacy: A Shared Responsibility
The legal frameworks governing European digital space – from the DSM Directive and AI Act to the Cyber Resilience Act – are often viewed as regulatory hurdles. In reality, for cultural institutions and AI developers alike, they represent modern curatorial protocols. When AI systems are trained with respect for opt-outs, provenance, and data integrity, they cease to be mere scrapers and become true guardians of heritage.
As we bridge the gap between historic cultural archives – whether physical, digital, or digitized – and neural networks, our duty is clear: to empower AI innovators with the legal clarity and proper stewardship needed to build boldly and confidently, while ensuring that the cultural tapestry of humanity enters the digital future uncorrupted. Through rigorous governance, tech optimism, and deep respect for historical and cultural truth and meaning, we can ensure that tomorrow’s technology, bearing a curatorial duty to respect its source, consent, and contextual truth, actively carries and weaves yesterday’s legacy into the future.
Key Legal Sources & Primary Instruments
Keywords:
#Legal White Paper on AI & Cultural Heritage #AI Enterprise Value Proposition #Actionable AI Vendor Procurement Contracts for Cultural Assets#Digital Cultural Heritage#Curatorial Stewardship & AI#Cultural Asset Ingestion#Generative AI Provenance#Cultural Data Governance#Heritage Preservation Technology