The Archivist and Artificial Intelligence: An Ally, Not a Replacement

On trust, provenance, and the future of the photographic document

There is something ironic about the situation archivists find themselves in when facing artificial intelligence. These are professionals whose work consists of preserving memory, documenting what has existed, ensuring that traces of the past remain accessible and understandable for future generations. And yet here they are, being offered tools capable of reading, analyzing, classifying, and describing thousands of documents in seconds — tasks that once took them weeks.

The question is not whether artificial intelligence will transform archival work. It already is. The real question is how — and above all, for whose benefit.

This text explores what AI can concretely bring to archivists and heritage institutions, without ever losing sight of what no algorithm will ever replace: professional judgment, documentary ethics, and human responsibility toward collective memory.

1. What an archivist does — and what AI does not

Before going further, things must be named clearly. The work of an archivist is not simply “classifying documents.” It is a complex intellectual act that draws on knowledge of history, law, diplomatics, palaeography, description standards (RDDA, ISAD(G), Dublin Core), and often multiple languages. It is also a work of contextualization: a document only makes sense when placed back within its chain of creation, its provenance, its original fonds.

The principle of provenance — respect des fonds in the French tradition — is at the heart of archival practice. It establishes that documents produced or received by the same creator must be kept together, in their original order when possible. It is this context that gives meaning to information. A document removed from its fonds becomes an orphaned piece of data — intelligible perhaps, but impoverished.

Artificial intelligence, in its current state, does not understand provenance. It can read a text, identify named entities, suggest an approximate date, detect a language. But it does not know why this document exists, who produced it, with what intention, in what relationship to the other documents from the same creator. That knowledge belongs to the archivist.

What AI can do, on the other hand

It can process volumes that exceed human capacity. A photographic collection of 50,000 images can be browsed, analyzed, and pre-described in a few hours. Recurring fields — approximate dates, visible locations, general subjects, condition notes — can be filled in automatically, leaving the archivist to validate, correct, and enrich.

It can identify redundancies, flag anomalies, suggest thematic groupings. It can transcribe handwritten texts with increasing accuracy. It can translate metadata from one standard to another. It can make searchable collections that previously were not.

None of this is trivial. It represents hours, weeks, sometimes years of repetitive work that can be redirected toward what truly matters: analysis, contextualization, and mediation with researchers and the public.

2. Automated description: promise and limits

Archival description is one of the most time-consuming and standardized tasks in the profession. For photographic collections in particular, it involves filling in dozens of fields for each image or series of images: title, date, location, subject, photographer, rights, condition, reference number, and so on. IPTC standards and frameworks such as RDDA or Dublin Core provide structure, but filling in that structure remains a considerable manual effort.

This is precisely where artificial intelligence can offer substantial time savings. Computer vision models are now capable of analyzing an image and producing a detailed textual description: identification of individuals (if known), recognition of locations, estimation of the period, detection of salient visual elements. This data can automatically populate metadata fields, subject to human validation.

“AI proposes. The archivist disposes.”

This simple formula captures the philosophy that should guide the integration of these tools into professional practice. Automated description is not a final description — it is a first draft, a working hypothesis, which the archivist will examine, correct, and validate based on their expertise and knowledge of the collection.

Risks not to be underestimated

The first risk is that of abusive normalization. An algorithm trained on Western, urban, or contemporary corpora will perform less well on colonial, rural, or historical archives. It may produce ethnocentric, anachronistic, or simply inaccurate descriptions. Archivists must remain vigilant, especially since these errors can replicate at scale if not corrected from the outset.

The second risk is that of the dispossession of professional judgment. If archivists are asked to validate thousands of automatically generated descriptions without being given the time to genuinely examine them, the problem has simply been displaced. Speed must not take precedence over rigor.

Finally, there is the question of responsibility. Who is accountable for an erroneous description produced by AI? The institution? The software provider? The archivist who validated it? These questions have not yet been resolved legally, and they deserve serious attention.

3. Photography: a special case

Photographic collections represent a specific archival challenge. Unlike textual documents, a photograph is not read — it is looked at. It engages visual interpretation, knowledge of the context in which it was taken, and an understanding of the aesthetic codes and photographic practices of an era. It may also carry emotional, political, or ethical dimensions that further complicate description.

Heritage institutions accumulate thousands, sometimes millions of images, often poorly documented, sometimes anonymous. Press agency collections, photographers’ personal archives, institutional holdings — all of this represents an immense visual heritage, partly inaccessible for lack of resources to describe it properly.

Artificial intelligence opens real possibilities here. Scene recognition, face detection, identification of iconic locations, estimation of dates from clothing styles or the built environment — all of this can help produce a first level of description that makes a collection searchable and usable.

But who took this photograph? And is it real?

Two questions arise with growing urgency in today’s digital context. The first concerns attribution: who made this image? Documentary photography rests on a bond of trust between the photographer, the institution that holds the image, and the public that consults it. That bond requires knowing where the image comes from, who produced it, and under what conditions.

The second question is that of authenticity. At a time when AI image-generation software can produce realistic photographs of scenes that never existed, the boundary between document and visual fiction has become porous. For archivists, journalists, and historians, this is a fundamental challenge.

Both questions — attribution and authenticity — are at the heart of an emerging technical standard that is beginning to transform the chain of trust around the photographic image.

4. C2PA: rebuilding trust in the digital image

C2PA — the Coalition for Content Authenticity and Integrity — is a consortium founded by Adobe, Microsoft, Intel, the BBC, Reuters, and dozens of other organizations. Its objective: to establish an open standard for certifying the origin and history of digital content, whether a photo, a video, or an audio document.

In practice, C2PA allows a set of cryptographically signed metadata — known as Content Credentials — to be embedded directly in an image file. These credentials attest to the image’s origin, the identity of its creator, the equipment used (camera, software), and all modifications made to it since its creation.

It is the digital equivalent of a documentary chain of custody — a concept archivists know well. Every intervention on the document is traced, timestamped, and identified. The integrity of the chain can be verified at any time by anyone with the appropriate tools.

Why this matters for heritage institutions

For an institution that holds press photographs, documentary images, or visual archives, C2PA represents a tool of primary importance. It makes it possible to distinguish an authentic photograph from a generated or manipulated image, to attest to an image’s provenance at the source — at the very moment of capture, if the camera integrates the standard — and to reconstruct the processing history of an image: cropping, color adjustment, format conversion.

When an institution publishes an image on its website, in a virtual exhibition, or in a digital publication, it can now attach a verifiable certification attesting to the image’s origin and the accuracy of its description. For the researcher or citizen consulting an online archive, this information becomes transparently accessible: a small icon, one click, and one can verify that this photograph was indeed taken by a given photographer, on a given date, with a given camera, without significant manipulation since. In a context of growing mistrust toward digital visual content, this ability to certify an image’s authenticity becomes a genuine tool for cultural mediation — and an irreplaceable argument for trust.

Current limitations

It would be naïve to present C2PA as a universal solution. The standard is still young, its adoption by mainstream cameras is ongoing, and images produced before its deployment obviously do not carry these origin certifications. For historical archives, complementary approaches will need to be developed — visual expertise, forensic analysis, cross-referencing of sources — to establish the authenticity of documents.

Furthermore, a digital signature can be stripped from an image. C2PA does not make an image unfalsifiable — it makes falsification detectable. That is no small thing, but it is not an absolute guarantee either. Trust remains, in the final analysis, a human and institutional construction.

5. The archivist of tomorrow: an augmented professional

What is taking shape is the profile of an archivist whose role evolves without disappearing. Repetitive and mechanical tasks — data entry, format conversion, mass indexing — are progressively being handled by automated tools. What remains, and what grows in importance, is the irreducibly human part of the profession.

The archivist of tomorrow will be the one who knows how to ask the right questions of AI, who can interpret its outputs with a critical eye, who can identify its blind spots and correct its biases. They will also be the one who knows how to explain to the institution what AI can and cannot do — who protects collections from misdirected technological enthusiasm.

They will ultimately be the one who maintains the link between the document and its human context. Because a photograph is not merely a JPEG file with IPTC metadata. It is the result of a human presence at a given moment in the world — a moment no one else witnessed in quite the same way.

New skills to develop

Archival education will need to evolve to incorporate these new tools. Not to train developers, but to train professionals capable of understanding what an algorithm does, of evaluating the quality of its outputs, and of making informed decisions about its use in any given context.

Knowledge of metadata standards — IPTC, XMP, Dublin Core, RDDA — becomes even more strategic in this context, since it is on these standards that automated systems rest. Understanding their logic, their strengths, and their limits means understanding what AI can truly do with the data it is given.

Knowledge of digital provenance issues — C2PA, chains of custody, image forensics — is likewise becoming essential for any institution managing contemporary photographic collections.

Conclusion: memory is a human responsibility

Artificial intelligence is an extraordinarily powerful tool. It can process what we cannot process, see what we do not see, suggest what we would not have thought of. In the archival context, it opens real possibilities for making accessible collections that have lain dormant for lack of resources.

But a tool, however sophisticated, carries no responsibility. It does not decide what deserves to be preserved. It does not understand why a photograph taken on a Montreal street in 1967 is an irreplaceable document of Quebec’s social history. It does not know that this image, poorly described, could disappear into digital oblivion even though it may be the only visual trace of a singular moment.

Collective memory is a human responsibility. The archivist is the guardian of that responsibility. Artificial intelligence can help exercise it better, more broadly, more effectively. But it cannot carry it in their place.

That may be the fundamental distinction: AI optimizes. The archivist judges. And in a world where images are manufactured as fast as they spread, where the boundary between the true and the plausible grows harder to trace each day, that judgment has never been more valuable.

Do you work within a heritage institution — archives centre, museum, historical society, library? Discover how DIGITUM Archiviste IA can transform the way you describe and showcase your collections.