Gains:
- Ability to produce metadata drafts in accordance with standards such as ISAD(G) or Dublin Core with artificial intelligence
- Ability to prevent hallucination and map subject terms to controlled vocabulary with leave-blank permission
- Ability to distinguish which fields, such as date, person name and reference code, require human approval and which should only be assigned by the institution
An archive or library, no matter how rich it is, cannot be found unless it is well defined. What makes a document “findable” is the descriptive information about it: who wrote it, when, where, what its subject is, what collection it belongs to. This information is called metadata (metadata — descriptive data about a document; “data about data”). Cataloging is the process of identifying documents with this metadata and saving them in a searchable system. Classification is to organize documents according to criteria such as subject, type or origin.
Generating metadata is extremely time-consuming; Entering date, subject, person and place information individually for thousands of documents in large collections can take years. AI is a strong draft generator in this business. In this unit, we will see how to generate metadata with AI, cataloging in accordance with standards (ISAD(G), Dublin Core), and both the power and pitfalls of automatic classification.
Archival standards: why they are needed
Metadata is not entered randomly; There are international standards so that records from different institutions can talk to each other:
- ISAD(G) (General International Standard Archival Description): Rules for describing archival material hierarchically (at the fund, series, file, document level). It contains fields such as reference code, title, date, scope, creator.
- Dublin Core: A common standard consisting of 15 core fields for describing digital resources (title, creator, subject, date, genre, format, source, language, etc.).
- Fonds: A set of records naturally formed as a result of the activities of a single person, family or institution. It is the basic organizing unit of archiving.
- Provenance (origin): Information about who a document comes from and which hand it passed through. In archiving, documents are kept together according to their origin.
Explicitly specifying the target standard when having the AI produce metadata ensures that the output is directly inputtable into the enterprise.
Another key concept is authority record—a record that defines a single, standard, approved form of the name of a person, place, or institution. In historical documents, the same person or place may appear with different spellings; The authority record binds them all into a single approved format so they don't get scattered in the search. AI suggests names from a document; but mapping this name to the standard form in the authority record is human work. Linking what the AI says "Ahmed" to "Ahmed Bey (1840-1912)" in the institution's authority record preserves the consistency of the catalogue.
Workflow: step by step
1. Prepare the source. Collect a transcript or image of the document and any context information.
2. Identify the standard and areas. Tell the AI which standard (ISAD(G), Dublin Core) and according to which fields it will produce output.
3. Generate draft metadata. AI fills in fields that can be extracted from the document (date, person, place, subject, language, genre); It leaves blank the areas it cannot remove.
4. Map to controlled vocabulary. Subject and place names are not free in most institutions, but are tied to a predefined list (controlled vocabulary / thesaurus). Map the subject terms suggested by the AI to this list.
5. Verify and confirm. Each field, especially date, name and subject, is compared by human beings with documents and records of authority.
Tip: Explicitly give the AI permission to "leave blank". Otherwise, it will "guess" fill in a date or creator that isn't in the document and insert fake metadata into your catalog. The instruction "Leave this field blank if it is not in the document" is the strongest shield against hallucination.
Fields and trust level in a table
Metadata field
How reliable is AI?
verification
language
high
sample
Document type
medium-high
control
Person/place names
Medium (reading error)
mandatory
date
Low-medium (risk of fabrication)
mandatory
Subject terms
Medium (free generating)
Map to checklist
Scope/summary
medium-high
Compare with source
Reference code
None (institution gives)
human throws
Summary: AI is fast and reliable in areas such as language and genre; generates drafts in fields such as date, name and subject, but human approval is required; AI never produces institutional fields such as reference code.
three mini cases
Case 1 — Draft metadata for 8,000 photos. A city archive was cataloging 8,000 historical photographs. The notes on the back of each photo were transcribed and given to the AI; AI generated date, location and topic outline. The human team verified it field by field and mapped it to the controlled list. An estimated 10-month job was reduced to 3 months; but every date and place name was visually checked.
Case 2 — Made-up history. An archivist had AI catalog an undated letter. YZ looked at the content of the letter and wrote the date "1912" precisely. However, there was no date in the document; AI predicted. If the archivist had not added the instruction "leave blank if not in the document", the catalog would have been filled with a made-up date. Lesson: blank permission is mandatory.
Case 3 — Free subject terms created confusion. AI assigned similar but different free terms such as "immigration", "immigration", "settlement" to similar documents. When it was not mapped to the controlled vocabulary, the same topic was tagged with three different terms and scattered in the search. Consistency was achieved by mapping the terms into a single authority list. Lesson: topic terms are not released, they are linked to the list.
Four copyable templates
1) Dublin Core outline:
Extract the Dublin Core metadata draft from the document transcription below. Fields: Title, Creator, Subject, Description, Date, Genre, Language, Scope (location/period). Rules: (1) Only use information CLEARLY in the text. (2) Leave the field that is not in the text BLANK, do not guess. (3) Mark the value you are not sure of with [?]. Transcription: [here]
2) ISAD(G) draft definition:
Generate draft for the following document in ISAD(G) fields: Title, Date(s), Description level (document/file), Scope and content (short summary), Creator, Language. Leave the reference code BLANK (the institution will provide it). Do not add any information that is not in the text; Leave the non-existent field blank. Document: [here]
3) Topic term suggestion (controlled):
Suggest 3-5 topic terms appropriate to the document below. First, rely on the facts IN the document. Below is our institution's controlled terminology list; FIRST, choose it from this list, if it is not in the list, mark your new term suggestion separately. List: [terms]. Document: [here]
4) Bulk consistency check:
Below are metadata records for documents from the same collection. Flag inconsistencies: records spelling the same place/person differently, different date formats, different terms for the same subject. Correction IMPOSITION; Just produce a list of inconsistencies for the human to review. Records: [here]
Weak prompt / Strong prompt
Weak:
Create a complete catalog record for this document.
The “complete” instruction pushes the AI to fill in every space, i.e., to invent history and creators that do not exist.
Strong:
The Dublin Core is drafted for this document. Use only information CLEARLY included in the transcription; leave the non-existent field blank, don't guess.Select topic terms from this controlled list: [list]. Mark the date and person names with [?] so I can verify it with the document. Transcription: [here]
Difference: "leave blank" permission instead of "complete" edition, controlled list adherence and verification marks protect the catalog from fabrication.
Common mistakes
- It means "fill it completely". This leads to fitting non-existent fields; "leave blank" permission should be given.
- Releasing subject terms. Terms that do not map to the controlled vocabulary will distract the search.
- Having AI predict history. Entering a fictitious date into an undated document pollutes the catalogue.
- Having the reference code generated by AI. This is a corporate, unique space; people throw.
- Not specifying the standard. If the standard is not mentioned, the output will be a messy draft that cannot be entered into the institution.
In summary
Metadata and cataloging are the invisible infrastructure that opens the archive to the researcher, and it requires a lot of effort. From the transcription, AI produces a rapid outline for fields such as date, person, place, subject, and genre; It can work according to standards such as ISAD(G) or Dublin Core. But the pressure to “fill in completely” pushes the AI to make things up; “leave blank” permission, mapping to controlled vocabulary, and human confirmation of dates/names keep the catalog trustworthy. AI prepares the draft; You are responsible for the accuracy of the catalogue.
Application task
Take a document transcription and request metadata from the AI with the “Dublin Core draft” template; Check that it leaves blank fields that do not exist. Then run the “topic term suggestion” template with your own (or sample) checklist. Finally, test a few records with a “bulk consistency check.” Note how many fields the AI tries to populate with prediction and how many subject terms need to be mapped to the list.
checklist
- [ ] I have clearly stated the target standard (ISAD(G)/Dublin Core).
- [ ] I gave the "Leave blank field blank" permission.
- [ ] I mapped the topic terms to the controlled list.
- [ ] I verified the date and person names with the document/authority record.
- [ ] I left the reference code to the institution, not to AI.