Gains:
- Ability to correct scanned documents by converting them to text with OCR and HTR and comparing the output with the actual image
- Ability to preserve the originality and integrity of the archive document and prevent AI from changing the content and fabricated restoration
- Ability to plan long-term digital preservation with provenance information and durable formats
An archivist is tasked with carrying fifty-year-old newspaper volumes, handwritten letters, or old photographs into the future. These documents wear out over time; To both preserve and make them accessible, digitization is done: the act of scanning a physical document and converting it into a digital copy. But screening alone is not enough; The digital object must be findable, readable and maintainable in the long term. In this unit, you will learn how to use artificial intelligence in the digitization, transcription and digital preservation workflow; but you will learn why you will hold the responsibility for originality and accuracy.
A few concepts. OCR (Optical Character Recognition) is a technology that converts the text in the image of a document into machine-editable text. HTR (Handwritten Text Recognition) does the same job for handwritten documents and is much more difficult. Digital preservation is a planned effort to ensure that digital objects remain readable for decades, even as file formats become obsolete and software changes. AI is a powerful but flawed assistant in all three areas.
Step by step: from digitization to preservation
1. Prioritize. You can't digitize everything at once. The most fragile, most requested and most unique documents are put first. This is a judicial decision; AI helps with list editing, but the institution determines the priority.
2. Scan and apply OCR/HTR. Scan the document at high resolution, then use OCR or HTR to extract the text. AI-based tools are getting better and better here, even with old and difficult texts.
3. Correct and verify the text. OCR/HTR output is never perfect. Errors are common, especially if there is old writing, stained pages, or a manuscript. It is essential to compare the output with the actual image and correct it.
4. Add metadata and protection information. Add to the digital object its provenance, scanning conditions, format information, and provenance — where the document came from and how it was processed. This information is critical for future verification and protection.
5. Plan for long-term protection. Choose file formats that are open, common and durable; back up; Make a migration plan in case formats become obsolete.
Tip: Don't accept the OCR output as "clear text". If a search system were to work through this text, incorrectly recognized words would cause the source not to be found. For critical collections, correcting the OCR output to the human eye determines the quality of all subsequent access.
Authenticity and integrity of the digital object
At the heart of archival work are authenticity and integrity: the assurance that a document truly comes from the claimed source and that its contents have not been altered. At this point, AI creates a two-pronged threat. First, during OCR/HTR correction or “enhancement,” the AI can silently modify the text; can add a word that does not exist under the name "correction". Second, if image "repair" or "restoration" is done with AI, non-original details may be produced.
The archivist's rule is clear: the digital preservation copy must be faithful to the original. Every improvement made with AI should be documented and the actual (raw) scan should always be preserved. "Beautifying" a document can destroy its historical evidentiary value.
Caution: It may be tempting to "improve" an old photo or document with AI; but AI fills in the missing parts by making them up. In an archival document, this means creating false historical evidence. The restoration output never replaces the original source; is stored as a separate, clearly labeled variant.
three mini cases
Case 1 — Newspaper archives became searchable. A library has digitized its 30-year-old collection of local newspapers. With OCR, 90,000 pages were transcribed and made searchable; A topic search that previously took days has now been reduced to seconds. However, the team noticed that OCR accuracy dropped to 80% on old and stained pages and prioritized correcting the top years with the human eye.
Case 2 — HTR opened the manuscript. An archive applied HTR to a collection of difficult-to-read 19th-century letters. The system gained great speed according to months of reading by experts; But each page of the printout was compared to the original by an expert paleographer (an expert on ancient writing), because there were errors in the names of people and places.
Case 3 — Fake “restoration” rejected. One team considered "repairing" a torn historical photograph with AI. AI completed the missing facial parts by fitting them. The archivist realized that these fabricated details would distort the historical evidence; The repaired image was stored separately alongside the original, only clearly labeled "AI derivative".
Access copy, preservation copy and format decision
A good practice in digitization is to separate copies of the same object for different purposes. A preservation master is the actual digital object to be carried into the future, stored in the highest resolution, uncompressed or lossless format; it is rarely touched. Access copy is a smaller, easily opened variant presented to the user. This distinction is important because the format that is practical for use (small, compressed) may not be suitable for long-term preservation; The ideal format for preservation (large, lossless) is cumbersome for daily use. In this workflow, AI can help in inventorying files, detecting missing variants, and keeping metadata consistent between two copies. However, the choice of format is a preservation decision: open, widely documented and durable formats should be preferred (rather than proprietary, closed formats), and a migration plan should be kept ready in case the formats become obsolete over time. Even if AI suggests a format, the institution's long-term preservation policy has the final say.
Tip: For each digital object, keep a brief record that answers the questions "where is the original source, in what format is the preservation copy, in what format is the access copy, and which is made available to the user?" Mixing these three layers makes it unclear which file is the trusted master going forward.
Four copyable templates
1) OCR output correction help:
Below is an OCR text and description of the actual page. Only fix openOCR errors (bad characters, compound/split words). Do not change content, add or remove words. If you are not sure, mark with "[?]". OCR text: [here]
2) Text-image consistency check:
Are there any additions in the corrected text below that were not expected to be in the original document (which may have been made up)? Mark each suspicious part and write why it is suspicious. Text: [here]
3) Protection metadata draft:
Generate the preservation metadata outline for the following digitized object: source, provenance, scan date/resolution, file format, transaction history. Just use the information I provided, leave the missing field blank. Information: [here]
4) Prioritization help:
Below is the list of collections awaiting digitization; Assist in prioritizing based on fragility, demand, and uniqueness. Give brief justification for each resource; Leave the final decision to me.List: [here]
Weak prompt / Strong prompt
Weak prompt:
Correct and beautify this scanned text.
“Beautify” invites AI to change the content, make up for the omissions; The originality of the archive document is damaged.
Powerful prompt:
Your role: digitization editor assistant. In the following OCR text, fix ONLY explicit recognition errors (bad character, incorrectly split word). Do not add, remove or change any words. Leave "[?]" where you're not sure. Show each correction you made in a separate list. Text: [here]
The powerful prompt limits the AI to only recognizing errors, prohibits it from touching content, and makes corrections traceable.
Workflow and risk table
Stage
Contribution of AI
Main risk
protection measure
scan
—
poor quality
high resolution
OCR/HTR
Convert to text
Recognition errors
human correction
correction
Error suggestion
Changing content
Comparison with original
restoration
Image derivative
fitting detail
Keep the raw original, label it
protection
Metadata draft
loss of origin
Provenance record
Common mistakes
- Accepting the OCR output without correction. Errors disrupt search and access.
- Calling AI “make beautiful.” It changes the content and destroys originality.
- Substituting AI restoration for actual evidence. Fabricated detail creates false history.
- Not storing the raw scan. No correction can be verified without the original.
- Omitting provenance information. The reliability of the document becomes untraceable.
In summary
AI in digital archives and digitization; It is a valuable aid that speeds up OCR/HTR transcription, metadata drafting and prioritization. But the essence of archival work is authenticity and integrity: OCR output should be corrected by the human eye, AI should never silently replace content, restoration derivatives should not replace the original evidence, and raw scanning and provenance information should always be preserved. Use AI as a tool that doesn't "improve" the document, but rather makes it accessible while staying true to it.
Application task
Get a short sample of OCR output (either your own production or from an actual scan). Have only the recognition errors corrected with the "OCR output correction help" template, then check whether the AI has added content with the "Text-image consistency check" template. Then create a record with the provenance and format information for the object with the "Preservation metadata draft" template.
checklist
- [ ] I compared the OCR/HTR output with the actual image and corrected it.
- [ ] I have checked that the AI does not add/change content.
- [ ] I kept the raw (master) scan separately and unchanged.
- [ ] I added provenance and format information to the metadata.
- [ ] I have clearly labeled restoration variants as "AI derivative".