Unit 12 / 12

End-to-End Project: Processing an Archive Collection with Artificial Intelligence

Gains:

  • Ability to process a collection in a 7-stage workflow, from ethical screening to publication, with a human checkpoint at each stage
  • Understand how early checkpoints prevent an error (wrong date/name) from propagating throughout the chain
  • Ability to apply the principles of preserving the master file, not saving on verification, and declaring the use of AI in a holistic project

Previous units have covered individual skills: digitization, transcription, metadata, scanning, translation, verification, pattern analysis, retrieval, preservation, and ethics. This final unit combines it all. A true archive project experiences these steps not separately but as an interconnected workflow. Here, we will see how you can process a collection from start to finish with an AI-supported but human-controlled process, step by step and with a control chain.

The goal is to position AI as a partner that “accelerates every step but has no final say on any step.” When you finish this unit, you will be left with a holistic method that transforms a messy pile of documents into a research-ready, verified, ethically clean collection.

Scenario: what we have

Let's say you receive a collection of 250 documents: some printed (printed correspondence, advertisements), some manuscript (letters, notebooks), from different dates, some containing sensitive personal data. Goal: to digitize, transcribe, catalog and make this collection available to researchers. Let's break down the process in seven stages.

Seven-step workflow

Phase 1 — Assessment and ethical screening. Get to know the collection first. Which documents contain sensitive personal data? What is the copyright status? Is there cultural sensitivity? This scanning determines which documents can be given to cloud-based AI tools and which will be processed only in a closed/approved environment. If this stage is skipped, the entire project is at ethical risk.

Stage 2 — Digitization. Scan documents at high resolution (at least 300-400 DPI, good contrast). Keep the original (master) files untouched; All processing will be done on the copies.

Stage 3 — Text extraction. Apply OCR for hardcopy documents and HTR for manuscript. Have the AI ​​clean up the raw output with “safe proofreading” instruction — optical/recognition errors only, no content comment, unreadable places [?]. Alternative reading and paleographer control for the Ottoman/manuscript.

Stage 4 — Metadata and cataloging. Have standard (Dublin Core/ISAD(G)) metadata drafted for each document; With permission from "leave non-existent field blank". Map topic terms to the controlled list. Verify dates and names with documentation.

Stage 5 — Enrichment and patterning. Extract entities (person/place/organization), normalize, produce relationship and timeline outlines. Verify each pattern in the document; There are no contrived connections and no chronology.

Stage 6 — Access tools. Produce finding aid outline for collection; Base the history section only on verified notes. Clearly mark access restrictions (privacy/copyright).

Stage 7 — Verification, declaration and publication. Verify critical fields (date, name, quote, reference) one last time. Declare use of AI. The expert gives final approval; The collection is opened for research.

Tip: Put a "human checkpoint" at each stage: An expert validates the AI's output before moving on to the next stage. Without these checkpoints, an error at the first stage (a misread date) will propagate throughout the chain and eventually leak into the catalog, pattern, and guide.

Control chain table

Stage

AI's job

human checkpoint

1. Ethical screening

Sensitive data marking

Installation decision

2. Digitization

(Image enhancement)

Quality, master protection

3. Text extraction

OCR/HTR + secure correction

Name/date/reading verification

4. Metadata

standard draft

Field confirmation, term matching

5. Pattern

entity, relationship, timeline

Validation in document

6. Access tool

Finding aid draft

History and constraint approval

7. Release

Final approval, declaration

three mini cases

Case 1 — Chain error caught. In one project, in phase 3, the date of a document read "1868" instead of "1888". If the stage 3 checkpoint had not caught this error, the date would have been spread across the catalog (4), the timeline (5), and the guide (6). Thanks to the checkpoint, the bug was fixed in one place. Lesson: early checking prevents the error from propagating up the chain.

Case 2 — Ethical screening saved the project. 18 of the 250 documents contained sensitive data of living persons. Ethical screening in phase 1 flagged these; these 18 documents were processed in a closed environment, the rest with standard tools. It would have been a serious breach if the scanning was skipped and it was all uploaded to the cloud. Lesson: ethics screening is the first step, not a fix added later.

Case 3 — Time and effort. Using the traditional method, it would take approximately 8-10 months to fully process a collection of 250 documents. With the AI-supported, checkpoint workflow, the process was reduced to approximately 3 months. But the savings came not from “AI did everything,” but from “AI did the mechanical work, focused on human verification.” The time devoted to verification has not decreased; The time devoted to mechanical work has decreased.

Four copyable templates

1) Project initial plan:

I have a collection of [number] documents: [mix of print/manuscript], [period], some of which may contain sensitive data. Create a 7-step workflow plan to digitize and catalog this collection: specify the AI's task and the HUMAN checkpoint at each stage. Put ethics screening first.

2) Stage transition checklist:

I finished stage [stage name]. Produce a checklist of critical points I need to verify before moving on to the next stage: what facts (date/name/reference) need to be verified, what errors can propagate up the chain? I have the final approval; You give me the checklist.

3) End-to-end quality control:

Below are all the layers of a processed document: transcription, metadata, extracted assets. FIGURE INCONCEPTS between layers: does a date in the transcription match the metadata, do the entities actually exist in the text? Imposition of correction; Generate a list of inconsistencies. Layers: [here]

4) Project closing and declaration:

In this collection project, I used AI in the following stages: [list]. For project documentation: (1) what AI support was used at which stage, (2) how was human verification done, (3) what limits/constraints are there — write an honest closing note and draft AI use statement that includes these.

Weak prompt / Strong prompt

Weak:

Process this collection thoroughly with AI and make it ready for research.

A single “do anything” instruction bypasses ethical screening, checkpoints, and verification; errors propagate along the chain.

Strong:

Set up a 7-stage workflow for this collection, with a human checkpoint at each stage. Let the 1st stage be ethics/privacy screening. The output produced by the AI ​​at each stage will be verified before moving to the next stage; mark critical facts (date/name/reference). At no stage does the AI ​​have the final say.

The difference: phased structure, ethics priority, checkpoints and the “people have the final say” principle make the entire project safe and defensible.

Common mistakes

  • Leaving ethical screening to the end. Privacy/copyright scanning is the first step; If it is added later, the violation has already occurred.
  • Bypassing checkpoints. An error that goes unverified spreads throughout the chain.
  • Not protecting the master file. If the original is not kept untouched, return becomes impossible.
  • Saving on verification. AI shortens mechanical work; If the verification time is shortened, the quality decreases.
  • Not declaring the use of AI. Transparency is part of the scientific credibility of the collection.

In summary

An end-to-end archive project connects individual skills into a chain of control: ethical scanning, digitization, text extraction, metadata, pattern, retrieval tool, and publication. At each stage, AI speeds up mechanical work; At each stage, a human checkpoint has the final say. Ethical screening is the first stage, the master file is protected, errors are caught early so they do not spread up the chain, and the use of AI is declared honestly. The savings come not from the AI ​​doing everything, but from the human being freed from mechanical work and focusing on verification. AI accelerates; Man ensures the accuracy and integrity of history.

Application task

Select a small collection (10-20 documents; your own material or a sample set). Create a 7-step workflow with the "project startup plan" template. Run at least two documents through this flow: ethical scanning, text extraction, metadata, and a pattern layer. Verify each stage with a “stage transition checklist.” Finally, document your use of AI with the “project closure and statement” template. Note which errors were caught at early checkpoints.

checklist

  • [ ] I did the ethics/privacy screening at the first stage.
  • [ ] I kept the master files untouched.
  • [ ] I put a human checkpoint at each stage.
  • [ ] I verified critical facts before they were propagated up the chain.
  • [ ] I declared the use of AI at the project closing.

Module Exam

1. What is the most appropriate positioning for artificial intelligence in history and archiving?

  • A) AI is a summary, editing and drafting tool; Comment, source criticism and final responsibility belong to the expert ✔
  • B) Artificial intelligence can decide the authenticity and meaning of the document on its own, like a historian
  • C) Artificial intelligence only works for translating text, it has nothing to do with other archive tasks
  • D) Since artificial intelligence is always more impartial than humans, historical interpretation should be left to it.

Description: Artificial intelligence; It is an assistant that edits, scans, translates and produces drafts of text. Source criticism, interpretation, cause-effect establishment and final decision belong to the expert; An unverified output may lead to a fabricated source, a false date, or an anachronism.

2. What is the hallucination produced by artificial intelligence in historical research?

  • A) Artificial intelligence processes a document very slowly
  • B) Artificial intelligence fabricates a non-existent source, quote or event as real ✔
  • C) A document is physically damaged
  • D) An archive catalog is not up to date

Explanation: Hallucination is when artificial intelligence confidently fabricates information, sources, quotes or events that do not actually exist. Although its form appears perfect, its content is not real; Therefore, in every reference catalogue, every citation must be verified at the original source.

3. What is the best approach to having an OCR output corrected by artificial intelligence?

  • A) Asking the text to be fluent and beautiful and to logically complete the missing parts.
  • B) Allowing it to correct all numbers and dates based on context
  • C) Asking him to correct optical recognition errors only, leave unreadable areas [?] and not change numbers/dates ✔
  • D) Asking the student to fill in the unreadable words with the most likely guess.

Explanation: The golden rule in OCR correction is to correct only obvious optical recognition errors (such as rn to m, l to 1); not interpreting or completing the content of the text. Unreadable areas are marked with [?]; The numbers and dates are not changed, they are verified with the image.

4. What should be asked from artificial intelligence when it comes to reading a word whose vowel is not written in an Ottoman text?

  • A) Give a single, precise reading, thus speeding up work
  • B) Selecting the word that will make the meaning most fluent and filling in the blanks
  • C) Completing the unreadable parts from his own general knowledge.
  • D) Offering possible reading alternatives and marking ambiguous areas instead of imposing a definitive reading ✔

Explanation: In Ottoman Turkish, short vowels are mostly not written; A single vowel/letter difference (abul and kabul, wali and vali) completely changes the meaning. Therefore, instead of imposing a definitive reading, possible alternatives should be asked, unreadable areas should be preserved, and the final decision should be left to the paleographer.

5. What is the most effective way to prevent hallucination when having AI produce a draft metadata (catalog record)?

  • A) Explicitly grant permission to leave that field blank if it is not included in the document ✔
  • B) Requesting that every field be filled in and not left incomplete.
  • C) Estimating a date based on the content of undated documents
  • D) Allowing artificial intelligence to generate the reference code

Description: Fill-in pressure forces the AI to make up the document by guessing the date or creator that is not in the document. The strongest shield against this is to explicitly grant permission to leave the field blank if it is not in the document; Additionally, topic terms are mapped to controlled vocabulary.

6. How should a document summary produced by artificial intelligence be positioned in historical research?

  • A) As reliable evidence that can be directly cited
  • B) As a discovery map that directs you to the actual document; Evidence is taken from the original document ✔
  • C) As an adequate resource that can replace the document itself
  • D) As a text that can be used in the study without ever returning to the document

Description: The summary is a discovery map, not a proof. There is something like this in this document, it says go and look; It does not say exactly what it says in the document. If an event, quote or data is to be included in the study, it must be taken verbatim from the original document.

7. What is the proper way to avoid anachronisms when translating a historical document for publication?

  • A) Replacing all period terms with today's closest equivalent
  • B) To convey the text as fluent and modern as possible
  • C) Leave the period titles and institution names unique and add an explanation ✔
  • D) Adapting proper names to their current usage

Explanation: Anachronism is the wrong transfer of elements from one period to another period. Changing period titles, institution names and personal names to today's terms transports the reader to a world that does not belong to the period. The correct way is to keep the period term and add explanation next to it.

8. How to make sure that an archive reference (fund name, file number) given by artificial intelligence is real?

  • A) Trusting the reference because its format seems correct
  • B) By asking the AI again if the reference is correct
  • C) Accepting the reference as true because it is convincing and detailed
  • D) By personally finding and verifying the reference in the archive catalog or reliable source ✔

Description: Artificial intelligence can produce references that are perfect in form but do not actually exist; A fictitious reference is in the same form as a real reference. The only way to prove that a reference exists is to find it where the original source is located (archive catalogue, library).

9. How should a pattern found when performing pattern analysis (NER, network, timeline) with artificial intelligence be evaluated?

  • A) As a hypothesis that needs to be verified in the document; starting point, not outcome ✔
  • B) As a definitive finding that does not require verification
  • C) It is an undisputed objective fact because it is numerical.
  • D) As a result that can be put into direct study because it found artificial intelligence

Explanation: Every pattern is a hypothesis, not evidence. Entities should be linked to the source, different spellings (of the same person/place) should be normalized with human approval, numbers should be questioned for missing data and sampling bias, and undated events should not be chronologized.

10. Why does the biographical/institutional history section in a finding aid require special attention?

  • A) This section is unimportant and can be published without verification
  • B) Since artificial intelligence can add fabricated events, false dates and titles in this section, it should only rely on verified sources ✔
  • C) Artificial intelligence can be used safely because it does not make any mistakes in writing history.
  • D) Only researchers fill this section, the archivist is not interested

Description: The history section is where hallucination most often creeps in: the AI can insert non-existent events, false dates, and made-up titles while writing a fluent text from general information about a person or institution. This section should be based on verified sources only.

11. How should artificial intelligence answer a researcher's fact question in a reference (consultation) service?

  • A) The answer to the question should be given directly and as a definite fact.
  • B) Must produce a believable date or name and convey it to the researcher.
  • C) Instead of giving facts directly, it should direct the researcher to the relevant collection and source ✔
  • D) It should provide a complete answer so that the researcher does not need to go to the source.

Explanation: In the reference service, artificial intelligence should not give a direct fact answer, but should direct the researcher to the source. Because the direct factual answer given by artificial intelligence may be fabricated and the researcher may mistake it for the source. Instead of saying "The answer is this", it should say "Look at these documents in this collection".

12. Which operations are acceptable and objectionable when processing a damaged document image with artificial intelligence?

  • A) All kinds of operations are free, including filling in the missing parts.
  • B) No action should be taken, the image should never be touched
  • C) Filling in the missing sections is the most valuable action because it completes the document.
  • D) Readability manipulations such as contrast/noise are acceptable; It is inconvenient to fill in the missing parts with guesses ✔

Explanation: Processes that improve readability (contrast, lighting, noise reduction) are acceptable because they do not change the content; However, filling in missing/unreadable parts with guesswork is the risk of producing what is not in the document, that is, historical forgery. The original is preserved, transactions are recorded.

13. What needs to be done before processing an archive document containing sensitive personal data?

  • A) Privacy scanning and anonymization if necessary or using a closed/approved tool ✔
  • B) Uploading the document directly to a public vehicle without even thinking about it
  • C) Ignoring privacy concerns for efficiency
  • D) Overcoming access restrictions for efficiency reasons

Disclosure: Uploading documents containing sensitive data identifying living persons (health, religion, ethnicity, address) to public, cloud-based tools can be a serious privacy violation. Privacy screening should be done first, anonymization or a closed/approved tool should be used if necessary.

14. What is the main purpose of a human checkpoint in an end-to-end archive project?

  • A) Slowing down the work of artificial intelligence and delaying the project
  • B) To prevent an error (incorrect date/name) in one stage from being chained to all subsequent stages ✔
  • C) Increasing human labor and giving up automation completely
  • D) Ensuring that artificial intelligence has the last word

Description: A human checkpoint at each stage ensures that the AI's output is verified before moving on to the next stage. Without this, an error at the first stage (a misread date) would propagate up the chain through the catalog, pattern, and guide. Early checking ensures that the error is corrected in one place.

15. What do transparency and authenticity require in the use of artificial intelligence in humanities and social sciences?

  • A) Keeping the use of artificial intelligence secret, ensuring that no one knows
  • B) To present every text produced by artificial intelligence as its own original idea
  • C) To honestly declare the use of artificial intelligence and not to present its output as one's own original work ✔
  • D) Sharing only the results without explaining the methods

Description: Transparency includes recording that a transcription, metadata or translation was produced with the help of artificial intelligence; Originality requires not presenting a comment produced by artificial intelligence as one's own original idea. This is essential for verification by subsequent researchers and for academic integrity.