Gains:
- Ability to keep the read character and the estimated completion separate and marked when transcribing inscriptions and manuscripts into draft text with OCR/HTR
- Ability to recognize forms of hallucination such as blurred text completion, false translation, and fabricated attribution, and to check the translation against the source and back-translation.
- Ability to understand how the context of language, writing and genre determines the reliability of artificial intelligence and that the final reading and interpretation belongs to the expert philologist
One of the most direct sounds left to us from the past is written sources: inscriptions carved in stone, clay tablets, papyri, parchment manuscripts, tombstones and seals. In this unit, we will see how epigraphy (the science of studying inscriptions engraved on hard surfaces such as stone, metal, ceramics), paleography (the expertise of reading and dating ancient handwritings) and AI accelerate the reading, translation and editing of these texts, but why every reading must be confirmed by an expert philologist.
Basic principle: AI helps read faint text, translate a language, and suggest possible complements; but the final reading of an inscription, the completion of missing letters and the meaning is the decision of the expert who knows the language and the period. It is in this area that the AI is at greatest risk of hallucination, as the model is very prone to filling in missing or faint text with something “plausible” but made-up.
Steps for working with written sources
1. Documentation. The inscription or manuscript is displayed in high resolution, contrast-enhanced format. For obscure surfaces, special lighting (side light) and techniques such as RTI are used; AI helps sharpen contrast and letter edges.
2. Character recognition. OCR (Optical Character Recognition) is used for printed text, and HTR (Handwritten Text Recognition) is used for handwriting. AI-based HTR is powerful at transcribing old manuscripts.
3. Transcription and completion. For faded or broken parts, the expert suggests possible restorations based on grammar and parallel texts. Completed letters are always indicated by signs such as square brackets; What is read and what is predicted are never confused.
4. Translation. The text is translated from the source language (Latin, Ancient Greek, Ottoman, cuneiform languages, etc.). AI gives first draft translation; expert nuance corrects terminology and historical context.
5. Comment. What the text says, who wrote it, and what event it refers to are historical interpretations and belong to the expert.
Tip: If you tell the AI to "read and complete" obscure text, it will safely make up the blanks. Instead, say "only give clearly readable characters, leave blank and mark the areas where you are not sure". Ask for completion in a separate step, with justification, and be sure to have it confirmed by the expert.
Hallucination: the biggest danger of this area
The most devastating mistake of AI in the humanities is fabrication. There are four typical forms in this field:
- Text fabrication: "Completes" the non-existent part of a faint inscription, adding a non-existent word.
- Fake translation: Translates a word he does not understand fluently but incorrectly; The reader cannot notice because he does not know the source.
- Fabricated reference: It says "This inscription was published in that publication", but that publication does not exist.
- Misdating/attribution: Confidently places a style of writing in the wrong period.
All of these can never be used without confirmation by an expert who sees the source and knows the language.
Attention: Just because a translation is fluent and convincing does not mean that it is correct. AI can even transform a cuneiform sign it doesn't understand into a confident sentence. Someone who cannot read the source text can never verify the accuracy of the AI translation on their own.
three mini cases
Case 1 — HTR opened the archive. One archive kept thousands of pages of Ottoman manuscripts closed to researchers because reading them required expertise and time. AI-based HTR transcribes notebooks into searchable draft text; Experts corrected the draft and unreadable areas were marked. It's not full text, but a powerful scanning tool has emerged.
Case 2 — Fabricated completion caught. A student had an AI read a broken Latin tombstone; The AI "completed" the missing line in a fluent sentence. The epigrapher showed that the proposed completion was physically impossible because there was no space left in the stone. The completion was reconstructed and marked based on parallel inscriptions.
Case 3 — Fake translation corrected. One team would report AI's translation of an Ancient Greek inscription; When the philologist checked, he found that the AI had translated a verb in the wrong tense, reversing the meaning of the text (whether it was an offering or a memorial). Going back to the source saved the comment.
Four copyable templates
1) Conservative transcription:
Your role: transcription assistant. In this inscription/manuscript image, give ONLY clear readable characters. Do not read areas that are unclear, broken or unsure; Mark it as "[unreadable]". COMPLETE THE MISSING PARTS. Also note the characters you suspect.
2) Reasoned completion proposal:
Below is a text with exact reading parts and [spaces]. Suggest POSSIBLE complements for the gaps, but for each suggestion: (a) justification (grammar, parallel text), (b) whether the gap fits the number of letters, (c) alternatives. These are suggestions; Not a definitive read. State that expert approval is required.
3) Draft translation (source protected):
Translate the [language] text below. Keep source text next to each sentence (align line by line). Mark the words you are not sure about and give an alternative translation. DO NOT make up the terms; If you don't know, write "uncertain". This is a draft; The expert philologist will check.
4) Translation verification (back translation):
Translate this translation BACK into the source language and compare it with the original source text. Mark the parts where the meaning is changed, added or dropped. Especially check the tense, subject and numbers.
Weak prompt / Strong prompt
Weak prompt:
Read this ancient inscription and tell us what it says.
In a faint inscription, the AI makes up the gaps, may mistranslate the language, and the person who does not see the source will not notice this.
Powerful prompt:
In this inscription image, first transcribe ONLY the clearly readable characters, leaving the obscured parts "[unreadable]". Then SEPARATELY, suggest possible complements for the gaps with justification but do not offer any definite. Finally, provide a draft translation and mark the words you are not sure about. All are subject to expert approval.
Difference: the first gives a fictitious "reading"; the latter keeps the read, predicted, and translated layers separate and verifiable.
Languages, scripts and the knowledge frontier of AI
Archaeological texts span a wide range of languages and scripts: Latin and Ancient Greek inscriptions, Ottoman and Arabic manuscripts, cuneiform tablets, Egyptian hieroglyphs, runic inscriptions, and more. AI proficiency on these languages and scripts is very uneven. While a draft translation may be reasonable for languages with abundant digital text, such as Latin and Ancient Greek, for texts that are sparsely documented, deciphered, or read by only a few experts, AI is almost entirely unreliable; produces convincing but completely fabricated "translations" in these areas.
There is also the issue of context and genre. The same word can have different meanings in a legal document, a prayer text, and a tombstone. The expert philologist recognizes the type of text (devotion, monument, contract, letter) and reads the word in the correct context; AI lacks this generic sensitivity and imposes the most common meaning. Abbreviations, formulas and stock expressions (especially very common in inscriptions) easily mislead the AI.
So the golden rule is: use AI as an accelerator for languages you know, languages you can control; Never rely on an AI translation alone in a language you cannot read. Someone who cannot read the source cannot verify the accuracy of the translation and therefore cannot turn that translation into a scientific claim.
Tip: When giving a text to the AI, specify its language, estimated period, and type (inscription/letter/document). Context knowledge partially reduces the AI's drift towards the most common but incorrect reading — but does not eliminate the need for confirmation.
Written resource task table
Quest
Role of AI
human decision
verification
image enhancement
Contrast/edge
Criteria selection
Comparison with original
OCR/HTR
draft text
indeterminate character
expert proofreading
completion
Recommendation (reasoned)
Accept/reject
Parallel text, location control
Translation
draft
nuance, term
Back translation, source
Comment
Context summary
historical meaning
expert, literature
Common mistakes
- Completing the faint text. AI fits the gaps; The reading and the prediction should be separated and marked.
- Trusting the translation without seeing the source. Fluent translation is not correct translation.
- Accepting the made-up attribution. Saying "It's on that broadcast" requires confirmation.
- Not checking the physical location of completion. The proposal must fit into the broken space.
- Leaving the interpretation to the AI. The historical meaning of the text is the matter of the expert philologist.
In summary
AI in epigraphy and ancient texts; It significantly speeds up image improvement, HTR/OCR draft, completion proposal and draft translation, and opens closed archives to research. But this is the area where the risk of hallucination is highest: reading is separated from prediction, complements are marked, translations are checked against source and back-translation, and the final reading and interpretation belong to the expert philologist.
Application task
Select an inscription or manuscript image (from an open access collection). Extract only clear characters with the "Conservative transcription" template, then get suggestions for spaces with "Reasoned completion suggestion". Produce a draft translation and back-translate it with the "Translation verification" template and mark where the meaning has shifted. Compare with published reading if possible.
checklist
- [ ] I marked the character read and the completion of the prediction separately.
- [ ] I checked that the completions fit in the broken area.
- [ ] I checked the translation against the source text and the back translation.
- [ ] I confirmed the publications/citations given by YZ in the actual catalogue.
- [ ] I left the final reading and interpretation to expert approval.