Unit 5 / 12

Source Scanning, Discovery and Summarization

Gains:

  • Ability to accelerate discovery by scanning large piles of documents with artificial intelligence, eliminating and summarizing what is relevant
  • Understand that the summary is a discovery map, not a proof, and be able to link each important point and quote back to the original document.
  • Ability to scan cases and check elimination with samples without having artificial intelligence produce results/comments

One of the biggest bottlenecks in historical research is finding "what works" in a huge pile of material. A researcher may have to scan hundreds of documents, newspaper volumes, registers, or correspondence files for a thesis or article. This resource review task traditionally requires reading page by page and can take weeks. AI opens up this bottleneck by quickly scanning and summarizing large chunks of text and marking relevant places. But here too the danger is clear: the AI ​​summary does not replace the source; it is merely a map that directs you to the real source.

In this unit, you will learn how to scan large historical material with AI, produce reliable summaries, find relevant passages, and avoid the most dangerous pitfall of synopsis: "trusting the synopsis without going to the source". The goal is to use AI for speed while never sacrificing accuracy; When used well, these two goals do not conflict, but rather nourish each other.

Summary does not replace the source

Let's lay out the basic principle from the beginning: The AI ​​brief is a discovery tool, not a proof. If a historian is going to include an event, a quote, or a piece of data in his study, he has to take it from the original document, not from the AI's summary. The summary says, "there's something like this in this document, go look at it"; He doesn't say "this is exactly what it says in the document". Because the summary process compresses, selects, sometimes distorts and sometimes adds something that doesn't exist at all.

You can use AI in three powerful jobs in scanning:

  1. Scan and sift: Quickly mark among hundreds of documents which are relevant to your topic.
  2. Summarizing: Outlining the outlines, people, and topics of a long document.
  3. Finding a passage: Marking the question "Where are the places that talk about this topic" in the document.

In all three, the output is a "redirect"; The evidence is taken from the original document.

An overlooked danger here is the limit of the AI's context window — the amount of text the model can process at one time. When you give a very large stack of documents at once, the AI ​​may “forget” or skim through the middle sections; tends to give emphasis to what happens at the beginning and end. Therefore, instead of fitting hundreds of documents into a single request, dividing them into logical groups and scanning each group separately gives more reliable results. Additionally, sample checking is essential to distinguish whether a document that AI "never mentions" is truly irrelevant or omitted.

Workflow: step by step

1. Clarify the research question. A clear question rather than a vague “summarize this to me”: “Where are the places in these 40 documents that mention the water rights dispute?”

2. Prepare the ingredients. Edit the texts to be scanned (transcription or OCR output). Attach the source's name/reference to each piece so that the summary can be linked to the source.

3. Give scan/summary instructions. Ask the AI ​​to rely solely on the text you give, linking each claim to the source (document/page) and not adding anything that doesn't exist.

4. Link the summary back to the source. Find and verify in the original document every important point marked by the summary. Take quotes verbatim from the document.

5. Question gaps and bias. What might AI have missed? Did he just highlight the obvious? Did he choose the one that suits your question and hide the opposite?

Caution: When generating a summary, the AI ​​may add a link or result that is not in a document because it "seems logical". Especially "what conclusion can be drawn from these documents?" Questions like these invite the AI ​​to make up interpretations and inferences. Keep the summary request factual; keep the comment to yourself.

Scanning tasks with a table

Quest

Contribution of AI

Risk

verification

Eliminate the relevant document

speed

Skip relevant

sample control

Document summary

outline

Warp/add

Compare with source

Finding a passage/quote

bookmarking

wrong location

find in document

Comparative summary

pattern marking

fitting link

Confirm every claim

Quoting

draft

word substitution

word for word

Summary: AI shows where it is; The actual document tells what happened.

three mini cases

Case 1 — 300 documents reduced to 2 days. A researcher was studying a town's 1900-1920 water rights dispute. Scanning a series of 300 documents by hand would take weeks. AI marked each document as “is it related to water rights”; 42 out of 300 documents turned out to be relevant. The researcher read only these 42 in depth. But he also checked 20 of the 258 documents marked as irrelevant as a sample; Found that 2 were actually related and fixed it. Lesson: elimination speeds up but must be controlled by sample.

Case 2 — Fake result. A student asked YZ to summarize 15 documents and ask for "the conclusion from these documents." AI wrote a cause-effect relationship that did not exist in the documents as "result" in a fluent language. When the student went through the documents, he saw that such a conclusion was not supported in any document. Lesson: don't have the AI ​​produce the interpretation and conclusion; Get a case scan, you decide the results.

Case 3 — Modified quote. An author would put a “quote” from the AI ​​abstract into his paper. When he compared it to the document, he noticed that the AI ​​had simplified the sentence and changed a few words. The quote was no longer exactly the same as what was written in the document. Lesson: each quote is taken verbatim from the document, not the abstract.

Four copyable templates

1) Source-dependent scanning:

Below are parts of the document, each labeled with its reference. My research question: [question]. For each document, simply mark "relevant/irrelevant/uncertain" and, if relevant, indicate which sentence/passage it is relevant for and the document reference. Adding comments, drawing conclusions. Documentation: [here]

2) Faithful summary:

Summarize the document below. Rules: (1) Write only what is in the text, DO NOT ADD inferences/comments/conclusions. (2) Link each important point to its place (paragraph/page) in the document. (3) If you need to quote, give it word for word, do not change it. (4) Write "unspecified" information that is not in the text. Document: [here]

3) Passage finder:

Find ALL mentions of this topic in the text below: [topic]. For each, give a brief quote (verbatim) and its location. Adding places that don't include the topic. Mark the match you are unsure of as [unclear]. Text: [here]

4) Comparative case review:

Below are two/three documents. List WHAT each document says on this subject [topic], separately, with its source. PRODUCING a "conclusion" or "overall evaluation" between documents; Just give each document its own statement as fact. Documentation: [here]

Weak prompt / Strong prompt

Weak:

Read these documents and summarize what they say to me about the topic, the results and their importance.

Wanting “consequences” and “significance” pushes AI to deviate from the fact and invent interpretations and connections; It does not connect to the source either.

Strong:

List the facts about [topic] in the following documents, linking each to the document reference. ADD a comment, conclusion, or importance rating. Give quotes word for word. If the topic is not mentioned in a document, write "not mentioned". Documentation: [here]

The difference: fidelity to fact, linking to source, ban on interpretation, and verbatim quotation transform the summary into a reliable map of discovery.

Common mistakes

  • Mistaking the summary for evidence. Summary leads; The evidence is taken from the original document.
  • Telling AI to "draw conclusions". Interpretation and cause and effect are the historian's job; AI scans cases.
  • Taking the quote from the abstract. AI simplifies/replaces quotes; Each quote is taken verbatim from the document.
  • Not supervising elimination at all. If those marked "irrelevant" are not checked against the sample, important documents will be missed.
  • Not asking for gaps. It is necessary to question what AI omits and what it biases.

In summary

Source scanning and summarization is one of the areas where AI saves the most time in history research: it sifts through hundreds of documents, summarizes them and marks relevant passages. But the summary is a map of discovery, not evidence. Link each key point back to the original document, take quotes verbatim from the document, do not have the AI ​​produce conclusions and interpretations, control elimination by sampling, and question what the AI ​​has left out. AI takes you to the right page; It is your job to read that page and make sense of it.

Application task

Collect at least 10 pieces of documentation related to your topic and attach references to each one. Eliminate those related to the "source based scan" pattern; then check some of those marked "irrelevant" yourself as a sample. Have a “faithful summary” summarize one of the relevant outputs and find and verify each key point in the summary in the document. Note how many points the AI ​​adds or distorts.

checklist

  • [ ] I started with a clear research question.
  • [ ] I have attached a source reference to each piece of documentation.
  • [ ] I linked each point of the summary back to the original document.
  • [ ] I took the quotes verbatim from the document.
  • [ ] I audited the elimination by sampling and questioned the gaps.