Unit 8 / 10

Literature Review and Knowledge Synthesis: RAG, Systematic Compilation and Hypothesis Generation

Gains:

  • Ability to understand that plain chat tools produce fake references and apply the discipline of verifying each source in the actual database.
  • Ability to reduce hallucination by connecting literature synthesis to real documents with the RAG (resource dependent generation) approach
  • Ability to interpret hypotheses as a beginning rather than evidence by using artificial intelligence for preliminary screening in systematic review and keeping the final decision and transparency in the hands of humans

Knowledge explosion in bioengineering is a real problem: hundreds of thousands of articles are published every year, accumulating thousands of studies on a single topic. Before establishing a hypothesis, choosing a method, or interpreting a result, an engineer needs to scan relevant literature; but this takes weeks manually. AI is one of the areas where it saves the most time in literature review, summarization and information synthesis. But it is also the most dangerous source of hallucinations: AI can produce non-existent articles, fabricated DOIs, and spurious findings. The main lesson of this unit is the discipline of connecting every claim to the original source when using AI in literature.

In this unit, you will learn how to use AI in safe literature review, why the RAG (Retrieval-Augmented Generation — AI finds and produces the answer from the real document database, not from its memory) approach is more reliable, systematic review support and hypothesis generation.

Why is the plain chat tool dangerous?

When you tell a language model to “find articles about This means fluid but made-up references: author names, journals, years, and DOIs as if they were real. Therefore, never violate two rules in literature review: (1) verify each reference in the actual database (PubMed, Google Scholar, journal page), (2) if possible, use RAG-based tools linked to real articles (such as Elicit, Consensus, Scite).

Caution: A testimonial from an AI may appear genuine — in the correct format, familiar journal, reasonable year. This doesn't mean it exists. Consider an article as existing only when you find and open it in PubMed/DOI. Including a fabricated reference in your publication is a mistake that can destroy your scientific reputation.

RAG: resource-dependent production

In the RAG approach, the AI ​​first finds relevant passages from a real document repository (either PDFs you upload or a scientific database), then generates the answer based on those passages and points to the source. This greatly reduces hallucination because the model has to "quote" rather than "imagine." Giving your own PDF library to a RAG tool and asking “what do these 40 articles say about so and so” is much more reliable than straight conversation; however, see the passage of each quote in the actual article.

Tip: Even when using RAG, explicitly instruct "cite source and answer only from the documentation I provide, do not add non-document information". Then find the passage the model points to in the actual document. These two steps make your synthesis defensible.

Systematic review and hypothesis generation

A systematic review—answering a specific question by scanning all relevant literature in a transparent and reproducible method—requires high methodological rigor: search strategy, inclusion/exclusion criteria, data extraction. AI speeds up the mechanical parts of this process (abstract screening, initial screening, data extraction draft), but the responsibility and transparency of each decision (whether to include or exclude a paper) rests with the researcher. AI also inspires hypothesis generation by pointing out gaps in the literature; but the hypothesis it produces is a starting point, not evidence.

three mini cases

Case 1 — Scanning speeded up 10x. A review team had to sift through 1,850 article abstracts. With a RAG-based tool, the inclusion/exclusion pre-qualification went from 3 days to 5 hours. But two researchers independently checked 120 articles that AI had deemed "excluded" and recovered 9 false eliminations; AI made the selection, the human made the final decision.

Case 2 — Fabricated reference caught. A PhD student asked YZ for 15 references for his thesis entry. When he checked on PubMed, he found that 6 of them did not exist at all and 3 were attributed to the wrong author. The plain chat tool produced a realistic but fake bibliography. Verification averted an academic disaster.

Case 3 — Hypothesis inspiration. A metabolic engineering team handed the AI ​​a corpus of 60 papers and asked "which pathways are understudied?" The AI ​​pointed to an untested enzyme combination. The team took this as a hypothesis, validated it in the literature, and designed an experiment; Ultimately, they came up with a real finding that increased productivity. Inspiration came from AI, proof came from experiment.

Four copyable templates

1) Source-related literature synthesis (RAG):

Your role: scientific review assistant. Answer ONLY from the documents I uploaded; Do not add any extra-document information. For each claim, give the source document name and relevant passage. To a question that does not have an answer in the documents, say "it is not in the documents". Reference, DOI, or finding FITTING.Question: [question]

2) Reference verification discipline:

I will give you a topic. Suggest possible keywords and search strategy (with PubMed syntax). But YOU article list is FAKE; I will make the call myself. Tell me what terms, filters and inclusion/exclusion criteria I should use instead.

3) Systematic compilation elimination aid:

I will give you article title+summaries. Suggest include/exclude/unclear for each based on the following inclusion criteria(include: [criteria]) and write your RATIONALE. I will manually review the "uncertain" ones. Note that the decision is preliminary, not final; Do not eliminate any article without justification.

4) Hypothesis generation (carefully):

I will give you a summary of findings on a topic. Suggest me 3 testable hypotheses based on GAPS in the literature and untested combinations. For each hypothesis: state the basis, how it will be tested, and that it is a START, not evidence. Do not rely on non-existent evidence.

Weak prompt / Strong prompt

Weak prompt:

Find and summarize 10 articles for me about enzyme engineering with CRISPR.

The model does not perform actual searches; will most likely generate fake references.

Powerful prompt:

Your role: build assistant. Based ONLY the 12 PDFs I uploaded, I synthesize the main findings on the subject of "CRISPR-based enzyme engineering". For each finding, show which PDF it came from and the full passage. Do not add any studies, DOIs, or results that are not in the PDFs. Mark any subtopics that are not covered as "not covered in these documents."

Difference: linking to actual documents, source+passage requirement, prohibition on fabrication, and honesty of scope.

Literature tasks and reliable approach

Quest

risky road

safe way

verification

Find an article

Saying "find" to the chat

PubMed + RAG tool

Confirmation in DOI

summarizing

Summary without documentation

Summary from loaded PDF

See the actual passage

elimination

Blind trust in AI

AI pre-qualification + human

Second reviewer

synthesis

free production

Source dependent synthesis

Quote check

hypothesis

mistaking it for "evidence"

starting point

experimental test

Common mistakes

  • Not verifying references. AI produces realistic but fake resources; each should appear in the database.
  • Mistaking plain chat for a call. The model generates from memory; actual search requires RAG/database.
  • Delegating the elimination decision completely. Incorrect elimination will break the compilation; Second reviewer required.
  • Mistaking a hypothesis for evidence. The hypothesis generated by AI is speculation until tested by experiment.
  • Exaggerating the scope. AI may give the impression that “I scanned everything”; You document the actual scope.

In summary

Literature review is the area where AI saves the most time but produces the most hallucinations. Plain chat tools can produce realistic but fake testimonials; so verify each source against the actual database and, if possible, use RAG-based tools that depend on actual documents. In a systematic review, AI performs preliminary screening, but the final decision and transparency belongs to the researcher. AI inspires hypothesis generation; That hypothesis is a starting point, not proof, until it is confirmed by experiment.

Application task

Pick a topic and tell the AI ​​(via a plain chat tool) “find 8 articles on this topic and summarize.” Verify each reference he gives against PubMed or DOI, one by one, and count how many are genuine. Then repeat the same task with a RAG tool or by uploading a few PDFs you have downloaded yourself and with the instruction "answer from these documents only". Write a brief comparison of the reliability of the two approaches.

checklist

  • [ ] I verified every reference the AI gave against the actual database.
  • [ ] I used a source-dependent (RAG/loaded document) approach in literature synthesis.
  • [ ] I audited the screening decisions with a second reviewer.
  • [ ] I considered the hypothesis produced by AI not as evidence, but as a beginning to be tested.
  • [ ] I saw the passage of each quote in the original article.
  • [ ] I have honestly documented the true extent and limitations of the scan.