Gains:
- Ability to plan the scanning scope with artificial intelligence support by generating keywords, synonyms and Boolean search strings from the research question
- Positioning artificial intelligence as a strategy generation tool, not a source finding tool, and being able to verify every source found in real databases (Google Scholar, Scopus, Web of Science, PubMed).
- Ability to understand the risk of artificial intelligence providing fabricated sources and fake statistics and acquire the reflex to verify every citation when determining a gap in the literature.
Every research begins with a literature review: you cannot make an original contribution without seeing what has been done before on your topic, what questions have been answered, and where there are gaps. But the literature is huge; Millions of articles are published annually. This is where AI is a great accelerator — but with a very distinct line of demarcation. The most important sentence of this unit is this: AI is used in literature review as a search strategist, not as a source database. Because AI does not find you real articles; It produces possible-looking article bylines, some of which are fabricated.
In this unit, we will learn how to derive a systematic search plan from your research question, use AI to enrich that plan, and validate each source you find against real databases.
Skeleton of the literature review
A good scan has three stages. First you break down the concepts: you break down your research question into its main components. For example, in the question "The effect of feedback on student motivation in distance education", there are three concepts: distance education, feedback, motivation. Then you create a term pool for each concept: synonyms, field jargon, English equivalents (distance education → distance education, online learning, e-learning, distance learning). Finally, you combine these terms with Boolean operators.
Boolean operators are logical conjunctions that tell search engines how to combine terms. There are three basic operators: AND (results in occurrence of both terms — narrows the scope), OR (any of the terms — expands the scope, collects synonyms), and NOT (excludes a term). Also quotes "..." look for complete phrases, asterisk * is the root contraction (motivat* → motivation, motivational). A search string usually looks like this:
("distance education" OR "online learning" OR e-learning)AND (feedback OR "formative assessment")AND (motivation OR engagement OR "self-efficacy")
AI is excellent at generating these strings, completing synonyms, and adapting them to the syntax of different databases (Scopus, PubMed, Web of Science require different spellings). But the AI isn't where you run the string; real databases: Google Scholar (large, free), Scopus and Web of Science (curated, citation-indexed), PubMed (medical/life sciences), field databases (ERIC education, PsycINFO psychology, IEEE Xplore engineering).
Attention: If you ask an AI chat tool "list 20 articles on this topic", it will return titles that look real, but some of them are made up. This is not screening but danger. Use AI to generate string, find the article in database.
Scope, inclusion and gap analysis
Knows good scanning limits. Inclusion/exclusion criteria determine the scope of your search: which year range, which languages, which study types (experimental, review), which context. AI can suggest these criteria in draft form, and you can sharpen them according to your field. If you are doing a systematic review, you can remind the PRISMA (reporting standard for systematic reviews) flow to the AI and have it produce a crawl log template.
A literature gap is a question or context that has not yet been adequately explored; Your original contribution often fills a gap. AI helps you name possible gaps based on the summaries you read — but take it as an assertion to be verified, not a hypothesis. AI may say "X has not been studied in this area"; However, it may have been studied, it is just that the AI does not know. You only confirm the gap with your actual scan.
Snowball method and gray literature
Keyword search alone is not enough; A good scan deepens with snowballing. Snowballing backwards, the key you find is to scan the bibliography of an article and find the ones before it; Snowballing forward is tracking who cited that article (“Cited by” link in Google Scholar) and arriving at subsequent studies. This method allows you to find studies that are central to the field but that you cannot capture with your keywords. When you provide a bibliography of an article, AI can help you prioritize which references are closest to your question; But again, it's up to you to find and read the article.
Additionally, academic databases do not cover everything. Gray literature — dissertations, institutional reports, conference proceedings, preprints — is critical in some fields and offsets publication bias (the tendency to publish only positive results). AI can suggest which gray literature sources (ProQuest dissertations, institutional repositories, preprint servers such as arXiv/SSRN) you should search; It is often imperative in systematic reviews to not limit your scope to journal articles only.
three mini cases
Case 1 — String expansion. When an education researcher searched for "flipped classroom" on Scopus, she got 300 results, but felt she was missing synonyms. He had AI produce a term pool; He added equivalents such as "inverted classroom" and "flipped learning". With the new string the result increased from 300 to 470; 40 relevant articles were missed in the initial search. The AI did not find the articles, but constructed the string that made them findable.
Case 2 — Return from made-up list. One graduate student asked AI directly for the "top 15 sources" for scanning and pasted the list into his thesis. Under the consultant's control, 4 citations could not be found in Web of Science, and 2 DOIs were incorrect. The student started over: this time he took the line from the AI, found 60 real results in Scopus, sifted through the title-abstract and selected 18. These 18 were all real.
Case 3 — Scope clarification. A health researcher realized that his search was so broad that he was overwhelmed by 2,000 results. Had draft AI inclusion/exclusion criteria produced: last 10 years, English, adult sample, randomized controlled trials. He added the "hospital environment only" filter with his own information. The result fell from 2,000 to 180; A scannable cluster was formed.
Copiable templates
1) Concept separation and term pool:
Your role: an information specialist (research librarian). My research question is:"[write the question]". Break this question down into its main concepts. Generate a term pool of English synonyms, field jargon, and alternative spellings for each concept. Give it in a table. The article is FAKE.
2) Boolean search string generation:
Construct a Boolean searchstring for Scopus using the following term pool: OR terms within concepts, AND between concepts, quote phrases, use * where appropriate. Adapt the same string to Web of Science and PubMed syntax. Pool: [paste]
3) Draft inclusion/exclusion criteria:
I'm doing a literature review on "[topic]". Suggest me a draft of inclusion/exclusion criteria: under the headings of year range, language, study type, sample, context. These are suggestions; I will change them according to my field. Write the reasons briefly.
4) Querying vacancy candidates (to be verified):
Below are summaries of the 10 articles I scanned. Draw out their common themes and suggest gap CANDIDATES as to WHICH QUESTIONS ARE LESS ADDRESSED. Present these as hypotheses that I will verify, not absolute truth. Summaries: [paste]
Weak prompt / Strong prompt
Weak prompt:
Give me a list of resources on artificial intelligence and education.
AI creates fake IDs; The subject is also too broad.
Powerful prompt:
Your role: information specialist. Subject: "The effect of gamification on mathematics anxiety in primary school". DON'T GIVE ME A LIST OF SOURCES. Instead:(1) parse the 3 concepts, (2) generate a repository of English synonyms for each concept, (3) Boolean string for Scopus, (4) draft applicable inclusion/exclusion criteria. The article title is fake.
Purpose
use AI
verification location
Term pool expansion
Yes, strong
Review with domain knowledge
Boolean string construction
Yes, strong
Test in database
Scope/criteria outline
Yes
Sharpen by area
Find an article
no
Scholar/Scopus/WoS/PubMed
Clearance commit
partially
Real scanning is a must
Common mistakes
- Mistaking AI for a database. Requesting an article citation produces spurious citations; This is the most common and dangerous mistake.
- Skipping synonyms. A single term search misses much of the literature.
- Leaving the scope unclear. Without year, language, and genre filters, you'll be overwhelmed with thousands of results.
- Asserting emptiness without verifying it. Saying "no one studied" without actual screening is risky.
- Relying on a single database. Different indexes cover differently; use several together.
Tip: To make your search reproducible, record the string you used, database, date, and number of results in a table. This journal is mandatory in systematic review; It is a good habit in every research.
In summary
Literature review is a systematic task: concept distillation, term pool, Boolean string, coverage criterion. AI is a powerful assistant in these strategic steps; produces synonyms, constructs strings, outlines criteria. But you find articles from real databases, not from AI, and verify each byline. Short rule of thumb: strategy from AI, source from database, validation from you.
Application task
Get your own research question. Parse concepts with AI and produce a term pool and a Boolean string. Run this string against a real database (Scholar or an index you have access to). Note the top 10 results and compare them with the bylines (if any) that the AI has previously suggested when asked directly; Mark how many actually exist.
checklist
- [ ] Have I broken down my research question into concepts?
- [ ] Have I created a synonym/English term pool for each concept?
- [ ] Have I tested the boolean string on the real database?
- [ ] Have I clarified my inclusion/exclusion criteria?
- [ ] Have I verified that every resource I found actually exists?