Gains:
- Ability to draft corpus measurements such as word frequency, collocation, vocabulary richness and keywords with artificial intelligence
- Knowing that artificial intelligence is unreliable in exact counting and being able to verify each count with a separate tool
- Ability to understand that a number does not say anything on its own and that the linguistic interpretation that establishes the meaning belongs to humans.
Literature and language studies are not just "commentary"; It also has a measurable and countable aspect. How many different words an author uses, which words he likes to use with which words, which words become common in a period - these are investigated through linguistics (the science that studies the sound, structure, meaning and use of language) and especially corpus linguistics. A corpus is a collection of texts, such as the complete works of an author, the annual archive of a newspaper, or the poems of a period. Corpus analysis looks for numerical patterns in this chunk.
In this unit, we will learn to use AI as a corpus analysis assistant: quickly extracting metrics such as word frequency (the number of times a word occurs in the text), collocation (word pairs that occur frequently together, e.g. "black" + "my fortune"), vocabulary richness, and seasonal variation. But a warning from the beginning: AI can make mistakes when counting alone; Special tools are needed for large and precise counts, and you are the linguist who establishes the meaning of each number.
Where is AI strong and where is it weak in corpus analysis?
Its strengths are: Recognizing patterns in small and medium texts, generating hypotheses about the vocabulary of a text, suggesting candidates for collocations, interpreting a result, describing a methodology. AI asks “what phrases stand out in this author?” It gives a quick start to the question.
Weakness: Accurate counting. AI is a language model, not a calculator; can give an approximate and sometimes incorrect answer to the question "how many times has it occurred" in a long text. For precise frequency, exact colocation statistics, or large corpus scanning of thousands of words, specialized software (e.g. corpus tools such as AntConc or a spreadsheet) is used. AI is valuable in interpreting the output of these tools; But don't make him do the raw count blindly.
Tip: If you need an exact number, use two ways: have a tool do the counting in small text and have the AI interpret it; Or have the AI count and verify it in another way. Never use the sentence "AI said it occurs 47 times" without confirmation.
A step-by-step corpus study
- Put a question. “How are color words distributed in this poet?” Counting is meaningless without a clear question like:
- Define the corpus. Which texts? Which edition? Which period? The border must be clear.
- Select the measurement. Frequency, collocation, richness of vocabulary?
- Measure and verify. Make the count, then confirm a second way.
- Comment. The number alone does not say anything; It is your comment that makes the finding that "color words doubled in the second period" meaningful.
A frequently used measure of vocabulary richness is the ratio of the number of different words to the total number of words (type/token ratio). If this ratio is high, the text is word-variety, and if it is low, it means it is repetitive. But this ratio is affected by text length; Be careful when comparing short and long texts directly.
three mini cases
Case 1 — Collocation demonstrated a style. A researcher examined adjective-noun pairings in 3 novels of a novelist. AI suggested that the word “pale” matches 5 different nouns; The researcher verified 4 of the texts and eliminated 1 as fabricated. The confirmed pattern became a thesis statement about the author's descriptive style.
Case 2 — Counting error caught. A student asked AI how many times the word “night” appeared in a 2,000-word text; The AI said "23". When the student threw the text into a spreadsheet and had it counted, the actual number was 17. It has become clear that AI's raw count is unreliable.
Case 3 — Periodic change sparked controversy. One group compared the proportion of abstract words in a poet's early and late poems. The first draft of AI suggested a trend; When the group went back to the text and verified the count, the trend turned out to be real, but the "causes" suggested by the AI were unverifiable speculation and were eliminated.
Four copyable templates
1) Word frequency outline (with verification):
List the 15 most frequently occurring content words (exclude function words such as conjunctions and prepositions) in the text below, with their approximate frequencies. WARNING: Note that counts MAY NOT be accurate and note "to be accurate, they must be verified with a separate counting tool." Text: [paste text here]
2) Colocation candidates:
Suggest pairs of words (collocations) that occur frequently together in the text below. Give at least one quote from the text for each pair. Double fabrication that has no equivalent in the text; If you are not sure, mark it as "needs verification".Text: [paste text]
3) Vocabulary comparison:
I will give you two texts. For each: approximate total words, observations about different word varieties, and prominent word domains (e.g. color, nature, emotion). State that the numbers are approximate. Leave the interpretation of the comparison to my verification. Text A: [...] Text B: [...]
4) Result interpretation:
I got the following results from a counting tool: [paste results].Generate 2-3 hypotheses about what these numbers might mean linguistically. Mark each hypothesis as "requires evidence" and tell me how I can test it. Avoid excessive comments.
Weak prompt / Strong prompt
Weak: "Which word occurs most in this text?"
Strong: "List the most frequent content words in this text with their approximate frequency; exclude function words. Make it clear that the numbers are not exact and need to be verified with a separate tool. Give an example sentence for each word."
The powerful prompt makes the AI accept the count limit and tells you upfront that verification is required. In the weak prompt, the AI gives a single number and you don't know that it could be wrong.
Measurement and reliability table
measurement
What shows
AI alone
right way
word frequency
frequent words
About, risky
Counter + AI commentary
colocation
those who passed together
Produces candidates
Confirmation from the text
Type/token ratio
Verbal diversity
Can be misleading
Account with tool
seasonal change
Trend over time
generates hypotheses
Measure and verify
Concluding comment
Meaning
Powerful but exaggerates
human judgment
Keyword and n-gram concepts
Two concepts make your work easier in corpus analysis. The first is the keyword (keyword in English - the word that occurs in a text or author much more often than expected compared to the general language, giving that text its "character"). An author's keywords reveal his interests and style; For example, if "sea" and "separation" occur more frequently than expected in a poet, this statistical protrusion becomes the starting point for an interpretation. To identify keywords, it is necessary to compare the frequencies of a text with the frequencies of a reference corpus (a general, large body of text used for comparison); This comparison is a definitive tool job, it is a calculation that AI cannot make reliably on its own.
The second is n-gram (a sequence of n consecutive words in a text; a sequence of two words is called "bigram", a sequence of three words is called "trigram"). N-gram analysis reveals an author's formulaic expressions and frequently used phrases. For example, if a writer uses phrases such as "kind of" or "however" extremely frequently, this indicates his habit of making sentences. The AI can suggest these phrases as candidates; but for true frequency and statistical significance it is necessary to return to the counting tool.
Hint: Tell the AI, "List the notable phrases (two or three words) in this text as candidates, and give an example from the text for each." Take the candidate list and measure the actual frequency with a separate tool. Position the AI as the “noticer,” the agent as the “counter,” and yourself as the “interpreter.”
Common mistakes
- Relying on the AI's raw count. AI is not a calculator; Verify the exact number by separate tool.
- Presenting the number without comment. "Passed 47 times" doesn't say anything; You create the meaning.
- Ignoring text length. Lyric diversity rates are length sensitive.
- Making thesis without verifying the collocation. Non-AI couples may suggest; Confirm from the text.
- Extreme comment. Don't accept the "reasons" suggested by the AI without evidence.
In summary
Corpus analysis opens up the countable facet of literature: frequency, collocation, vocabulary, periodic change. AI is powerful in this area at recognizing patterns and interpreting results; but it is unreliable in absolute counting. Verify the number with the separate tool, confirm collocations from the text, and establish the meaning of each number yourself. Number is not evidence, it is the door to evidence.
Application task
Choose a short text (500-1000 words). Identify a content word (e.g. "night" or "road"). First, count yourself in the text. Then have the AI count it. Compare the two results; If there is a difference, investigate the reason. Then write a one-paragraph comment on the function of that word in the text — turning the number into meaning.
checklist
- [ ] I put a clear linguistic question.
- [ ] I defined the boundaries of the corpus (text, edition, period).
- [ ] I verified the AI's count a second way.
- [ ] I confirmed the collocations from the text.
- [ ] I did not leave the number without comment; I created the meaning myself.