Gains:
- Ability to implement never taking the citation count and citations from the AI memory while clustering and summarizing the data taken from the real citation database with AI
- Ability to understand that metrics such as citation, h-index, and impact factor indicate use but not quality, and their incomparability across fields.
- Ability to consider the ethical limits of using bibliometric criteria alone in evaluating individuals and balance them with qualitative evaluation
A university library measures the research output of its institution; a researcher to find the most influential studies of a field; a manager may want to evaluate the scientific visibility of a department. Behind this work is bibliometrics: the field that numerically examines scientific publications and the citation relationships between them. In this unit, you will learn how to use artificial intelligence in citation analysis, publication mapping and impact assessment; but you will learn how to navigate the limits and ethical pitfalls of numerical metrics.
Let's define the concepts. Citation is a scientific reference made by one publication to another publication; The number of citations is considered an indicator of how much a work has been used. The h-index is a measure that combines both the productivity and impact of a researcher (if h number of publications have each received at least h citations, the h-index is h). Impact factor is a measure that shows the average citation level of a journal. These metrics are useful, but none are direct measures of “quality”; Keeping this distinction in mind is the essence of bibliometrics.
Step by step: responsible bibliometric analysis
1. Clarify the question. "What are the basic studies of this field?", "Which are the most cited publications of our institution?", "How do these two fields intersect?" Each question requires a different analysis.
2. Use reliable data. Citation analysis is only as reliable as the data source. AI can make up citation counts or publication lists; Therefore, it is essential to work with data from a real citation database. Use AI to interpret and summarize real data, not to generate data.
3. Have the patterns marked. AI helps summarize topic clusters, collaborative networks, or trends over time in a large set of publications. This directs the human analyst's attention to important points.
4. Interpret the criterion in context. A low citation count does not indicate that the work is worthless; The field may be small, the topic may be new, or the publication may be very new. Never interpret the number without context.
5. Observe the ethical boundary. Bibliometric measures should not be used alone to make decisions about hiring, promoting, or funding individuals. Numbers do not measure the value of people; Abusing them is both unfair and harmful.
Tip: When having AI perform a citation analysis, provide the citation counts and publication citations from the real database; Let the AI simply organize, group and interpret this data. "What is this writer's h-index?" Asking directly may result in a made-up number.
Limits of criteria and manipulation
Because bibliometric metrics appear powerful, they are open to abuse. Citation numbers can be inflated in a variety of ways: excessive self-citation (citing one's own work), citation rings (groups of people reciprocally citing each other), or the field's own citation culture. Moreover, the criteria are not comparable across fields: the citation environment of a medical article and a philosophy article is completely different.
This problem is even more evident in the humanities and social sciences; because significant outputs in these fields (books, translations, reviews) are underrepresented in citation databases. To measure the impact of a historian or literary scholar solely by the number of citations is to misunderstand the nature of the work. The librarian is a guide that explains the limits of these criteria to administrators and researchers.
Caution: Do not present a bibliometric measure as a final judgment about an individual's academic merit or the quality of a department. Principles of responsible research evaluation emphasize that measures should complement qualitative evaluation, not replace it.
three mini cases
Case 1 — Area mapped. A librarian wanted to show a research group basic work in their field. It pulled data from 200 publications from a real citation database; YZ summarized this data into topic clusters and highlighted the 10 most cited studies. The analyst verified each cluster and identifier in the database; the group quickly grasped the structure of the area.
Case 2 — Fake h-index caught. A manager asked the AI about a researcher's h-index; The AI said "42". The librarian looked at the actual database: the correct value was 19; The AI had made up the number. This showed that metrics should never be taken from AI memory but from real data.
Case 3 — Incorrect comparison prevented. One unit wanted to allocate funds by comparing two chapters based on their average number of citations. The librarian explained that one of the departments was medicine (high citation culture) and the other was humanities (book-heavy, low citation culture), and that this comparison would be unfair. The evaluation was supported by qualitative criteria that took into account field differences.
Visual maps and reading traps
A frequently used tool in bibliometrics is network maps (e.g. co-citation or co-authorship networks), which show relationships between publications or authors by nodes and links. These maps make visible the structure, clusters and bridgeworks of an area; AI can aggregate the underlying real data and help interpret these maps. But visual maps, while powerful, can also be misleading: a large, centrally visible node may not be the "most important" but just the "most connected"; The size of a cluster can only reflect the volume of publication, not the value of that space. Additionally, how the map is drawn (which threshold, which time period, which database) completely changes the outcome. Therefore, when presenting a bibliometric visualization as a basis for a decision, it is imperative to clearly state which data it is based on, which parameters and over what time period. If an AI-generated comment qualifies a node on the map as a “leader” or “most influential,” check to see if that characterization is supported by real data.
Tip: Don't fall into the "center = top" trap when interpreting a network map. Centrality only indicates connection density; The original contribution or quality of a work is not visible on the map and can only be understood by reading the content.
Four copyable templates
1) Summarizing actual citation data:
Below is the publication list and citation numbers I took from a real citation database. Divide them into topic clusters and summarize each cluster in a single sentence. Just use the data I provided, do not add/make up new publications or citation counts. Data: [here]
2) Trend and pattern marking:
Highlight notable trends (increase, decrease, collaboration concentration) in the year-based publication/citation data below. Base every observation on data, do not speculate. Data: [here]
3) Criteria interpretation warning:
Interpret the following bibliometric criterion, also stating its limits: on which subject does it provide information, on which subject does it not provide information, on which subject is it open to misuse? Criteria and context: [here]
4) Area difference control:
Can the citation data of the following two groups be directly compared? Evaluate the differences in citation culture between the fields and explain whether the comparison is fair. Data: [here]
Weak prompt / Strong prompt
Weak prompt:
List this author's most influential articles and citation counts.
The AI generates made-up titles and citation counts from its global memory without accessing the real data. Result: a convincing but spurious analysis.
Powerful prompt:
Your role: bibliometrics assistant. Below are the publication citations and citation numbers I pulled from the real citation database. Sort this data from most cited domain to lowest, divide it into topic clusters, and summarize each cluster. Do not deviate from the data I have given, do not make up any numbers or labels. Add a short note about the limits of the criteria. Data: [here]
Powerful prompt; It limits AI to real data, prohibits fabrication, and reminds us of benchmark limits.
Criteria and limits table
criterion
What shows
What does not show
Attention
Number of citations
Usage/visibility
quality, accuracy
Depends on area culture
h-index
Productivity+impact
value of the individual
Sensitive to career length
Impact factor
Journal average
Value of single article
Journal level criteria
collaboration network
Who worked with
Contribution rate
Role in the network unclear
Common mistakes
- Requesting the citation count from AI. Numbers are made up outside the real database.
- Thinking that the criterion is a measure of quality. Citation indicates usage, not quality.
- Comparing areas directly. Attribution cultures are different, it is not fair.
- Measuring the humanities by attribution. Book-heavy outputs are underrepresented in databases.
- Evaluating individuals with criteria. The hiring/promotion decision should not be based on a single criterion.
In summary
In bibliometrics and citation analysis, AI is a powerful assistant that clusters real data, flags trends and summarizes; But never take citation numbers and bylines from the AI memory, because it makes it up. Metrics show usage and visibility, not quality; cannot be compared across domains; is underrepresented in the humanities and cannot be used alone to evaluate individuals. The librarian is a responsible guide who explains the limits of these criteria and balances them with qualitative evaluation.
Application task
Take a small list of publications (10-15 records, with citation counts) from a real citation database. Have the AI cluster and summarize with the “Summarize real citation data” template, then validate the citations and counts against the database. Then print out the limits of a metric of your choice (e.g. h-index) with the "Criteria interpretation warning" template and plan how you will explain it to a manager.
checklist
- [ ] I provided the citation numbers and citations from the real database, I did not have them made up by AI.
- [ ] I presented the metrics as indicators of usage/visibility, not quality.
- [ ] I took into account the difference in citation culture between fields.
- [ ] I took into account the underrepresentation of the humanities.
- [ ] I did not use the criterion alone in evaluating individuals.