Unit 2 / 11

Artificial Intelligence in Cataloging and Metadata Generation

Gains:

  • Ability to shorten cataloging time by producing draft metadata in accordance with standards such as MARC and Dublin Core from real imprint information
  • Ability to verify fields with high risk of fabrication, such as publication year, ISBN, author name, with the original source and map subject terms from the controlled vocabulary
  • Ability to make author and subject formats single-approved with authority control and maintain catalog consistency

Cataloging is the invisible but most fundamental job of librarianship. The discovery of a resource (book, article, map, sound recording, digital file) is possible with the correct metadata. Metadata means “data about data”: structured information that describes a resource's title, author, year of publication, subject, language, and physical attributes. If the metadata is good, the user finds the resource; If it is bad or incomplete, the resource is lost even if it exists. In this unit you will learn how to use artificial intelligence in draft metadata generation, but how to ensure standards compliance and accuracy.

Let's first define the two concepts. MARC (Machine-Readable Cataloging) is a record format that libraries have used for decades that places each piece of information in numbered fields; For example, field 245 holds the title and field 100 holds the main author. Dublin Core is a simplified metadata standard for digital resources, consisting of 15 basic fields (title, creator, subject, date, genre, etc.). These standards ensure that records from different libraries can talk to each other.

Step by step: generating draft metadata with artificial intelligence

1. Collect the source's identity information. The cover of the book, the title page, the contents or the text of a digital file. AI only works from the information you give it; If you provide incomplete information, it produces incomplete or fabricated metadata.

2. Specify the target standard. Give the AI ​​a clear target, such as “by Dublin Core fields” or “by MARC 245, 100, 260, 650 fields.” If you do not specify a standard, it produces an arbitrary format.

3. Produce the draft. AI drafts the title, responsibility statement, publication information, and topic proposal from the information you provide.

4. Exercise authority control. An authoritative record is one that establishes the single, standard form of an author's name, institution, or subject (e.g., "Atatürk, Mustafa Kemal, 1881-1938"). AI can write a name in different ways; You check and correct the approved authority format.

5. Verify and confirm. Compare each field to the original source. In particular, precise information such as publication year, ISBN and author name are very susceptible to being forged by AI.

Hint: Give the AI ​​a photo or text of the source's actual credit page; Don't say "guess the name of the book". Predictive metadata undermines the reliability of the catalog. AI writes confidently when making predictions, and this will mislead you.

Controlled vocabulary and consistency

Consistency is everything in cataloguing. If the same topic is written with two different terms in two different records, the user will find one and miss the other. That's why libraries use a controlled vocabulary — an approved, standard list of terms. When suggesting terms from free text, AI can choose its colloquial equivalent rather than the official term in this list. For example, the AI ​​might write “heart attack”; whereas the approved term may be "myocardial infarction".

The solution is to give the AI ​​the controlled vocabulary and say "just select from this list". So instead of making up terms freely, the AI ​​matches the most appropriate one from the approved list.

Caution: If a subject term suggested by AI is not in the controlled vocabulary, do not add that term to the catalogue. The decision to add new terms is up to the expert responsible for authority management; Adding arbitrary terms disrupts the consistency of the catalogue.

three mini cases

Case 1 — Donation collection accelerated. A public library received a donation of 1,200 books. The cataloging specialist had the AI ​​read the imprint page of each book and produce a draft of the Dublin Core; Time per recording decreased from 9 minutes to 3.5 minutes. The expert devoted the remaining time to authority checking and verification; The total time was reduced by approximately 55%, but no records entered the system without being verified.

Case 2 — Fake ISBN caught. An expert had AI produce drafts of 30 books. During the check, he found that the ISBN in 4 records was fake: AI had written a 13-digit number that seemed reasonable, even though he did not know the real number of the book. The verification step prevented four incorrect ISBNs from becoming persistent in the catalogue.

Case 3 — Authority inconsistency corrected. A library found an author as "J. Smith" in some records, "John Smith" in others, "Smith, John" in others. AI was matched to the authority file, bringing all records into a single approved format ("Smith, John, 1965-"); If this job was done manually, it would take days, but with the AI ​​suggestion, it was reduced to a few hours, but each match was approved by the expert.

Retrospective conversion and old records

Many libraries have incomplete or outdated records that were entered years ago with different rules. Moving these to the current standard is called retrospective conversion and is very laborious when done manually. The AI ​​can read an old record and generate a draft to move it to the current field structure: for example, it can parse a free-text citation and distribute it to the correct fields. However, there are two dangers here. First, AI can make up missing information to “make up”; second, it may perpetuate an error in the old record by copying it to a new format rather than correcting it. Therefore, in retrospective transformation, the output of AI should be checked systematically, not by sampling, but by critical fields (author, year, identifying numbers). A good practice is to display pre- and post-conversion records side by side, making each change visible; so the expert can see at a glance what has been moved, what has changed and what may have been fabricated.

Attention: The draft metadata produced by the AI ​​should not be transferred to the system as a batch without reviewing it one by one. A mass error corrupts hundreds of records at once, and retrieving is much more troublesome than producing. It is safest to produce and verify in small batches.

Four copyable templates

1) Dublin Core outline:

Your role: assistant cataloguer. Fill in the Dublin Core fields from the following byline text: Title, Creator, Subject, Description, Publisher, Date, Genre, Format, Language. Use only information that is clear in the text; Leave the missing field "[no information]", do not guess. Imprint text: [here]

2) MARC field outline:

Propose outlines for the following MARC fields from the book information below:245 (title/credit), 100 (main author), 260/264 (publication),300 (physical description), 650 (subject). Mark any field you are unsure of as "[needs verification]". Information: [here]

3) Topic matching from controlled vocabulary:

Below is a summary of the source and a list of approved topic terms. Match only the 3-5 most suitable terms from the list to the source. If there is no suitable term in the list, write "no suitable term in the list", do not make up new terms. Summary: [here] / Term list: [here]

4) Authority form control:

Match the author names below with the approved forms in the authority list I provided. For each name: format in record -> approved format. If there is no equivalent in the list, write "no authority record". Names: [here] / Authority list: [here]

Weak prompt / Strong prompt

Weak prompt:

Extract the cataloging information for the following book: "Artificial Intelligence and Society".

Only the title is given; The AI ​​is forced to make up the author, year, publisher, and ISBN. There is no target standard. Result: a convincing but probably false record.

Powerful prompt:

Your role: assistant cataloguer. Below is the full text of the book's imprint page. Fill in the Dublin Core fields from this. Only use information that is clearly stated in the text; Leave the non-text field blank and write "[no information]". No dates, names or ISBNs are made up. Imprint text: [full text here]

The powerful prompt limits the AI to real byline information and forces it to honestly leave blanks instead of making them up.

Metadata field and validation table

area

Risk of fabrication

How to verify

Title

low

Compare with imprint page

Author/authority format

medium

Search in authority file

Year of publication

high

Confirm with imprint/colophone

ISBN

very high

One-to-one control with barcode/source

Subject term

high

Match in controlled vocabulary

Common mistakes

  • Just requesting metadata from the title. Incomplete information drives AI to make things up.
  • Not specifying the target standard. Arbitrary formatting disrupts the harmony of the records.
  • Accepting the ISBN and year without verifying it. These areas carry the highest risk of fabrication.
  • Taking the subject term from outside the controlled vocabulary. Consistency and availability are compromised.
  • Bypassing authority control. The same author disperses in different formats, the user misses the source.

In summary

In metadata generation, AI is a powerful assistant that quickly extracts imprint information and significantly shortens cataloging time. But the price of productivity is stringent validation against the risks of fabrication: make the target standard (MARC, Dublin Core) clear, run the AI ​​only with real byword knowledge, map subject terms from controlled vocabularies, and validate authority forms. Instead of making up the gaps, the AI ​​should honestly leave it blank; You keep the final approval.

Application task

Give the text of the imprint page of a real book you have to the AI ​​with the "Dublin Core draft" template and get a draft. Then have the same book produced with just the title ("Weak prompt") and compare the two outputs: which fields were fabricated? Then, with the "Topic mapping from controlled vocabulary" template, give a short list of approved terms to get the topic matched and check if the AI ​​goes off the list.

checklist

  • [ ] I gave the real identity information of the source to the AI, I did not leave it to guessing.
  • [ ] I have clearly specified the target metadata standard (MARC/Dublin Core).
  • [ ] I have verified the publication year and ISBN verbatim with the original source.
  • [ ] I mapped the topic terms from the controlled vocabulary.
  • [ ] I have confirmed the forms of authority with the approved record.