Gains:
- Ability to reduce legal risks by evaluating the copyright status of the input and the data conditions of the tool and questioning the rights of the output
- Ability to protect personal data through anonymization and minimization and ensure privacy by being transparent to the user
- Ability to establish a governance framework for the institution that includes approved tools, prohibited data and verification rules, and test its use with librarianship values.
The librarian and information manager is someone who not only organizes and presents information, but also protects it: protecting the rights of the creator, the privacy of the user, and the responsibility of the institution. Artificial intelligence challenges all three fields with new and complex questions. Is uploading a document to AI a copyright violation? Does giving user data to an AI tool violate privacy? How should the organization manage the use of AI? In this closing unit, we will establish a framework that brings together the ethical and legal threads we have touched upon throughout the module.
Let's define the concepts. Copyright is the legal right granted to the creator of a work over the use of his work; Unauthorized duplication or use may be considered infringement. Personal data is any information that directly or indirectly identifies a person and is subject to legal protection. Governance is the framework of rules, policies, and accountability that governs an organization's use of AI. These three concepts form the legal and ethical basis for responsible AI use.
Step by step: responsible and legal use of AI
1. Evaluate the copyright status of the entry. Before loading a text, book, or article into an AI tool, consider the copyright status of that work and the terms of use of the tool. Giving a copyrighted work to a tool that uses the data you upload for its own purposes may constitute a violation of rights.
2. Question the rights of the output. The copyright status of the text produced by the AI may be uncertain, and the AI may unknowingly reproduce copyrighted content in the training data. Before publishing an AI output, evaluate its originality and potential violations.
3. Protect personal data. Users' identity, reading/search history and personal information should not enter any AI tool without explicit consent and strong protection. Anonymization and data minimization (only as much data as necessary) are essential.
4. Ensure transparency and consent. Users should know how their data and interactions are processed; Where AI is involved, this should be clearly stated. The use of covert AI undermines trust.
5. Establish a governance framework. The organization should establish in a written policy which tools are approved, what data can be entered, who is responsible, and how the outputs will be verified.
Tip: Create a simple “AI usage rules” one-sheet for your organization: which tools are approved, what data is never entered (personal/proprietary/confidential), how each output is verified, and who is responsible. A clear one-page guide is more enforceable than a lengthy policy.
Ethical principles and values of librarianship
The deep-rooted values of the librarianship profession serve as a compass in the age of AI: equal access to information, freedom of thought, user privacy, intellectual honesty, and impartiality. AI can either support or threaten these values. It may expand access but reinforce prejudice; may speed up service but may compromise privacy; It may facilitate discovery but may spread misinformation.
The librarian in charge tests each AI use against these values: "Does this use equalize access or exclude a group? Does it protect privacy? Does it provide accurate and honest information?" AI is a tool; The values with which you will use it are the essence of the profession. Technology changes, but these values are the constant compass of librarianship.
Caution: "Everyone is using it" or "it's fast" does not justify the use of an AI. Even if a use is legal, it may be unethical; Do not adopt a practice that is harmful in terms of copyright, privacy and justice, just because it is easy. The ultimate responsibility lies not with the tool, but with the professional using it.
three mini cases
Case 1 — Copyrighted download has been stopped. One team planned to load an entire copyrighted reference book into a general AI tool to generate summaries. The information manager found this to be risky, both in terms of copyright and the tool's data usage terms; instead, it was studied only in brief, authorized excerpts and within the closed system of the institution.
Case 2 — Personal data anonymised. A library wanted to analyze user feedback with AI; but the feedback contained usernames and contact information. The team anonymized the data before analysis; only anonymous, collective tendencies were processed. Privacy preserved, insight regained.
Case 3 — Governance gap closed. In one organization, employees were using different AI tools without supervision, sometimes entering confidential documents. The information administrator published a one-page policy defining approved tools, prohibited data types, and verification rules; Uncertainty and risk have decreased significantly.
Personnel competence and corporate culture
A governance policy is only as strong as the people who implement it. Even the best-written rule remains on paper if employees do not understand the risks of AI. The enduring basis for responsible AI use is therefore staff competence: each employee's understanding of what a hallucination is, what data should never be entered, and why an output should be verified. Libraries and information institutions have a dual role in this regard; It must both train its own staff and spread this literacy to users. A good corporate culture is one that reports an AI bug rather than hiding it, and finds it normal to say "I'm not sure about this output, let's verify." In an environment where mistakes are hidden for fear of punishment, a fabricated source or a breach of privacy spreads unnoticed. Additionally, because technology and tools change rapidly, policy and training should be updated regularly rather than one-time. Governance; It is a living system where rules, tools, education and culture work together. At the center of this system stands the competent person, who takes responsibility and makes the final decision, not the vehicle.
Tip: Before adopting a new AI tool, do a small pilot: try it with a limited team, on real but low-risk tasks, recording errors and risks. This both builds competence and reveals policy vulnerabilities before rolling out the tool throughout the organization.
Four copyable templates
1) Copyright preliminary assessment:
Help me evaluate the following content for copyright before loading it into an AI tool: can the content be copyrighted, which way of use is risky, which alternative (excerpt, permission, public domain) is safer? Content type: [here]
2) Personal data scanning:
Does the following text contain personal data (name, contact, identity, reading/search history, information implying health/belief)? Mark each one, classify it and suggest how to anonymise it. Text: [here]
3) Draft AI usage policy:
Draft one-page AI usage guidelines for our organization: (1) approved tool policies, (2) data types never entered, (3) output validation rule, (4) accountability and transparency, (5) copyright and privacy policies. Keep it short and actionable. Context: [here]
4) Ethical control query:
Test the following planned AI use against librarianship values: equity of access, privacy, accuracy/integrity, risk of bias, and flag concerns if any.Usage: [here]
Weak prompt / Strong prompt
Weak prompt:
Should I put it on AI to summarize this book?
The question does not take into account copyright, data conditions of the tool and alternatives; Answering "yes" may result in a violation of rights.
Powerful prompt:
Your role: assistant copyright and privacy advisor. I'm considering giving the following content to an AI tool. List possible problems in terms of copyright status, vehicle data use and personal data risk; suggest safer alternatives (short quote, permission, public domain, closed system). Claim legal certainty, show risks.Content and context: [here]
Powerful prompt; It considers the decision in terms of copyright, data condition and privacy aspects and puts secure alternatives on the table.
Governance dimensions table
Size
basic question
precaution
copyright
Is there any input/output infringement?
Permission, public domain, short excerpt
privacy
Is personal data protected?
Anonymization, minimization
transparency
Does the user know?
Clear information, consent
Responsibility
Who is accountable?
Written policy, human approval
Values
Are access and integrity protected?
Ethical check, bias check
Common mistakes
- Uploading copyrighted content without thinking. There is a risk of copyright infringement and data condition.
- Entering personal data into the public tool. Invasion of privacy and legal risk.
- Hiding the use of AI. Without transparency, trust is damaged.
- Confusing “legal” with “ethical.” A legal use may be unethical.
- Not establishing a governance framework. Uncontrolled use creates unpredictable risks.
In summary
Copyright, confidentiality and ethics are the basis of responsibility of librarianship in the age of AI. Evaluate the copyright status of the entry and the data conditions of the tool; question the rights of the output; protect personal data with anonymization and minimization; be transparent to the user and obtain consent; Establish a clear governance framework in the organization. The deep-rooted values of librarianship (equal access, privacy, integrity, impartiality) are the compass that tests every use of AI. Technology changes; Responsibility and values do not change, and ultimate accountability always lies with the human being.
Application task
Prepare a one-page usage rules document for your organization (or a fictitious organization) with the “AI usage policy draft” template: approved tools, data never to be entered, validation rule, liability. Then, test a real AI use you are planning for access, privacy, accuracy, and bias with the “Ethical control query” template and write how you will address any concerns that arise.
checklist
- [ ] I evaluated the copyright status of the entry and the data conditions of the tool.
- [ ] I anonymized personal data and used only what was necessary.
- [ ] I have been transparent to the user about the use of AI and respected consent.
- [ ] I have created a written AI usage policy for the organization.
- [ ] I tested the usage with librarian values and kept the responsibility on the human.
Module Exam
1. Which of the following is the most accurate positioning for artificial intelligence in librarianship and information management?
- A) AI is a summary, metadata, search and recommendation assistant; The responsibility for accuracy, source reliability and ethical decisions lies with humans ✔
- B) Artificial intelligence can reliably verify the authenticity and accuracy of a source
- C) Artificial intelligence only works in book cataloging, it has nothing to do with other librarianship work
- D) Since artificial intelligence is more impartial than humans, subject and classification decisions should be left to it.
Description: Artificial intelligence; It is an assistant that drafts metadata, summarizes it, translates it, marks the pattern and points it to the source. The competent expert is responsible for high-consequence tasks such as verification of accuracy, topic and classification decisions, source credibility, and copyright and privacy decisions; An unverified output may lead to the dissemination of a fabricated source or a privacy violation.
2. Why is 'hallucination' particularly dangerous for the librarian?
- A) The output of artificial intelligence is produced very slowly and wastes time
- B) Artificial intelligence can convincingly produce non-existent sources and quotes, and the essence of the profession is the authenticity of the source ✔
- C) Hallucination occurs only in image processing, it has nothing to do with text
- D) Hallucination is only seen in very old artificial intelligence tools, not in current tools.
Explanation: A hallucination is when artificial intelligence produces a non-existent source, quote or information in an extremely convincing way. Since the essence of the profession is the authenticity and reliability of the source, if a fabricated imprint or quote is transferred to a bibliography or user without verification, serious misinformation will occur.
3. Which of the following is the area with the highest risk of fabrication when producing metadata with artificial intelligence and how should it be addressed?
- A) The language area is the most at-risk area and does not need to be verified
- B) The title field carries the highest risk and should be filled with prediction
- C) ISBN and publication year carry the highest risk of fabrication and should be verified verbatim with the original source ✔
- D) No field is risky because artificial intelligence always produces metadata correctly
Explanation: Exact, numerical fields such as ISBN and publication year are very easily forged by AI; AI can generate a reasonable-looking number instead of an ISBN it doesn't know. These fields must be verified verbatim with the original source (barcode, imprint/colophone).
4. What is the main purpose of using 'controlled vocabulary' when assigning a topic term?
- A) Ensuring terminology consistency and availability across records; Preventing artificial intelligence from making up free terms ✔
- B) Automatically add every new term requested by artificial intelligence to the catalog
- C) Eliminate subject terms completely and search only by title
- D) Slowing down cataloging and writing each record with different terms
Description: A controlled vocabulary is a list of approved and standard terms. If the same subject is written with different terms in different records, the user will find one and miss the other. By giving the AI this list and telling it to 'just select from the list', consistency and findability is maintained; A new term that is not listed by artificial intelligence should not be added to the catalog.
5. What is the correct approach when getting a document's classification number from artificial intelligence?
- A) The exact number given by artificial intelligence should be accepted directly without looking at the official table
- B) Classification numbers should not be used at all, sources should be randomly shelved
- C) The master class proposal should be considered preliminary, the exact steps should be verified with the official classification table ✔
- D) Since artificial intelligence cannot classify, this task should be skipped completely
Explanation: AI can often find the right major class, but may miss sublevels (e.g. medical history instead of medicine) or produce incorrect numbers that seem reasonable. Therefore, the master class proposal should be considered a start and the exact steps should be verified with the official classification table.
6. Which statement is true when comparing semantic search and keyword search?
- A) Semantic search should replace keyword search in all cases
- B) Semantic search finds semantically related resources; Keyword search for exact title or ISBN is more reliable ✔
- C) Keyword search finds sources with different terms that are semantically related better than semantic search
- D) Semantic search only works on numerical data, not on text
Description: Semantic search finds different-word but related resources, such as 'sleep problem' and 'sleep disorder', because it matches the meaning of words (via embedding vectors). In contrast, when looking for an exact ISBN, full title, or statute number, a keyword/domain search is more reliable. Best practice is to use both together.
7. What is the librarian's duty against the risk of a 'filter bubble' in personalized resource recommendations?
- A) Strengthen the bubble by only recommending the most similar resources to the user
- B) To give up making suggestions completely
- C) Sharing the user's reading history publicly
- D) Consciously adding sources with different perspectives and horizons and stating this transparently ✔
Explanation: The filter bubble is when the recommendation system constantly locks the user into similar content and keeps them away from new and different perspectives. The value of the library is equal and diverse access; Therefore, the librarian consciously breaks the bubble by adding different-looking, eye-opening resources and stating this transparently.
8. Which practice is correct in terms of 'authenticity and integrity' when processing an archive document with artificial intelligence?
- A) 'beautify' the document with artificial intelligence, complete it by making up the missing parts and replace the original
- B) Deleting the raw original and keeping only the version improved with artificial intelligence
- C) OCR should only correct recognition errors, content should not be modified, the raw original should be preserved and derivatives should be clearly labeled ✔
- D) Accept the OCR output as 'clean text' without any corrections and open it for search.
Description: The essence of archival work is the assurance that the document comes from the claimed source and that its content has not been altered. OCR correction should only remove recognition errors, not change the content; AI 'restoration' is not a substitute for actual evidence as it can make up missing parts. The raw (original) scan should always be preserved, derivatives should be clearly labeled.
9. Which of the following is necessary for responsible construction of a library information bot?
- A) The bot produces answers to each question quickly from its general knowledge and does not cite sources
- B) The bot relies only on verified corporate information, cites facts and directs sensitive questions to an expert/human ✔
- C) The bot directly gives precise advice on medical and legal questions
- D) Hiding from the user whether he is talking to a bot or a human
Explanation: The advice bot should rely only on the verified current information and source list of the institution, not on its general memory, should give the facts by citing the source, should not give advice on medical/legal/sensitive issues, should direct it to the expert, and should transfer it to the human when necessary. Thus, the bot does not replace the human, but rather becomes a layer that eases the burden.
10. In bibliometrics, what is the correct method to find out an author's citation count or h-index?
- A) Asking the AI directly, because the numbers it gives are always correct
- B) Getting the numbers from the actual citation database; Using artificial intelligence only to organize and interpret this data ✔
- C) Estimate the numbers and use them in broadcast decisions
- D) Considering the number of citations as a direct measure of the quality of an author's work.
Explanation: Artificial intelligence can make up values such as citation count and h-index without accessing real data (for example, it can give a value of 42 when it is actually 19). Therefore, numbers and citations must be taken from a real citation database; AI should only be used to cluster, sort and interpret this real data.
11. Which is true about the limit of bibliometric measures (number of citations, h-index, impact factor)?
- A) Metrics directly measure quality and are comparable across domains
- B) Criteria can be used safely alone in hiring and promotion decisions.
- C) Metrics indicate usage, not quality; cannot be compared across domains and must be balanced with qualitative assessment ✔
- D) The outputs of the humanities are fully and completely represented in citation databases
Disclosure: These metrics indicate usage and visibility, they are not direct measures of quality. Because the citation culture of the fields is different (such as medicine and philosophy), they cannot be directly compared; Book-heavy outputs in the humanities are underrepresented in databases. Therefore, criteria should not be used alone to evaluate individuals, but should be balanced with qualitative evaluation.
12. What is the key advantage of the RAG (access-augmented manufacturing) approach over a general AI?
- A) Significantly reducing the hallucination by linking the answer to the real documents of the institution and citing the source ✔
- B) Always respond perfectly, regardless of the quality of the document base
- C) Making hallucination completely and permanently impossible
- D) Opening all documents to everyone without the need for access control
Explanation: When answering a question, RAG first finds the relevant pieces from the institution's own documents and produces the answer based only on these pieces, citing the source. This way the answer is linked to real and traceable documentation and the hallucination is greatly reduced. However, RAG does not completely eliminate the hallucination, as the answer may still be wrong if the wrong piece is brought up.
13. In the humanities and social sciences, what is the most reliable method for verifying a quote from an AI to a user?
- A) Make sure the quote looks believable and fluent
- B) Just asking the AI to reproduce the same quote
- C) Verify the quote by finding it in the author's actual work ✔
- D) Accepting that it is sufficient that the quote is attributed to a famous name
Description: Artificial intelligence can generate completely fabricated quotes attributed to a real author; The fact that the quote seems word for word accurate and convincing is not proof of its accuracy. The most reliable method is to verify the quote by finding it in the author's actual work; The logic of 'this author might have said something like this' is not enough.
14. Which principle is correct in using artificial intelligence in terms of copyright, privacy and ethics?
- A) Copyrighted and personal data can be entered into public tools without thinking because it is fast
- B) If a use is legal, it is definitely ethical and no further evaluation is required.
- C) When artificial intelligence produces an output, the responsibility completely passes to the tool, not to the human.
- D) Copyright and privacy must be respected from the very beginning; anonymization, transparency and consent are essential and the ultimate responsibility lies with the human ✔
Explanation: Impulsive input of copyrighted content and personal data (user ID, reading/search history) into general artificial intelligence tools may result in a violation of rights and privacy; Anonymization, data minimization, transparency and consent are essential. Additionally, even if a use is legal, it may be unethical; 'everyone uses it' or 'it happens fast' does not justify a practice and the ultimate responsibility lies with the expert, not the tool.