Gains:
- Recognize direct and indirect identifiers, metadata and digital traces that might give away the source and ask 'how many people would this description fit?' Ability to apply anonymization with test
- Gain the discipline of choosing only security-verified tools for sensitive work, applying the principle of least data, and cleaning up traces
- Ability to understand that KVKK and professional confidentiality necessitate this discipline both legally and morally and that once a leaked identity cannot be taken back
In journalism, a source is a person who speaks, often at great risk: a public employee who exposes corruption, a worker who describes wrongdoing, someone who could lose his job, his freedom, or even his safety if his identity is revealed. Source protection — the principle of keeping the identity of the informant confidential — is one of the profession's most sacred obligations and is protected in most legal systems. The age of artificial intelligence (AI) both simplifies this task and introduces new risks: the tool you use to understand a document may hide that document; If you fail to anonymise, the traces you leave on the AI may reveal the source. The principle of this unit is clear: No data that could give away the source enters an unsecured AI tool; Anonymization and digital hygiene are done before using AI, not after. Once leaked, the identity cannot be retrieved.
What are the traces that give away the source?
Resource protection is not just "no name". The following traces may also reveal identity:
- Direct identifiers: name, telephone, e-mail, TR ID, address, unit of employment.
- Indirect descriptors: descriptions such as “the only female engineer working in a circle of three”; It refers to a single person in a small group.
- Document metadata: author name, edit history, creation date, GPS location, printer track embedded in a file.
- Contextual clues: who has access to the document; timing of the leak; a specific event in the narrative.
- Digital trace: from which device, from which network, with which account was contacted.
Critical risk for AI: sticking any of these traces into a tool takes the data out of your control. Most free/public AI tools store inputs or can use them in model refinement.
Step by step: Using AI while saving resources
Step 1 — Classify the vehicle. Only tools with verified security and data processing conditions (preferably corporate, non-data storage or offline) are used for sensitive work. Sensitive data is not entered into public/free tools.
Step 2 — Clean metadata. Delete metadata before exporting the document to the tool; If possible, convert the document to plain text or manually summarize the content and anonymize it instead of a screenshot.
Step 3 — Anonymize. Generalize all direct and indirect identifiers: “Source A,” “a public official,” blur dates and locations. Change recipes that refer to single people in small groups.
Step 4 — Minimum data policy. Give the AI only what the job requires. Instead of summarizing the entire document, give the relevant non-identifying section.
Step 5 — Verify and save. Test the adequacy of anonymization with a check step (template below). Keep communication with the source on separate, secure channels.
Step 6 — Clean the tracks. Once the job is done, clear the history/session in the tool; Securely store or delete sensitive files.
Attention: Saying "I anonymized" is not enough; In a small group, even a single indirect description reveals identity. You can use anonymization to ask “how many people would this recipe fit?” Test it with the question. If the answer is "one", the resource is not protected.
KVKK and professional confidentiality
In Türkiye, KVKK (Personal Data Protection Law) regulates personal data; Information such as health and political opinions are special data and require higher protection. Giving the source's data to an arbitrary tool may result in both ethical and legal liability. Professional confidentiality is a duty of honor independent of the law.
three mini cases
Case 1 — Metadata leak prevented. A reporter was about to summarize a leaked report; At the last moment, he realized that the organizer's name was embedded in the file's metadata. He used it after converting the document to plain text and removing the name. If the raw file were uploaded, the metadata would directly give away the source.
Case 2 — Indirect identifier danger. A reporter used a phrase when consulting YZ, describing a source as "the only disabled employee at headquarters." Whether or not the AI stored the text somewhere, this description pointed to one person. The editor noticed, changed the description to "an agency employee." In the small group, the indirect description was as dangerous as the name.
Case 3 — Choosing a safe vehicle. When analyzing a highly sensitive batch of documents, an investigative team chose to use a no-data/offline solution rather than a generic AI tool; additionally anonymized the documents. Analysis was a little slower, but at no point was the security of the resource compromised.
Copiable templates
1) Anonymization proficiency test:
Mark each element in the text below that might reveal the identity of the source: direct identifiers (name, title, unit) and indirect identifiers (descriptions that point to a single person in the small group, specific events, date/location). For each sign, “how many people would this description fit?” Ask the question and suggest safer generalizations. Change text. Text: [here]
2) Checking before document sharing:
Make a security checklist I should do before feeding a document to an AI tool: metadata cleansing, direct/indirect identifier scanning, minimal data policy, tool security class, history cleaning. Add a single sentence "how to" for each item.
3) Source communication risk assessment:
I will contact a sensitive source. Create a checklist of digital security measures that will protect your identity and communications (channel selection, leave no trace, metadata, device hygiene). Give concrete and actionable suggestions; do not promise absolute security.
4) Pre-publication source protection control:
Include in the following ready-to-publish news any details that might identify the source: indirect descriptions, timing of attribution, number of people who had access to the document, specific event details. Highlight each risky spot and suggest how to blur it. News: [here]
Weak prompt / Strong prompt
Weak prompt: "Summarize this document." (by uploading a sensitive, metadata source document as is)
Result: Metadata and indirect identifiers are out of your control; The source may be compromised without you being aware of it.
Strong prompt: First convert the document to plain text and delete the metadata, generalize the direct/indirect identifiers, give only the relevant part: "Summarize the following anonymized text; also mark any remaining identity clues in it."
Difference: The strong approach cleans data without giving it to the vehicle, shares minimal data, and uses AI as an additional layer of control; the resource is protected.
comparison chart
track type
example
Risk
precaution
Direct identifier
Name, phone, unit
high
Delete/generalize
indirect identifier
"The only female engineer"
High (hidden)
"How many people does it fit?" test
metadata
File author/date/GPS
high
Convert to plain text, clear
Contextual clue
Leak timing
medium
Blur time/place
digital trace
Device, network, account
Medium-High
Safe channel, hygiene
Common mistakes
- Forgetting metadata. Uploading the file as is and leaking embedded author/date/location information.
- Just delete the name. Ignoring that indirect identifiers also reveal identity.
- Exporting sensitive data to the public tool. Entering source data into tools that store data/use it in model training.
- Sharing too much data. Increasing the risk surface by providing more documents/details than the job requires.
- Not cleaning the tracks. Leaving session history and sensitive files when the job is done.
Safe vehicle selection: what to look for?
Deciding whether an AI tool is “safe” for a sensitive job is a technical literacy that a journalist must acquire. The main questions to look at are: Does the tool store the data you enter, for how long and where? Are your inputs used in model training (yes in most free consumer tools, often no in enterprise versions)? In which country is the data processed and under what law? Can the tool work offline (on your own device, without sending data to the internet)? Does your organization have a data processing agreement with this tool? Entering sensitive resource data into a tool without knowing the answers to these questions is putting the resource at risk with blind trust.
A general hierarchy is helpful: for the most sensitive documents, solutions that run offline or on your own infrastructure are preferred; contracted enterprise tools that do not store data for medium precision; Generic consumer tools may only be used for non-identified, publicly available or anonymized content. Doing this classification at the beginning prevents the "entering the wrong data into the wrong tool in a hurry" mistake. Additionally, in a corporate newsroom, this decision should not be left to individual reporters but should be standardized with a written tool approval list; because a single mistake by a single person can destroy the source of an entire investigation. Vehicle security is the technical pillar of resource protection and should be taken as seriously as anonymization.
In summary
Resource protection poses new risks in the age of AI: tools can hide data, metadata and indirect identifiers can give away identity. Safe way; selecting the tool according to its security class, cleaning metadata, anonymizing all direct and indirect identifiers (the "how many people fit this description?" test), applying the principle of least data and cleaning traces. KVKK and professional confidentiality require this discipline both legally and morally. Once leaked, the identity cannot be retrieved.
Application task
Take a real or fictional source document/note. First follow the "Anonymization proficiency test" prompt; Mark each direct and indirect descriptor that appears and ask "how many people would it fit?" Evaluate with the question. Then securely anonymize the document and manually apply the "Document pre-sharing check" list once.
checklist
- [ ] I use only security-verified tools for sensitive work.
- [ ] I clear the metadata before exporting the document.
- [ ] Uses all direct and indirect descriptors as "how many people does it fit?" I anonymize it with the test.
- [ ] Follows the least data principle; I only share the necessary part.
- [ ] I clear session history and sensitive files after work; I observe KVKK and professional confidentiality.