Unit 8 / 12

Interview and Survey Analysis: From Qualitative Data to Theme

Gains:

  • Ability to divide interview transcripts and open-ended survey responses into themes with artificial intelligence and support them with evidence citations
  • Ability to summarize quantitative survey data and avoid misleading interpretation (small sample, leading question, selection bias)
  • Ability to work ethically and KVKK compliant by protecting participant confidentiality and anonymizing transcripts

The richest insights in consulting often lie not in charts but in what people say. A manager interview, open-ended responses to a customer survey, notes from a field visit... These are non-numerical (qualitative) data and it takes hours to read, code, and divide into themes. Artificial intelligence greatly speeds up this work: It breaks down 30 interview transcripts into themes in minutes, finding recurring patterns. But there are two dangers: the model may “make up” a theme (add a comment that is not in the transcript), and participant confidentiality may be easily violated. In this unit, we will learn to analyze qualitative data in an evidence-based, safe manner.

Qualitative analysis: from coding to theme

Qualitative analysis has three steps. Coding first: each statement is given a short label (“price complaint”, “delivery praise”). Then theming: similar codes are collected into a theme (“operational dissatisfaction”). At the end, commentary: the relationship between themes and the "so what" layer are established. AI is very fast in the first two steps; The third step requires the counsel's judgment.

The most critical assurance is this: each theme must be supported by a verbatim quote. This confirms that the theme actually comes from the data; The quote-unquote theme carries the risk of hallucination.

Your role: qualitative research analyst.Below is an anonymized transcript of 12 customer interviews (labeled P1, P2...).Task:1) Extract recurring themes (up to 7 themes).2) For each theme: how many participants saw it + 2 verbatim quotes (with participant code).3) Show separately statements that support and CONFLICT the theme.Rules:- Do not add any opinions not in the transcript; Each theme must be proven with quotation. - Do not produce identification information other than participant codes. - List single "minority but important" opinions separately at the end (dissenting voices).

Another powerful technique is to ask the model to code with fixed categories that you specify, rather than having it set up its own code scheme first. For example, if you give predefined themes in a customer satisfaction study such as "price, product quality, delivery, support, ease of use", different transcripts become comparable and the model is prevented from making up different names each time. Free theme extraction is suitable for discovery, whereas fixed category coding is suitable for comparison and measurement; For most projects, using both together works best: free exploration first, then digitize with a fixed schema.

Your role: qualitative coding assistant. Fixed categories: [price, quality, delivery, support, ease of use, other]. Task: Assign each open-ended response below to one or more of these categories. - Write the trigger phrase in the response (verbatim) next to each assignment. - Tick whether the response is positive or negative in sentiment. Adding comments not included in the answer.

Tip: It is very valuable to ask for “minority but important” opinions separately. A once-worded but critical alert (e.g. a single large customer saying "we are considering renewing the contract") gets lost in the majority themes; remains visible when listed separately.

Privacy and anonymization

Interview transcripts are personal data and are often sensitive (employee complaints, manager criticisms). Before uploading them to an AI tool, it is essential to anonymize them: names, titles, internal distinguishing details are replaced by the code number (K1, K2...). The participant is often promised that "what you say will be reported anonymously"; The technical equivalent of this phrase is anonymization.

Data type

Risk

precaution

Participant name/title

identity disclosure

Replace with code number

Internal detail

indirect diagnosis

Generalize or mask

sensitive vision

Risk of retaliation

Individual matching with bulk report

Audio/video recording

highest sensitivity

Render in approved tool, then delete

You can have the model perform anonymization as a preliminary step before uploading the transcripts to the tool; But be sure to visually check the output:

Your role: privacy editor. Prepare the following interview transcript for analysis; do the following masking:- Make all contact names K1, K2...- Generalize indirectly identifying details such as title, department, location ("The only foreign engineer in R&D" → "an engineer").- Remove company/customer names. Give only the masked text; Briefly list what type of information you are masking and why.

Caution: Saying "Anonymous" is not enough. The phrase "The only female manager responsible for marketing" identifies the person even though her name is not written. Anonymization should also include all indirectly identifiable details. KVKK violation and loss of participant trust start from here.

Interpreting quantitative survey data

Artificial intelligence produces a summary and breakdown of closed-ended (numerical) survey results. But quantitative analysis has its own pitfalls, and the model can ignore them:

  • Small sample: Saying "61% of customers are satisfied" from a survey of 18 people is misleading; If the sample is small, the raw number and uncertainty should be discussed, not the percentage.
  • Leading question: "How satisfied are you with our excellent service?" The question embeds the answer; The result is biased.
  • Selection bias: If only the most satisfied or angriest customers responded, the result would not be representative of the universe.

Before summarizing the results of this survey, conduct a methodological check: - Is the sample size sufficient to present the results in percentages? - Are there leading/leading statements in the questions? Tick.- What is the response rate and risk of selection bias? Provide audit report only; We will then summarize it in a separate message.

three mini cases

Case 1 — Made-up theme caught. A consultant assigns 20 interviews to the model; the model extracts the theme “participants want to work remotely.” The consultant asks for a quote; The model can only find 1 ambiguous quote. The theme did not actually exist, the model made it up from the general trend. If there were no citation requirement, an incorrect finding would be included in the report.

Case 2 — Indirect diagnosis. In an employee satisfaction project, the transcript includes "The only foreign engineer in R&D said this." There is no name, but the person is clear; Once the report goes to the manager, there is a risk of retaliation. The consultant had to generalize from the beginning all the distinguishing details ("an engineer").

Case 3 — Small sample trap. In a product test, 15 out of 22 users liked it; the model says "68% satisfaction" and is placed on the slide. The investment board objects: "22 people do not represent the universe." The consultant honestly rewrites the result as "15 out of 22 users positive; indicative, confirmatory large sample required"; trust returns.

Weak prompt / Strong prompt

Weak prompt:

Read these interviews and extract the main findings.

The model produces themes without quotes, without evidence; The risk of fabrication and confidentiality is high.

Powerful prompt:

Your role: qualitative analyst. Transcript below, anonymized (K1..Kn).Task: Up to 7 themes; number of participants for each + 2 verbatim quotes (with code); supporting and contradictory statements separately; minority but critical opinions separate list. Rules: adding opinions not in the transcript; identity generation; Do not give a theme without quotes; mark the theme you are not sure of as "weak evidence".

Saturation and weight of themes

In qualitative analysis, two concepts determine the quality of the interpretation. The first is saturation: data are said to be "saturated" when new interviews no longer produce new themes; This is a sign that the sample is sufficient. If brand new themes still emerge at the 8th interview, more interviews are needed. Secondly, the weight of the themes: how many people see a theme and the importance of that theme are different things. A minor inconvenience mentioned by five people and the risk of contract cancellation mentioned by one person are not equal in frequency, but the latter may be much more critical in importance. Therefore, ranking themes by frequency alone is misleading; Evaluate each theme in terms of both "how many people" and "how important". Having the AI ​​say “rate these themes in two separate columns of frequency and potential business impact” makes this distinction visible; You make the business impact judgment.

Common mistakes

  • Quoteless theme acceptance. Each theme must be evidenced by a verbatim quote from the transcript; Otherwise there is a risk of hallucination.
  • Making anonymization superficial. It is not enough to delete the name; Indirectly identifying details should also be masked.
  • Presenting small sample with percentage. With a small number of participants, raw numbers and uncertainty should be discussed, not percentages.
  • Ignoring the leading question. The outcome of leading questions is biased; Check the methodology.
  • Losing outliers. Once spoken, the critical warning is swallowed by the majority themes; list separately.
  • Mixing quantitative and qualitative. Process open-ended comments with numerical results in different ways.

In summary

Qualitative data is the richest source of insight in consultancy, and artificial intelligence saves hours in coding and theming. But each theme should be evidenced with verbatim quotes, participant confidentiality should be fully protected, including indirect identification, and methodological pitfalls such as sampling, question bias, and selection should be controlled in quantitative surveys. The model is a fast reader and labeler; it is the consultant's job to verify the evidence, maintain confidentiality, and establish the "so what" layer.

Application task

Prepare several interview transcripts (or open-ended survey responses), either real or representative, and anonymize them first (including name and indirect identification). Extract themes, quotes, conflicting statements, and dissenting voices with strong qualitative prompting. Check the citation of at least one theme in the transcript to verify that it actually exists. Then take a small quantitative survey and evaluate sample and question bias by requesting a methodological audit.

checklist

  • [ ] I anonymized the transcripts, including name and indirect identification.
  • [ ] I evidenced each theme with a verbatim quote.
  • [ ] I showed conflicting expressions and contradictory voices separately.
  • [ ] In the small sample, I used raw numbers and uncertainty instead of percentages.
  • [ ] I marked the leading questions.
  • [ ] I evaluated the risk of selection bias.
  • [ ] I reported sensitive comments collectively without matching them to the individual.