Unit 3 / 11

Sampling, Representation and Bias

Gains:

  • Being able to clearly define the universe and evaluate the representativeness of the sampling method and who is systematically excluded
  • Ability to distinguish between bias in data and stereotypes carried by artificial intelligence and control both
  • Ability to express findings limited to a sample, without exaggeration, and write an honest limitations section.

The most crucial concept of social research is the sample — that is, the subgroup you choose because you cannot examine the entire group you are interested in (the universe/population: all the people or units that the research wants to talk about) one by one. To whom the findings of a study can be generalized (i.e., "whether this result is valid only for these people or for a broader group") depends entirely on the representativeness of the sample. The result of a wrong or biased sample is wrong, no matter how well it is analyzed. In this unit, you will learn how to use AI to think about sample design, detect bias — systematic drift of data or interpretation in one direction, and honestly express the limits of generalization.

It is necessary to distinguish two types of prejudice here. The first is bias in the data: the sample overrepresents one group and underrepresents another group (e.g. those who can only participate in the online survey, i.e. those with internet access). The second is the AI's own bias: the AI ​​carries the patterns of the texts it has been trained on and can produce stereotypical generalizations about a social group. A good social researcher must be alert to both; AI can help you recognize both, or it can magnify both.

Step by step: thinking about representation

1. Describe the universe. Who do you want to tell about your finding? "All voters in Türkiye" and "young voters in a neighborhood in Istanbul" are very different universes. A sample cannot be established without writing this clearly.

2. Select the sampling method. Methods such as random (everyone has an equal chance of being selected), stratified (dividing the universe into groups and selecting from each group), snowball (one participant recommending another) give different powers of representation. AI can explain the pros and cons of each method in plain language.

3. Ask who is left out. Every sampling misses someone. Online survey misses those without internet, phone survey misses those without landline, daytime interview misses employees. AI produces a good checklist for “who does this method systematically exclude?”

4. Map sources of bias. List sources such as selection bias (who is chosen), response bias (who answers), recall bias (misremembering the past).

5. Write the generalization limit. Phrase the finding as “young people in this sample,” not “all teens.” AI can flag overgeneralizations in your draft.

Tip: A good research report is strengthened, not weakened, by the "limitations" section. Writing honestly about who your sample does not represent makes your study more reliable. Have the AI ​​draft this section, but confirm any limitations yourself.

three mini cases

Case 1 — Hidden selection bias. To measure satisfaction with a municipality's services, one team put a survey on the municipality's website and found 78% satisfaction. When I asked AI "who is this sample missing?", it turned out that the segment that already uses the site (that is, has access to the service and is probably more satisfied) is over-represented. The finding was corrected to “78% among site users.”

Case 2 — Stereotype of AI. When a researcher had AI summarize “rural elderly people's view of technology,” AI added a stereotypical frame that was not in the data, such as “elderly people are afraid of technology.” When the researcher realized this and said "just rely on the statements in the transcript", it was seen that there was actually a wide variety of attitudes in the data.

Case 3 — Stratified sampling correction. In a sample of 500 people, women were 30%, compared to 51% in the population. AI explained the weighting logic in plain language, and the team weighted the results according to the population proportions, resulting in a more representative picture.

Four copyable templates

1) Criticism of the sample:

The universe of my research: [description]. My sampling method is: [method]. Your task: list who this sample might systematically under- or over-represent. Consider choice, response, and recall errors separately. Give a mitigation recommendation for each risk. Do not exaggerate my finding; Show boundaries honestly.

2) Generalization control:

Examine these finding sentences: [text]. Mark each overgeneralization. Rewrite statements like “all/everyone/young people/women” with a narrower statement that the data actually supports. For example, instead of “young people do X,” “Y% of teens in this sample reported X.” Flag any claim of uncertain origin.

3) AI bias capture:

I gave you the summary below. Check: are any judgments, adjectives or generalizations you added REALLY present in the given text? Mark and remove any expressions that are not in the text but come from your own patterns. Just stick to what's in the text.

4) Draft limitations section:

Write a draft "Limitations" section based on the following method information: sample size [n], method [x], context [y]. Explain how each limitation may affect the findings. Neither exaggeration nor underestimation; Be balanced and honest. I will confirm every item.

Weak prompt / Strong prompt

Weak prompt:

According to my survey results, most young people have this opinion. Write this in a strong sentence.

This produces overgeneralization without ever questioning the sample. AI easily slips into a claim that the data does not bear, such as "Young people in Türkiye..."

Powerful prompt:

My sample: 220 undergraduate students at a single university, online survey, voluntary participation. I would like to express a finding: 61% of the participants expressed opinion X. Write this in a no-frills academic sentence that clearly states the limits of the sample (representation, volunteering bias). Limit generalization to this sample.

The difference: the second sentence produces a defensible claim that the data actually supports.

Sampling methods and representativeness

Method

How does it work

power of representation

Typical risk

simple random

Equal chance for everyone

high

transportation cost

stratified

Divide and select into groups

high

Group definition error

quota

Filling group rates

medium

In-group bias

easy

Don't ask anyone who can be reached

low

Severe selection bias

snowball

By participant suggestion

low

Intranetwork similarity

Common mistakes

  • Establishing a sample without defining the universe. Representation cannot be evaluated without knowing to whom you will generalize.
  • Generalizing convenience sampling to the entire population. The most common mistake; "survey respondents" are confused with "everyone".
  • Never asking who is left out. Every method kidnaps someone; Look for this consciously.
  • Not noticing the stereotype added by AI. The model may add stereotypes that are not present in the data.
  • Hiding limitations. Hiding the frailty of the sample makes the study fragile, not reliable.

In summary

The value of a finding is as much as the representativeness of the sample on which it is based. AI; is a powerful aid in explaining sampling methods, mapping out who is underrepresented, flagging overgeneralizations, and drafting an honest limitations section. But beware of two biases: the bias of the data itself and the stereotypes carried by the AI. Define the universe clearly, ask who is left out, limit generalization to the sample, and write down limitations honestly.

Application task

Describe the population and sampling method for a real or hypothetical study. Extract who might be under/over-represented from the AI ​​with the “sampling critique” template and identify at least three sources of bias. Then translate a finding statement into a narrow statement that the data actually supports with the “generalization check” pattern. Finally, write an honest half-page limitations section.

checklist

  • [ ] I have clearly defined the universe (to whom I will generalize).
  • [ ] I evaluated the representativeness of the sampling method.
  • [ ] I listed who was systematically left out.
  • [ ] I expressed the findings limited to the sample and without exaggeration.
  • [ ] I removed the stereotypes/generalizations that AI added.