Unit 7 / 11

Data-Based Decision and Public Data Analysis

Gains:

  • Recognizes the quality, confidentiality and interpretation hazards of working with data and establishes a secure analysis flow.
  • It selects the right indicator for a purpose, distinguishing between misleading single indicators and measurements that are open to manipulation.
  • It recognizes misleading chart techniques (axis not starting from zero, chosen range) and observes honest visualization.

The public produces data every day: application numbers, solution times, budget items, population movements, service requests, audit results, satisfaction surveys. Most of this data sits in Excel spreadsheets, forms, and recording systems, but does not translate into decisions — because analyzing it requires expertise and time. Data-based decision-making (an approach to allocating resources and setting policy based on measured evidence, rather than opinion or habit) increases both efficiency and fairness in the public sector: resources go where they are needed most based on numbers. Here, AI is a powerful assistant in making sense of a table, finding patterns and anomalies, extracting summary statistics and translating the findings into simple sentences. In this unit, you'll see how to shorten the path from raw public data to a defensible decision with AI and stay disciplined in the most critical aspect—the correct interpretation of the number.

Three dangers of working with data

Data is powerful, but it has three traps, and AI can both mitigate and magnify these traps:

  1. Correlation ≠ causation. Just because two things change together doesn't mean one causes the other. Ice cream sales and drownings are on the rise — because they are both linked to summer heat, not one causing the other. AI fluently constructs “because” sentences; Question every claim of causality.
  2. Biased/missing data. From whom and how was the data collected? Complaint data reflects only those who can complain; digital application data those with access to the internet. If you don't see the missing group, the decision will be biased.
  3. Number without context. "Increased by 30%" alone is meaningless: according to what, in what period, what is the absolute number? Ask the AI ​​for context.
Attention: While the AI ​​analyzes the table you provide, it sometimes adds "reasonable" but fictitious numbers from outside the table or silently fills in the missing cells. Compare each output number to the source table; Give the AI ​​the rule "use only the values ​​in the table I give you, if missing say 'no data'".

Step by step data analysis flow

  1. Clarify the question. What question are we looking for answers to for the decision? ("In which neighborhood does the solution time take the longest?")
  2. Know the data. What are the columns, what is the unit, what is the period, are there any missing/incorrect values?
  3. Extract personal data. Anonymize/aggregate if individual identification is not required for analysis (see Unit 11).
  4. Summary and pattern emerge. Have the AI ​​perform basic statistics, grouping and anomaly marking.
  5. Question the comment. Correlation or causation? Alternative explanation? Missing group?
  6. Visualize and simplify. A simple summary and graphic that the decision maker can understand.
  7. Document the decision and its reasoning. What decision was made based on what data? audit trail.

three mini cases

Case 1 — The source went to the right place. In the past, a municipality distributed its park maintenance budget "equally to each neighborhood." When 18 months of demand and usage data were analyzed with AI, the per capita demand of the three neighborhoods was 2.4 times the average. The budget was redistributed according to need; overall satisfaction increased, without additional expense.

Case 2 — Biased data was noticed. An agency would look at digital application data and conclude that “citizens are no longer coming to the toll booth.” But when the age breakdown was requested in the analysis, it was seen that applications over the age of 65 were still mostly coming from the counter. If only digital data were looked at, this group would be ignored; box office service was preserved.

Case 3 — Contrived causation stopped. While interpreting a table, YZ said, "Accidents decreased because the number of inspections increased." The analyst reminded that the weather also changed and traffic decreased in the same period; Attributing it to a single cause was misleading. The report was corrected to say "there is a relationship but causation has not been proven."

Four copyable templates

1) Getting to know the table:

Your role: data analyst. Examine the table below and, based ONLY the values ​​in this table: (1) summarize what each column is, its unit and period, (2) any cells that appear missing/incorrect, (3) summarize the top 3 patterns that stand out. Adding numbers from outside the table; if missing, say "no data".TABLE: [paste data]

2) Summary statistics with context:

Answer the following question from the table below: [question]. MUST give the result with context: absolute number + rate + comparison period. When you say "increased by X%", specify to what extent and in what period. Just use the given data.TABLE: [data] QUESTION: [question]

3) Causality questioner:

Interpret the following finding, but CAUTION: do not confuse correlation with causation. Generate at least 3 alternative explanations for this relationship (third variable, reverse direction, coincidence). Tell me what additional data is needed to prove causality. FINDING: [relationship]

4) Simple summary for the decision:

Simplify the following analysis for a decision maker who is not a data expert: no more than 5 items, each item a finding and a “what does this mean” explanation. Also write down the uncertainties and limits of the data clearly. Adding new information.ANALYSIS: [text]

Weak prompt / Strong prompt

Weak: "Analyze this chart and tell us what we should do."

Strong: "Your role is data analyst. ONLY use the table below; do not add numbers from outside, say 'no data' to the missing cell. First, get to know the columns, units and period. Then answer the question 'which neighborhood has the longest solution time' with absolute number + ratio + comparison period. If you find a pattern, come up with 3 alternative explanations before explaining it with causality. Finally, list the points and boundaries of the data (missing group, short period) that leave the decision to us."

Difference: strong prompt establishes data fidelity, context, causality discipline and decision-boundary together; The output is a defensible analysis.

Where can AI be trusted in the data business?

business

AI trust

rule

Table summary, grouping

high

Data given only

Anomaly/pattern marking

medium

human confirmation

Causality interpretation

low

Alternative explanation is required

Filling in missing data

too low

Don't make it happen; "no data"

Simplification/visual summary

high

Meaning must be preserved

Indicator design and misleading chart trap

The most critical step in making decisions with data is choosing the right indicator (the numerical expression that makes a phenomenon measurable—for example, “average processing time” or “cost per application”). The wrong indicator leads to the wrong decision, even with accurate data: while the “total number of resolved claims” appears to be increasing, the “waiting time per citizen” may be worsening. AI can list candidate indicators for a topic and what each measures and conceals; but you decide which indicator truly represents the public interest. Also, be aware of the risk of misleading visualization (distorting perception with techniques such as the axis not starting from zero, disproportionate scale, selected time interval) in the graphs that the AI ​​suggests or describes; honest visual in public data is part of governance.

Mini case — wrong sign, wrong boast. One unit was breaking records in the "number of files closed per month" indicator; but citizen complaints increased. Looking at the accurate indicator, the “first contact resolution rate,” the rate dropped from 58 percent to 44 percent — files were being closed and reopened quickly. When the indicator changed, the management priority also changed.

Mini case — axis that does not start from zero. On the graph AI described in a presentation, the vertical axis starts at 40, a 2 percent increase seemed like a dramatic jump. When the axis was moved to zero, the real picture emerged and the decision was eliminated from the exaggerated perception.

Template that supports indicator selection:

Task: Suggest 5 candidate indicators for the following purpose.Purpose: [e.g. improving citizen service quality]For each indicator: what it measures | what can it hide | openness to manipulation | recommended reading frequency.Next: write 3 warnings (those that are misleading on their own) to read these indicators together.Rule: Do not produce numbers; just suggest indicator design.

Caution: Do not fall into the misconception that "number is objective". What number you measure, how you display it, and what range you choose are all a matter of interpretation. Transparent and honest public interpretation is as important as having accurate data.

Common mistakes

  • Mistaking correlation for causation. "Increased together" is not the cause; Look for alternative explanations.
  • Having AI fill in missing data. The analysis collapses with a made-up value; leave the missing missing.
  • Generalizing biased data. Complaint/digital data does not represent everyone; Query the missing group.
  • Presenting the number without context. Next to the ratio there should be an absolute number and comparison period.
  • Unnecessarily analyzing personal data. Do not use individual identification if aggregated data is sufficient.
  • Not documenting the decision. What decision was made based on which data should be recorded for auditing.

In summary

Data-based decision making is the most powerful way to distribute resources fairly and efficiently in the public sector; AI is a fast assistant in this business that makes sense of the picture, finds patterns and simplifies the finding. But don't make the AI ​​make up numbers, don't make it fill in missing data, don't let it confuse correlation with causality, and avoid generalizing biased/incomplete data. There is strength in numbers; The person who interprets it correctly and documents the decision is stronger.

Application task

Get a real (personal data stripped) table from your unit. Introduce the data with the "Getting to know the table" template, then have an important question answered for the decision with the "Contextual summary statistics" template. Apply the "Causation questioner" template to an emerging relationship and find at least one alternative explanation. Check if the AI ​​is adding numbers from outside the table.

checklist

  • [ ] The AI only used the given table; He did not add numbers from outside.
  • [ ] I did not fill in the missing data; I left it as "no data".
  • [ ] I presented each result with context (absolute number + rate + period).
  • [ ] I produced an alternative explanation to the causality claims.
  • [ ] I evaluated the risk of biased/missing data and the missing group.
  • [ ] I documented the decision and the data on which it was based.