The numbers are correct. Could the conclusion still mislead?
Information judgment · Lesson 4: Notice what statistics leave out
Estimated time: 20–25 minutes. Suggested preparation: Check sources and context.
Teaching note: All data are fictional and calculations can be checked. They describe no actual product or population.
What you will learn
- Report relative changes alongside absolute changes and denominators.
- Distinguish averages, typical outcomes and individual differences.
- Check sample selection, grouping and time windows.
- Write the narrower conclusion supported by data instead of merely calling it misleading.
1. An impressive headline
A new workflow cuts errors by 50%. An amazing result!
Would you adopt it immediately? Write the three numbers you most want to see and the criteria needed to call the result amazing. The calculation may be correct while scale, cost, error severity and the comparison remain unspecified.
2. Read relative and absolute change together
Suppose the old workflow produced 20 errors in 1,000 tasks and the new workflow produced 10 errors in another 1,000 tasks.
- Old error rate: 20 ÷ 1,000 = 2%.
- New error rate: 10 ÷ 1,000 = 1%.
- Absolute reduction: 2% − 1% = 1 percentage point.
- Relative reduction: (2% − 1%) ÷ 2% = 50%.
Both descriptions are correct. In this observation the new workflow had 10 fewer errors per 1,000 tasks. That does not guarantee future outcomes or attribute the difference entirely to the workflow without comparable conditions.
A one-percentage-point reduction is not a 1% relative reduction. With a zero starting value, the usual relative-growth formula divides by zero and cannot be used.
A small absolute difference is not automatically unimportant. Scale, consequences, costs and uncertainty all matter.
3. The average does not describe everyone
Five team members produce 10, 10, 10, 10 and 60 items per month. The mean is 20, the median is 10 and the range is 10 to 60.
“Average output is 20 items” is correct, but does not mean most members produce 20 or that 20 is everyone's minimum. One high value strongly influences the mean.
The mean describes the total divided equally across people; the median locates the middle ordered value. Neither supplies all information. Examine distributions, ranges, groups and raw counts when relevant.
4. Denominators, groups and selection can change the story
| Presentation | What may be missing | What to add |
|---|---|---|
| “Complaints rose from 10 to 20” | Service volume may have risen from 100 to 1,000 | Rates of 10% and 2%, with consistent complaint definitions |
| “Completers have a high pass rate” | People leaving after enrollment are excluded | Enrollment, completion and pass counts |
| “A has a higher overall success rate” | A may receive more easy tasks | Comparisons within difficulty groups |
| “Growth this month” | Seasonality, a low baseline or longer-term changes | A longer series using consistent definitions |
| “Significant improvement” | Significant may be used informally | Actual differences, sample sizes and uncertainty |
Group-level trends can differ from the total because totals weight groups by their sizes. Do not choose whichever aggregate or subgroup favors your view; first define the comparison you need.
5. Ask four questions
- What was measured? Identify the measure, unit, period and definition.
- Who was counted? Check denominators, sampling, withdrawals and missing data.
- Compared with what? Are baselines, periods, tasks and populations comparable?
- How certain is the change? Consider data volume, possible influences and the conclusion's scope.
Read chart values too. A truncated vertical axis can exaggerate visual differences without necessarily constituting fabrication; appropriate scales depend on the chart. Inspect labels, ranges and values before assessing the visual presentation.
This lesson uses descriptive calculations. A few numbers alone do not establish statistical significance, and before–after changes do not automatically establish causal effects.
6. Exercise one: what does a 50% reduction mean?
Fictional records: the old workflow produced 40 errors in 2,000 tasks; the new workflow produced 20 errors in another 2,000. Task difficulty and allocation are unspecified.
Calculate both rates, the absolute reduction and the relative reduction. Rewrite “The new workflow is amazing and should be adopted everywhere.” Identify a condition to check before adoption.
Explore the answer and explanation
Explanation
The rates are 2% and 1%. The absolute reduction is one percentage point and the relative reduction is 50%. The observed counts differ by 20 across these two equally sized batches.
A rewrite: “In the two recorded batches, the new workflow's error rate was 1%, compared with 2% for the old workflow. Task comparability, costs and further observations should be examined before adoption.”
Avoid an ambiguous “1% reduction” or a guarantee of 20 fewer errors every time. Check task allocation, staff experience or whether error-recording definitions changed.
7. Exercise two: do completers represent all enrollees?
A fictional program enrolled 100 people; 40 completed it and 36 of those passed a test. The other 60 have no test records. Advertising says “Participants have a 90% pass rate.”
Calculate the pass rate among completers and the proportion of all enrollees with recorded passes. Does this establish a true pass rate of 36% for everyone, or that all non-completers failed?
Explore the answer and explanation
Explanation
36 ÷ 40 = 90% describes completers. 36 ÷ 100 = 36% describes enrollees with recorded passes.
Missing scores are not known failures. If the question concerns how many enrollees have the ability required to pass, the full result is unknown. If a predefined certification rule treats non-attendance as no certification, a certification rate can be reported with that definition explicit.
A rewrite: “Of 100 enrollees, 40 completed and 36 passed. Completers had a 90% pass rate; 36% of enrollees have recorded passes, while non-completers have no test data.” This neither conceals withdrawals nor invents scores.
8. Exercise three: ahead overall, ahead in each group?
Two fictional methods handle tasks of different difficulty. Entries show successes / tasks:
| Method | Easy tasks | Hard tasks | Total |
|---|---|---|---|
| A | 81 / 90 | 1 / 10 | 82 / 100 |
| B | 19 / 20 | 16 / 80 | 35 / 100 |
Calculate overall and within-group success rates. Assess “A has a higher overall rate, so choose A at any difficulty.” If each handled 50 easy and 50 hard tasks, using observed group rates as provisional estimates, what success counts would you expect?
Explore the answer and explanation
Explanation
Overall rates are 82% for A and 35% for B. On easy tasks A has 90% and B 95%; on hard tasks A has 10% and B 20%. B has a higher observed rate in both groups, but A received more easy tasks.
With the same 50/50 mix, A's estimated successes are 50 × 90% + 50 × 10% = 50; B's are 50 × 95% + 50 × 20% = 57.5. This is a model-based expected value. Actual success counts are integers, and future results need not equal the estimate.
Overall rates describe what happened under this task mix; group rates help compare similar difficulty. Allocation and other differences still need checking, so the table does not establish B's causal advantage. Groups are small and uncertainty estimates are absent.
9. Review your statistical interpretation
Omissions can reflect simplification, oversight or deliberate selection. Missing information alone does not establish motives. Explain how it affects the conclusion.
10. A numbers-checking record
The next lesson will examine reasoning that exceeds the evidence, including correlation versus causation and cases versus general conclusions. Continue using the lesson navigation below.
Entries stay in this browser, separately for each language. Export a copy for backup.