Registered non-governmental non-profit organisation certificate No. 1052p

Articles

Everyday statistics: how to read numbers properly

15 min read

Mean, median and mode, spread, percentages and percentage points, samples, correlation and misleading charts: simple examples and practice.

We meet numbers every day: an average mark at school, the step count on a phone, a headline saying something has “doubled”, a slick chart on social media. Most of these numbers are correct, but they are easy to misunderstand. Statistics is not just for scientists; it is the skill of reading numbers properly. In this article we go through the key ideas without complicated formulas: the mean, median and mode (and when each one is the fairest), spread, the difference between percentages and percentage points, samples and how representative they are, why correlation is not causation, misleading charts and the danger of small samples. Then we show how to calculate all of this in Excel, give you a list of questions for checking the numbers in the news, and finish with a practice exercise and a checklist.

An important note: every number in this article is an illustrative example, not the result of a real measurement or study. They are made up purely to show the idea.

Mean, median and mode: when each one is fairest

In everyday speech “average” means one thing, but in statistics there are at least three ways to describe the “central value”.

The mean

Add up all the values and divide by how many there are. An illustrative example: Dilshod notes his daily step count from his phone for a week:

  • Monday 4000, Tuesday 4500, Wednesday 5000, Thursday 5000, Friday 5500, Saturday 6000, Sunday 21000.

On Sunday he went hiking in the mountains with friends. The total is 51000; divided by seven days, the mean is about 7286 steps. Yet on none of his “ordinary” days did Dilshod go above 7000. One unusual day pulled the mean upwards.

The median

Put the values in ascending order and take the one in the middle. Dilshod’s steps in order: 4000, 4500, 5000, 5000, 5500, 6000, 21000. The middle value, the fourth one, is 5000. The median describes his typical day far better. If there is an even number of values, take the mean of the two in the middle.

The mode

The value that occurs most often. An illustrative example: Nilufar, a teacher, marks a test for 10 pupils on a five-point scale: 5, 4, 4, 3, 5, 4, 2, 4, 5, 3. The most frequent mark is 4 (four times). For this class the mode is 4, the mean is 3.9 and the median is 4. The mode is especially useful for categories: to know which school uniform size is needed most often, you want the mode, not the average size.

Which one to choose

  • Values close together with no extreme outliers: the mean works well.
  • A few very large or very small values: the median is fairer. It is not “sensitive” to extreme values.
  • For the question “what occurs most often?”: the mode.

A good habit: look at the mean and the median together. If they differ a lot, the data contains outliers and a single number does not tell the whole story.

Spread: the average does not tell you everything

Two sets of data can have the same mean and still be completely different. An illustrative example: Zarina can get to work on two different bus routes. For five days she noted how long she waited for each, in minutes:

  • Route 1: 9, 10, 10, 11, 10.
  • Route 2: 2, 18, 5, 20, 5.

Both have a mean wait of 10 minutes. But the Route 1 bus arrives at almost the same time every day, while the Route 2 bus comes either straight away or after 20 minutes. If Zarina must not be late for an important meeting, Route 1 is the better choice. This difference is what spread shows.

Range

The simplest measure: subtract the smallest value from the largest. Route 1 has a range of 11 - 9 = 2 minutes; Route 2 has 20 - 2 = 18 minutes. The drawback is that the range only looks at the two extreme values, so one unusual day can make it much larger.

The idea of standard deviation

The standard deviation shows how far values typically sit from the mean. You do not need to memorise the formula, because Excel works it out for you. The idea is this: take each value’s difference from the mean, square it, average those squares and take the square root of the result. In Zarina’s data the standard deviation is about 0.7 minutes for Route 1 and about 8.3 minutes for Route 2. A small number means a consistent result; a large one means an unpredictable one.

The everyday takeaway: when you quote an average, quote the spread too if you can. “I sleep 7 hours on average” and “I sleep between 6.5 and 7.5 hours every night” are two different levels of precision.

Percentages and percentage points: a small word, a big difference

These two ideas are confused in the news more often than any others.

A percentage point is the simple difference between two percentages. A percentage change is the change relative to the starting value.

An illustrative example: the share of pupils in a class getting top marks was 20% in the first term and 30% in the second.

  • The share rose by 10 percentage points (30 - 20 = 10).
  • The share rose by 50% (10 divided by 20 is a half).

Both statements are true, but they leave different impressions. Anyone who wants to grab attention usually picks the figure that looks bigger. So when you read “up 50%”, ask: from what to what?

Percentages of a small base

“Doubled” or “up 100%” sounds very impressive. But if a neighbourhood library signs up two new readers this month instead of one, that is also “100% growth”. Whenever you see a percentage, ask for the absolute numbers too: from how many to how many?

Percentages with no base

When you see “40% more”, ask: more than what? Than last year, than another product, than a competitor? A percentage with no stated basis for comparison means almost nothing.

Samples, representativeness and the danger of small samples

It is often impossible to ask everyone, so only some people are asked. The whole group we want to learn about is the population; the part we actually asked is the sample. For a dataset built from a sample to be useful, the sample has to resemble the population, in other words be representative.

An unrepresentative sample

An illustrative example: a school’s parents’ committee wants to know how many hours pupils sleep. The survey is posted only in one class’s Telegram group at 11 pm, and 15 people reply. The problems:

  • The people who replied are the most active parents in the group. Less active ones did not take part at all.
  • Those who replied were still awake at that late hour, which may by chance raise the share of “late to bed” families.
  • The result describes not the whole school but “parents who read the group in the evening and felt like replying”.

Another common situation is self-selection: people are more likely to answer a voluntary survey, especially if the topic made them very happy or very annoyed. That is why online reviews and polls often under-represent “middle” opinions.

The danger of small samples

The smaller the sample, the more the result depends on chance. An illustrative example: Sardor tried a new sleep app for three nights and slept well on two of them. Concluding that “the app helps 67% of the time” would be laughable: over three nights, the weather, how tired he was that day or simple luck could all have played a part.

Signs of a small sample:

  • “Neat” percentages with no absolute numbers (“80% of participants”: is that 4 people out of 5?).
  • A very striking result that does not match other sources.
  • Observations over a few days or a few people presented as a “pattern”.

Correlation is not causation

Correlation means two measures change together: when one goes up, the other goes up (or down) too. But that does not mean one causes the other.

An illustrative example: for a month, Malika recorded her steps and her hours of sleep. She noticed that on days she walked a lot, she slept longer. Does that prove “walking makes you sleep longer”? Not yet. There are other explanations:

  1. A third factor. Most of her high-step days fell at the weekend. At the weekend she did not have to get up early, so she slept longer. Both measures were actually being driven by the “weekend” factor.
  2. Reverse causation. After a good night’s sleep she felt more energetic the next day and walked more.
  3. Chance. With one month of data, two measures can move in the same direction purely by coincidence.

A classic example: in summer more ice cream is sold and there are more cases of sunburn. Ice cream does not cause sunburn; hot weather causes both.

Establishing a cause usually needs an experiment: for example, one of two groups changes something and the other does not, while all other conditions are kept as similar as possible. When you see a headline saying “X leads to Y”, ask: is this an experiment or just an observed link?

Misleading charts

A chart helps you grasp numbers quickly, but it is also easy to build one badly, sometimes by accident and sometimes on purpose.

The truncated axis

An illustrative example: class 7A has an average mark of 3.9 and class 7B has 4.1. If the vertical axis of a column chart starts at 0, the two columns look almost the same height, which reflects reality. If the axis starts at 3.8, the 7A column is drawn 0.1 units tall and the 7B column 0.3 units tall. As a result, 7B looks three times better, even though the difference is only 0.2 marks.

The rule: on a column chart the vertical axis must start at zero, because the height of the column represents the value itself. On a line chart, starting the axis above zero is sometimes reasonable, as it shows the direction of change. But in that case the axis values must be clearly labelled.

Other tricks

  • Uneven time intervals. The horizontal axis shows January, February and then suddenly June; the gaps look equal but are not.
  • A cherry-picked period. Only the convenient months are shown to suggest growth, and the rest are “cut off”.
  • 3D pie charts. The slice at the front looks bigger than the one at the back.
  • No axis labels or units. If you do not know what was measured and in what units, the chart only creates an impression.
  • Dual-axis charts. When two measures are drawn on one chart against different axes, it is easy to suggest a link that does not exist.

To learn more about building charts and preparing data, read our article on data analysis.

Calculating it in Excel and Google Sheets

You can calculate all of the measures above in a spreadsheet in seconds. If you are not yet comfortable with spreadsheets, start with our article on Excel and Google Sheets basics.

Suppose cells B2:B8 contain Dilshod’s step counts for seven days.

  • =AVERAGE(B2:B8): the mean.
  • =MEDIAN(B2:B8): the median. There is no need to sort the values first.
  • =MODE.SNGL(B2:B8): the mode. If no value repeats, it returns the #N/A error. In Google Sheets =MODE(B2:B8) also works.
  • =MIN(B2:B8) and =MAX(B2:B8): the smallest and largest values.
  • =MAX(B2:B8)-MIN(B2:B8): the range.
  • =STDEV.S(B2:B8): the standard deviation for a sample.

Note: depending on your computer’s settings, the separator between formula arguments may be a comma or a semicolon. In some language versions the function names are translated too.

To check a chart’s axis:

  1. Create a column chart: select the data, then Insert → Chart.
  2. In Excel, right-click the vertical axis and open Format Axis → Axis Options → Bounds → Minimum. It should be 0.
  3. In Google Sheets, double-click the chart and check the Min. value under Customise → Vertical axis.

Questions for reading the numbers in the news critically

Next time you see a number in a headline, ask these questions:

  1. Who measured it, and why? Did a party with an interest in the result carry out the study itself?
  2. Where is the original source? Which report or study does the number come from, and is there a link to it?
  3. Who was asked, and how many? Does the sample resemble the population, and is it large enough?
  4. Which “average”? The mean or the median? Is anything said about the spread?
  5. A percentage of what? Percentage points or a relative change? What are the absolute numbers?
  6. Which period? Why was this particular period chosen? What happened before and after it?
  7. Is a link being presented as a cause? Could there be a third factor or reverse causation?
  8. Is the chart honest? Where does the axis start, are the units labelled, are the intervals equal?
  9. What do other sources say? Does the result agree with other reliable sources?

You do not have to answer every question. But if two or three of them have no answer, that is reason enough to treat the number with caution.

Practice exercise and checklist

Try out what you have learnt on your own data. This exercise uses a personal sleep diary, which stays with you alone.

  1. For two weeks, write down how many hours you slept each night (for example, 7.5 or 6). If you like, add your step count too.
  2. Open a spreadsheet: put the date in column A, hours of sleep in column B, steps in column C and the type of day (“work” or “weekend”) in column D.
  3. Calculate AVERAGE, MEDIAN, MIN, MAX and STDEV.S for your hours of sleep. Are the mean and the median close?
  4. Calculate the range. Think back to what happened on your shortest and longest nights.
  5. Calculate the mean separately for workdays and weekend days (with a filter or the AVERAGEIF function). Is there a difference?
  6. Create a column chart of your hours of sleep. Then copy it and change the minimum of the vertical axis to a value just below the mean. Compare the two charts: which one makes the difference look bigger?
  7. If you can see a link between steps and sleep, find a third factor that could explain it.
  8. Write a three-sentence conclusion: how much you usually sleep, how consistent it is and what it might depend on. Mention that this is a small, two-week sample.

Checklist before you use a number

  • It is clear which “average” is used, and outliers have been checked.
  • The spread (range or standard deviation) is shown alongside the average.
  • Absolute numbers and the basis for comparison are given with any percentage; percentages and percentage points are not mixed up.
  • Who is in the sample and how big it is are stated.
  • A link is not claimed as a cause.
  • On column charts the axis starts at zero, and axes and units are labelled.
  • The source is given, and illustrative examples are marked as illustrative.

Statistics teaches you neither to fear numbers nor to trust them blindly. This skill comes in handy whether you are reading the news, preparing a work report or making sense of your child’s marks.

If you would like to learn to work with data in more depth, take a look at our association’s free Data Analytics and Data Science programmes, along with our other training programmes. When you are ready, submit an application and our specialists will get in touch.

Back to articles

More articles

14 min read

Using AI tools responsibly at work

What AI chat assistants are good for at work, how to write a good prompt, how to check the output and which data never to type in.

15 min read

Algorithmic thinking: the step before coding

Breaking problems down, patterns and abstraction, conditions and loops, pseudocode and flowcharts, simple tasks and tracing by hand, plus practical exercises.

Start learning today

Enrollment is open. Leave an application — our specialists will contact you and help you choose the right field.

Message us on Telegram