Math
statisticsmedium~20 min

Mean, Median, and Standard Deviation

You will be able to compute the mean, median, mode, and range of a list quickly and in the right order, predict how an outlier or a change to the data moves each statistic, decide which of two data sets has the larger standard deviation without computing anything, and work backwards from a target mean to the score that produces it.

Introduction

A data set is a list of numbers, and a statistic is one number that summarizes the list. The mean and the median both describe where the center of the data sits, but they answer that question in different ways: the mean balances the values, the median splits them in half. The range and the standard deviation describe how spread out the values are. Almost every question in this topic is a question about which of these numbers moves, and by how much, when the data change.

On the digital SAT these questions live in the Problem-Solving and Data Analysis domain and usually come with a short list, a frequency table, or a pair of dot plots. You will be asked to find a median from an even-length list, to find the score needed on one more test to reach a target average, or to say which of two distributions has the greater standard deviation. The ACT asks the same things with a little more arithmetic, including means from frequency tables and missing values in a list with a known average.

The standard deviation deserves one reassurance up front: you will never be asked to compute it. The tests only ever ask you to compare two spreads, and that is a judgment you can make by looking at how far the values sit from the mean. The computation that does matter is the one behind every mean question, namely that a mean times a count is a sum. Once you can move between means and totals without hesitation, the rest of the topic is bookkeeping.

Game plan

How to attack these questions on test day.
  1. 1

    Convert every mean into a total before doing anything else

    The moment a question mentions a mean, write sum = mean times count. A mean of 85 over five tests is a total of 425; a group of 12 people with a mean age of 30 has a combined age of 360. Missing-value questions, needed-score questions, and combined-group questions all become one subtraction or one division once the totals are on the page, and you never fall into the trap of averaging two averages.

  2. 2

    Sort, count, then locate the median by position

    Rewrite the list in order, count the values, and use position (n + 1) / 2. For 9 values that is the 5th; for 10 values it is halfway between the 5th and 6th, so average them. Cross values off from both ends in pairs if the list is long. When the data come as a frequency table or a dot plot rather than a list, add the frequencies from the left until the running count reaches the middle position, and that category holds the median.

  3. 3

    Compare standard deviations by distance from the mean, not by range

    When two dot plots or lists are shown and you are asked which has the larger standard deviation, find each mean by eye, then ask where the values sit relative to it. A set with most of its values stacked at the center has the smaller standard deviation; a set with values pushed out toward both ends has the larger one, even if the two ranges are identical. Never start computing squared deviations. The test does not require it, and the arithmetic burns a minute you need elsewhere.

  4. 4

    Ask which statistics move and which stay put

    For any change to a data set, run through mean, median, range, and standard deviation and decide the effect on each one. Adding a constant to every value shifts the center but not the spread. Altering or removing the largest value moves the mean and the range but rarely the median. Adding a value equal to the mean leaves the mean alone. Questions of the form which of the following must be true are built from exactly these facts, and the wrong choices are the ones that claim a statistic moves when it does not.

  5. 5

    On the digital SAT, let the calculator do the arithmetic and you do the reasoning

    The built-in calculator will find the mean or median of any list you type into it, so use it to check a hand computation or to handle a long, ugly list. It cannot decide for you which measure a question wants, which set has the larger spread, or what total a target mean requires. Those decisions are what the question is really testing, so make them before you reach for the keyboard.

Theory

The mean of a data set is its sum divided by its count. Think of it as a balance point: if you laid the values out on a number line as equal weights, the mean is where the line would balance, which is why every value pulls on it in proportion to how far away it sits. The mean carries one algebraic fact that matters more than any other on these tests: mean times count equals sum. Any question that gives you a mean and a count is really handing you the total, and any question that asks for a mean is really asking for a total divided by a count. Rewriting the problem in terms of the sum is the first move in almost every mean question.

The median is the middle value once the data are written in order from least to greatest, and ordering is not optional. If the count is odd there is a single middle value; with 7 values it is the 4th. If the count is even there are two middle values and the median is their average; with 8 values it is the average of the 4th and 5th, and it may be a number that does not appear in the list at all. In general the median sits at position (n + 1) / 2, so for n = 6 that is position 3.5, meaning halfway between the 3rd and 4th values. Two more summary numbers travel with these. The mode is the value that occurs most often, and a set can have no mode, one mode, or several. The range is the largest value minus the smallest, so it depends on only the two extreme values and ignores everything in between.

The mean and the median respond differently to an outlier, a value far from the rest. Because the mean is a balance point, an extreme value drags it toward that extreme; because the median depends only on position in the ordered list, an extreme value counts as one more value on one side and nothing more. In the data set 12, 14, 15, 15, 16, 18 both the mean and the median are 15. Replace the 18 with 60 and the mean jumps to 22 while the median stays at 15, because 60 is still simply the largest value. That is the whole story behind a rule the tests check constantly: a high outlier pushes the mean above the median, a low outlier pulls the mean below the median, and in a roughly symmetric data set the two agree. When a question asks which statistic would change if the largest value were doubled, the answer is the mean and the range; the median and usually the mode do not move.

The standard deviation measures spread. Informally, it is the typical distance of a value from the mean, computed from every value in the set. A set whose values cluster tightly around the mean has a small standard deviation, a set whose values scatter far from it has a large one, and a set whose values are all identical has a standard deviation of exactly zero. Neither the SAT nor the ACT asks you to compute it. They ask you to compare two data sets, usually shown as dot plots, histograms, or short lists, and to say which has the larger standard deviation. To do that, locate the mean of each set by eye and ask how far the values sit from it. Values piled up at the mean lower the standard deviation; values piled up at the extremes raise it. The size of the mean itself is irrelevant to the comparison, and so is the count. Two sets with the same range can have different standard deviations if one keeps most of its values near the center while the other pushes most of them to the ends.

Changing the data changes the statistics in predictable ways, and questions test whether you know which ones move. Adding the same constant to every value slides the whole set along the number line: the mean, median, and mode all increase by that constant, while the range and the standard deviation do not change at all, because every value is exactly as far from every other value as it was before. Multiplying every value by a positive constant k multiplies the mean, the median, the range, and the standard deviation all by k. Adding a new value above the current mean raises the mean, adding one below it lowers the mean, and adding a value equal to the mean leaves the mean alone while slightly shrinking the standard deviation, since the set now has one more value at distance zero. Removing an outlier moves the mean toward the remaining values, shrinks the range and the standard deviation sharply, and usually leaves the median close to where it was.

Working backwards from a mean is the most common calculation in this topic. If four test scores of 82, 90, 77, and 85 need to average 85 after a fifth test, the fifth score x satisfies 82+90+77+85+x5=85\frac{82 + 90 + 77 + 85 + x}{5} = 85. Do not try to average the averages. Multiply the target mean by the final count to get the required total, 85 times 5 = 425, subtract the total you already have, 334, and the difference, 91, is the needed score. The same total-first thinking handles a missing value in any list and handles combined groups: the mean of two groups combined is the combined sum divided by the combined count, which equals the average of the two means only when the groups are the same size or the two means happen to be equal.

Worked figures

Four summaries of a data set, before and after an outlier

Set B is Set A with its largest value, 18, replaced by 60. The sum rises by 42, so the mean rises by 42 / 6 = 7, from 15 to 22. The median and the mode do not move, because 60 is still simply the largest value and 15 still occurs most often. The range, which depends only on the two extremes, jumps from 6 to 48.

Table 1
StatisticSet A: 12, 14, 15, 15, 16, 18Set B: 12, 14, 15, 15, 16, 60
Sum90132
Mean (sum / 6)90 / 6 = 15132 / 6 = 22
Median (average of 3rd and 4th)(15 + 15) / 2 = 15(15 + 15) / 2 = 15
Mode1515
Range (largest - smallest)18 - 12 = 660 - 12 = 48

Where the mean and the median sit when one value is extreme

The six values of Set B are plotted in order. Five of the six sit below the mean line at 22, and only the outlier sits above it. The five values below the mean are a total of 10 + 8 + 7 + 7 + 6 = 38 below it and the outlier is 38 above it, which is exactly what makes 22 the balance point. The median stays on the dashed line at 15, the same place it was before the 18 became a 60, because the 3rd and 4th values in order are still 15 and 15.

Figure 1
01020304050607001234567median = 15 (also the mean of Set A)1214151516outlier 60position in the ordered listvalue
  • mean of Set B = 22

Same mean, same median, different standard deviation

Both classes have 10 students, a mean score of 80, and a median score of 80. In Class P six of the ten scores are exactly at the mean and none is more than 10 points away; in Class Q six of the ten scores are 20 points from the mean. Class Q has the larger standard deviation. Notice that its range, 100 - 60 = 40, is also larger than Class P's range of 90 - 70 = 20, but the bars show something the range cannot: where the values pile up.

Figure 2
012345660708090100test scorenumber of students
  • Class P (clustered at the mean, mean 80)
  • Class Q (pushed to the ends, mean 80)

What moves when the data change

Adding a constant slides every value the same distance, so the center shifts and the spread does not. Multiplying stretches everything, including the spread. Adding or removing a single value changes the mean and the spread in the direction you would expect, while the median, which only counts positions, moves by at most one step in the ordered list.

Table 2
Change made to the data setMeanMedianRangeStandard deviation
Add the same constant c to every valueincreases by cincreases by cunchangedunchanged
Multiply every value by a positive constant kmultiplied by kmultiplied by kmultiplied by kmultiplied by k
Add one new value equal to the meanunchangedunchanged or moves toward the meanunchangeddecreases slightly
Add one new value far above the largestincreasesunchanged or moves up one positionincreasesincreases
Remove a single high outlierdecreasesunchanged or moves down one positiondecreasesdecreases

Worked examples

Try each one before opening the solution.

Example 1: Mean and median of an even-length list

A student records the number of pages read on each of eight days: 7, 12, 9, 15, 9, 20, 11, 13. Find the mean and the median of the data set.

Show solution
  1. Add the values: 7 + 12 + 9 + 15 + 9 + 20 + 11 + 13 = 96.

    Pair values that make round numbers to avoid slips: 7 + 13 = 20, 9 + 11 = 20, and 12 + 9 + 15 + 20 = 56, so the total is 20 + 20 + 56 = 96.

  2. Divide by the count: mean = 96 / 8 = 12.
  3. Sort the list from least to greatest: 7, 9, 9, 11, 12, 13, 15, 20.

    The median is a position, so the order is not optional. Skipping this step is the single most common median error.

  4. With 8 values the middle is position (8 + 1) / 2 = 4.5, so the median is the average of the 4th and 5th values, which are 11 and 12.
  5. Average them: median = (11 + 12) / 2 = 11.5.

    The median does not have to be a value that appears in the list.

  6. Sanity check: the mean, 12, is slightly above the median, 11.5, which fits a list whose largest value, 20, sits farther above the center than the smallest value, 7, sits below it.

Answer: The mean is 12 and the median is 11.5.

Example 2: Find the score needed for a target mean

Marcus scored 82, 90, 77, and 85 on four tests. What score on a fifth test would give him a mean of exactly 85 over all five tests?

Show solution
  1. Translate the target mean into a target total: five tests with a mean of 85 must add up to 85 times 5 = 425.

    Mean times count equals sum. Everything else in the problem is bookkeeping.

  2. Add the scores he already has: 82 + 90 + 77 + 85 = 334.
  3. Subtract to find what is missing: 425 - 334 = 91.
  4. Check: (334 + 91) / 5 = 425 / 5 = 85.

    His current mean is 334 / 4 = 83.5, so he must score well above 85 on the last test to pull the average up. Scoring exactly 85 would leave the mean at 419 / 5 = 83.8.

Answer: He needs a score of 91 on the fifth test.

Example 3: Decide which set has the larger standard deviation without computing it

Set P consists of the values 40, 48, 50, 52, 60. Set Q consists of the values 40, 40, 50, 60, 60. Which set has the larger standard deviation?

Distances from the mean
SetValuesMeanDistance of each value from the meanRange
P40, 48, 50, 52, 605010, 2, 0, 2, 1020
Q40, 40, 50, 60, 605010, 10, 0, 10, 1020
Show solution
  1. Find each mean. Set P: (40 + 48 + 50 + 52 + 60) / 5 = 250 / 5 = 50. Set Q: (40 + 40 + 50 + 60 + 60) / 5 = 250 / 5 = 50.

    Same mean, same median of 50, same range of 20, same count of 5. None of those decides the question, which is the point.

  2. List each value's distance from the mean for Set P: 10, 2, 0, 2, 10.
  3. List the distances for Set Q: 10, 10, 0, 10, 10.
  4. Compare the two lists. Set P has three values within 2 of the mean and only two far away; Set Q has four values at distance 10. The typical distance from the mean is larger in Set Q.

    The standard deviation is built from exactly these distances, so the set whose values sit farther from the mean on average has the larger standard deviation.

  5. Conclude that Set Q has the larger standard deviation, even though the two sets share a range and a mean.

Answer: Set Q has the larger standard deviation.

Example 4: Remove an outlier and track every statistic

The data set 3, 5, 6, 6, 8, 44 gives the number of minutes each of six customers waited. Find the mean and the median, then find both again after the outlier 44 is removed, and describe what happens to the range and the standard deviation.

Show solution
  1. Mean of the original set: (3 + 5 + 6 + 6 + 8 + 44) / 6 = 72 / 6 = 12.

    Notice that 12 is larger than five of the six values. A mean that exceeds most of the data is the signature of a high outlier.

  2. Median of the original set: the list is already in order, and with 6 values the median is the average of the 3rd and 4th values, (6 + 6) / 2 = 6.
  3. Remove 44. The new sum is 72 - 44 = 28 and the new count is 5, so the new mean is 28 / 5 = 5.6.
  4. New median: with 5 values the median is the 3rd value of 3, 5, 6, 6, 8, which is 6.

    The median did not move at all. Removing the largest value only takes one value off the top of the ordered list.

  5. Range: before, 44 - 3 = 41; after, 8 - 3 = 5.
  6. Standard deviation: before, one value sat 32 away from the mean of 12; after, every value is within 3 of the new mean of 5.6, so the standard deviation drops sharply.

Answer: Before: mean 12, median 6. After removing 44: mean 5.6, median 6. The range falls from 41 to 5 and the standard deviation falls sharply; only the median is unaffected.

Practice

Check your understanding 1

The number of emails a worker received on six days was 5, 12, 8, 3, 12, and 14. What is the median of this data set?

Check your understanding 2

Nadia's mean score on her first three quizzes is 78. What score does she need on the fourth quiz for her mean score over all four quizzes to be 80?

Check your understanding 3

List A consists of the values 2, 5, 5, 5, 8. List B consists of the values 2, 2, 5, 8, 8. Which of the following correctly compares the standard deviations of the two lists?

Check your understanding 4

A data set contains 20 values. Each value is increased by 7 to form a new data set. Which of the following must be true about the new data set compared with the original?

Check your understanding 5

The data set 15, 18, 20, 20, 22, 95 lists the prices, in dollars, of six items. If the item priced at 95 dollars is removed from the data set, which of the following describes the effect on the mean and the median?

Common mistakes

Finding the median without sorting. The median is the middle of the ordered list, not the middle of the list as written. A list such as 5, 12, 8, 3, 12, 14 has a median of 10, not 8 or 12. Rewrite the values in order every time, and count them so you know whether you need one middle value or the average of two.

Averaging averages when working backwards. If three quizzes average 78 and the target over four quizzes is 80, the fourth quiz does not need an 82, and it does not need an 80. The earlier quizzes are each 2 points short, so the shortfall is 6, and the fourth score must be 86. Convert every mean into a total with sum = mean times count, then subtract totals. This also protects you when two groups of different sizes are combined, since the combined mean is not the average of the two means.

Assuming the same range means the same standard deviation. The range depends only on the two extreme values, while the standard deviation depends on how far every value sits from the mean. Two lists can share a mean and a range while one keeps most of its values near the center and the other keeps most of them at the ends; the second has the larger standard deviation. When comparing spread, look at where the values pile up, not at the endpoints.

Expecting the median to move whenever the data change. Doubling the largest value, replacing it with an outlier, or removing it changes the mean and the range, but the median depends only on the middle position of the ordered list and usually stays exactly where it was. The reverse mistake also appears: assuming the mean is unaffected by one extreme value. The mean is a balance point, and one far-off value can move it a long way.

Moving the spread when you shift the data. Adding the same constant to every value raises the mean and the median by that constant and leaves the range and the standard deviation unchanged, because the values are all the same distances apart as before. Only multiplying every value by a constant scales the spread. On a which-must-be-true question, check each statistic separately and ask whether the change moved the values relative to each other or only relative to the number line.