Math
statisticsmedium~20 min

Statistics, Spread, and Reading Data

You will be able to say what a random sample and a margin of error do and do not let you conclude, tell an observational study from an experiment and know which one supports a cause-and-effect claim, read a scatterplot and interpret the slope and a residual of a line of best fit in context, pull conditional probabilities out of a two-way table by choosing the right denominator, and know why a margin of error, a fitted line, and a table each answer only the question they were built for.

Introduction

Statistics questions on these tests look easier than they are, because they rarely ask for heavy computation. Instead they ask what a number means. Given two classes with identical averages, which one had more consistent results? Given a survey of 400 students, what can you honestly claim about the 3,000 students who were not surveyed? Given a line drawn through a cloud of dots, what does its slope say in the units of the problem? These are reasoning questions dressed as arithmetic.

On the SAT this material lives in the Problem-Solving and Data Analysis domain, which is a large share of the math section and is deliberately built around tables, scatterplots, and short survey descriptions. The ACT spreads the same content across its statistics and probability questions and adds more direct calculation, such as a probability from a table. Both tests reward students who slow down for one second and identify what the denominator is supposed to be, or who was actually sampled.

The good news is that the standard deviation is never something you have to calculate on these tests. You only ever have to compare spreads qualitatively, and you can do that by eye. Likewise, no one asks you to derive a margin of error; you only need to know what it describes. The previous lesson covered the mechanics of mean, median, and how an outlier moves them. This lesson takes the next step: what a sample lets you say about a population, what a fitted line lets you predict, and how to read a table without picking the wrong total.

Game plan

How to attack these questions on test day.
  1. 1

    In a two-way table, find the condition first and circle its total

    The words given that, among, and of those who tell you which row or column you are trapped in. Before you write any fraction, find that row or column total in the margin of the table and write it as the denominator. Only then look for the numerator, and count it inside that same row or column. If the question has no conditioning phrase at all, the denominator is the grand total in the bottom corner.

  2. 2

    Turn a margin of error into an interval, then eliminate by wording

    Subtract and add the margin to the estimate to get the interval, for example 62 plus or minus 4 gives 58 to 66 percent. Then read the answer choices for three tells. A choice that says exactly is wrong, because a margin of error exists precisely because the estimate is not exact. A choice about the sample is wrong, because the interval is a claim about the population. A choice about a bigger population than the one sampled is wrong, because a random sample of one school says nothing about the state. Also remember that a larger sample shrinks the margin; it never grows it.

  3. 3

    Ask who could have been chosen before you accept a conclusion

    A sample supports a claim about a population only if every member of that population had a chance to be selected. Volunteers who reply to an online poll, the first 50 people through a door, or everyone in one classroom are not random samples of a school or a city, so results from them cannot be generalized. Separately, ask whether the treatments were assigned at random. Only random assignment supports cause and effect; a survey, however large, can show association only.

  4. 4

    Write the slope of a line of best fit with both units attached

    A slope of 5.36 on a scatterplot of hours studied against score means about 5.36 points per hour, so each additional hour is associated with a predicted increase of about 5.4 points. Say predicted and associated, because the line is a summary, not a law. A prediction is the y-value on the line; a residual is the actual y-value minus the predicted one. When a question asks how far a data point is from the line, that subtraction is the whole problem.

Theory

Sampling determines what you are allowed to conclude. A sample selected at random from a population is representative of that population, so results from the sample can be generalized to it, but only to it. A random sample of students at one school supports a claim about that school and says nothing reliable about students statewide. A sample that is not random, such as people who volunteer to answer, the first 50 shoppers to enter a store, or every student in one class, is biased toward whatever kind of person was easy to reach, and no conclusion about the wider population is justified no matter how many people were surveyed. When a question describes how a sample was chosen, the first thing to identify is the population that every sampled individual was drawn from, because that is the largest group the result can describe.

A margin of error attached to a survey result gives a plausible interval around the estimate. If 62 percent of a random sample of 400 students support a later start time with a margin of error of 4 percentage points, then the true proportion of all students at that school plausibly lies between 58 and 66 percent. The margin does not mean that exactly 62 percent support it, and it is not a statement about the 400 students in the sample, whose answers are already known. Larger random samples produce smaller margins of error, because more data pin the estimate down more tightly; quadrupling the sample roughly halves the margin. You will never be asked to compute a margin of error on these tests, only to interpret one and to know which direction it moves.

Whether a study can support a cause-and-effect claim depends on a different question: were the treatments assigned at random? In an observational study, the subjects already belong to their groups, so some other difference between the groups may explain the outcome; a survey showing that people who drink coffee sleep less cannot say that coffee causes the lost sleep. In an experiment, the researcher randomly assigns subjects to treatments, which spreads every other difference evenly across the groups, so a difference in outcomes can be attributed to the treatment. The two questions are independent. A random sample with random assignment supports both generalization and causation; a random sample without random assignment supports generalization only; random assignment among volunteers supports causation for subjects like those volunteers but not generalization to a population.

A scatterplot shows paired data, one dot per individual, with the explanatory variable on the horizontal axis and the response on the vertical axis. If the dots trend upward the association is positive, if downward it is negative, and if the dots hug a straight path the association is strong. The line of best fit is the single line that comes closest to the whole cloud of points, and its slope is the predicted change in the response for each one-unit increase in the explanatory variable, carrying the units of both axes. Its y-intercept is the predicted response at an input of zero, which is meaningful only if zero is inside or near the range of the data. For any data point, the residual is the actual y-value minus the y-value the line predicts, so a point above the line has a positive residual and a point below it a negative one. Two cautions are tested constantly: predicting far outside the observed range of x is extrapolation and is unreliable, and a strong association from observational data does not establish that one variable causes the other.

A two-way table sorts individuals by two categorical variables at once, with the row totals, the column totals, and the grand total along the edges. Every probability from such a table is a count divided by a total, and the only real skill is picking the right total. An unconditional probability, such as the probability that a randomly chosen student passed, uses the grand total as the denominator. A conditional probability, signaled by the words given that, among, or of those who, restricts attention to one row or one column, so the denominator becomes that row or column total. Reversing the condition changes the answer: the probability that a student passed given that she studied in a group and the probability that a student studied in a group given that she passed use the same numerator but different denominators, and confusing them is one of the most common errors on the whole test.

Spread describes how far the values of a distribution sit from its center, and it is a separate fact from the center itself. Two data sets can share a mean, a median, a range, and a count and still differ in standard deviation, so an average alone never describes a distribution. These tests never ask you to compute a standard deviation, only to compare two of them, and the previous lesson on mean, median, and standard deviation covers how to make that comparison by looking at how far the values sit from the mean.

Worked figures

Two independent questions decide what a study can claim

Random selection of subjects is what lets a result be generalized to a population. Random assignment of treatments is what lets a difference in outcomes be blamed on the treatment. A study can have either, both, or neither, and each combination supports a different conclusion.

Table 1
Subjects were selectedTreatments were assignedGeneralize to the population?Conclude cause and effect?
At randomAt randomYesYes
At randomNot at random (observational)YesNo, association only
Not at random (volunteers)At randomNo, only to subjects like theseYes, for those subjects
Not at randomNot at randomNoNo

A larger random sample shrinks the margin of error

The same survey question, with 62 percent support each time, produces a narrower plausible interval as the random sample grows. Quadrupling the sample roughly halves the margin. You are never asked to compute these values, only to know that the margin moves in this direction.

Table 2
Random sample sizeApproximate margin of errorPlausible interval around 62 percent
1008 percentage points54 to 70 percent
4004 percentage points58 to 66 percent
1,6002 percentage points60 to 64 percent
6,4001 percentage point61 to 63 percent

A scatterplot and its line of best fit

The eight dots show a strong positive association between hours studied and score, and the line is the line of best fit, approximately y = 5.36x + 46.6. Its slope says each extra hour is associated with about 5.4 more predicted points. The dots sit above and below the line, which is expected: the line summarizes the trend rather than passing through every point, and the vertical gap at each dot is its residual.

Figure 1
405060708090100012345678910(4, 70)(8, 90)hours studiedtest score
  • line of best fit: y = 5.36x + 46.6

Two-way table: the denominator is the whole question

Among the 60 students who studied with a group, 42 passed, so that conditional probability is 42/60 = 0.7. Among the 70 students who passed, 42 studied with a group, so that conditional probability is 42/70 = 0.6. Same cell, different condition, different answer.

Table 3
Study methodPassedDid not passTotal
Studied with a group421860
Studied alone283260
Total7050120

Worked examples

Try each one before opening the solution.

Example 1: Two conditional probabilities from one table

In a group of 120 students, 60 studied with a group and 60 studied alone. Of those who studied with a group, 42 passed; of those who studied alone, 28 passed. Find the probability that a student passed, given that the student studied with a group. Then find the probability that a student studied with a group, given that the student passed.

Study method and exam result for 120 students
Study methodPassedDid not passTotal
Studied with a group421860
Studied alone283260
Total7050120
Show solution
  1. Read the first condition: given that the student studied with a group. That restricts you to the group-study row, whose total is 60.

    The condition fixes the denominator before you look at anything else. Underline it.

  2. Count the numerator inside that row: 42 of the group-study students passed.
  3. Form the fraction: 42/60 = 7/10 = 0.7.

    Both 42 and 60 are divisible by 6, which gives 7/10.

  4. Read the second condition: given that the student passed. That restricts you to the passed column, whose total is 42 + 28 = 70.

    The condition has flipped from a row to a column, so the denominator changes even though the question sounds almost identical.

  5. Count the numerator inside that column: 42 of the students who passed had studied with a group.
  6. Form the fraction: 42/70 = 3/5 = 0.6.

    The numerator was 42 both times. Only the denominator changed, and that is where the whole question lives.

Answer: The probability of passing given group study is 42/60 = 0.7, and the probability of group study given passing is 42/70 = 0.6.

Example 2: Interpret the slope of a line of best fit and find a residual

For eight students, the hours spent studying and the score earned produced a line of best fit of approximately y = 5.36x + 46.6, where x is hours studied and y is the score. Interpret the slope in context, predict the score for a student who studies 5 hours, and find the residual for the student in the data who studied 5 hours and scored 72.

y = 5.36x + 46.6 with the prediction at x = 5
405060708090100012345678910x = 5actual (5, 72)predicted (5, 73.4)hours studiedtest score
  • y = 5.36x + 46.6
Show solution
  1. Read the slope with both units: 5.36 score points per hour of study.

    Slope is always a change in y per one unit of x, so it carries the y-unit over the x-unit.

  2. State the interpretation: each additional hour of study is associated with a predicted increase of about 5.4 points in the score.

    Say predicted and associated. The data are observational, so the line does not show that studying causes the higher score.

  3. Predict for x = 5: multiply first, 5.36 times 5 = 26.8.
  4. Add the intercept: 26.8 + 46.6 = 73.4, so the predicted score is about 73.
  5. Find the residual as actual minus predicted: 72 - 73.4 = -1.4.

    A negative residual means the point sits below the line. A residual of this size is a normal gap between a summary line and one data point, not an error in the model.

  6. Check the limits of the prediction: for x = 40 hours the line gives 5.36 times 40 + 46.6 = 214.4 + 46.6 = 261, an impossible score, because 40 hours is far outside the observed range of 1 to 8 hours.

    This is extrapolation. The line describes the data it was fit to and nothing beyond it.

Answer: Each extra hour is associated with about 5.4 more predicted points; a student who studies 5 hours is predicted to score about 73.4; the actual score of 72 gives a residual of about -1.4.

Example 3: Read a poll result with its margin of error

A polling firm selects 400 residents of a city at random and finds that 55 percent of them favor building a new park, with a margin of error of 3 percentage points. State the interval the poll supports, name the population it applies to, and say what would happen to the margin of error if 1,600 residents had been polled instead.

Show solution
  1. Find the lower end of the interval: 55 - 3 = 52 percent.
  2. Find the upper end: 55 + 3 = 58 percent. So it is plausible that between 52 and 58 percent favor the park.

    This is a statement about the true proportion, which the poll estimates, not about the 400 people already asked.

  3. Identify the population: the sample was drawn at random from residents of this city, so the interval describes residents of this city.

    Not residents of the state, and not residents of other cities. A random sample generalizes only to the population it was drawn from.

  4. Compare sample sizes: 1,600 is 4 times 400, so the margin of error would shrink to roughly half, about 1.5 percentage points.

    The test only asks the direction: more data, smaller margin. The exact size of the shrink is never required.

  5. Rule out the overreach: the poll does not show that exactly 55 percent favor the park, and it cannot say why anyone favors it, because nothing was assigned or changed by the pollster.

Answer: It is plausible that between 52 and 58 percent of the city's residents favor the park; the result applies to this city only; polling 1,600 residents would cut the margin of error to about 1.5 percentage points.

Practice

Check your understanding 1

A survey of 120 students recorded how each studied and whether each passed an exam. Of the 60 students who studied with a group, 42 passed and 18 did not. Of the 60 students who studied alone, 28 passed and 32 did not, so 70 students passed in all. If one of the students who studied alone is selected at random, what is the probability that the student passed?

Check your understanding 2

A researcher surveys 400 randomly selected students at a large high school and finds that 62 percent support a later start time, with a margin of error of 4 percentage points at a 95 percent confidence level. Which conclusion is best supported by these results?

Check your understanding 3

A researcher wants to estimate the mean number of hours per week that students at a university spend at paid jobs. Which of the following sampling methods would best allow the result to be generalized to all students at the university?

Check your understanding 4

A scatterplot shows the age x, in years, and resale value y, in thousands of dollars, of 20 used cars of the same model. The line of best fit is y = -2.4x + 31. Which of the following is the best interpretation of the number -2.4 in this context?

Common mistakes

Using the grand total for a conditional probability. When a question says given that, among, or of those who, the denominator is a row total or a column total, not the total number of individuals. Before you write the fraction, underline the conditioning phrase and find its total in the table margin, then count the numerator only within that row or column.

Swapping the direction of a conditional. The probability of passing given group study and the probability of group study given passing are different numbers built from the same cell. Read the condition first, fix the denominator, and only then look for the numerator; doing it in that order makes the swap almost impossible.

Assuming equal means imply equal distributions. Two data sets with the same mean can have wildly different standard deviations, and a question that gives you both means and asks about consistency, reliability, or variability is asking about spread, not center. Look at how tightly the values cluster around the mean rather than at the mean itself, and do not let a shared range convince you the spreads match.

Treating a strong association as proof of causation. A scatterplot showing that towns with more libraries have higher test scores does not show that libraries raise scores; the data are observational and some third factor, such as town wealth, may drive both. Only a randomized experiment supports a causal claim, and only a random sample supports generalization to a population.

Trusting a sample that chose itself. Volunteers who answer an online poll, the first people through a door, and everyone in one class are all easy to reach and all unrepresentative, and a large count of them does not repair the bias. When a question describes the sampling method, decide whether every member of the population could have been selected before you accept any generalization.

Overstating what a margin of error says. A result of 62 percent plus or minus 4 points does not mean exactly 62 percent, and it does not mean between 58 and 66 percent of the sample. It means the true value for the population from which the sample was randomly drawn plausibly lies between 58 and 66 percent. Also remember the direction of the sample-size effect: increasing the sample size shrinks the margin of error, it does not grow it.

Predicting outside the data. A line of best fit fit to 1 to 8 hours of study has nothing to say about 40 hours, and a question that asks for such a prediction is usually testing whether you recognize extrapolation. Compare the requested input to the range of x-values in the data before you plug it in.