SAT
mathhard~40 min

Advanced: Problem Solving and Data Analysis

Working with the mean as a sum: missing values, removed values and weighted means of combined groups; how outliers and skew move the mean, median and standard deviation; recognizing exponential data by ratios; conditional probability from percents; what random sampling and random assignment each let you conclude; how the margin of error changes with sample size; and unit conversions with squared and cubed units.

Introduction

The hardest data questions on the SAT are not harder arithmetic; they are questions about what a number means and what it licenses you to say. A mean is a sum in disguise, and the test asks you to use it that way. A margin of error is a statement about a population, and the test offers you four conclusions of which three overreach. A study either randomly assigned its treatments or it did not, and only one of those permits the word "causes".

This lesson covers the reasoning behind each. The mean times the count is the sum, which unlocks every missing-value and combined-group question. The shape of a distribution tells you whether the mean or the median is larger. A constant ratio in a table is exponential and a constant difference is linear. A conditional probability is a fraction of a fraction. Random sampling supports generalization; random assignment supports causation; and a larger sample shrinks the margin of error by the square root of the factor.

These are the questions that separate a strong data score from a perfect one, and most of them can be answered without a calculator once the idea is clear.

Game plan

How to attack these questions on test day.
  1. 1

    Turn every mean into a sum: mean times count

    A class of 20 with mean 78 has a total of 1,560 points, and that total is what changes when a value is added or removed. To find a missing value, compute the total the mean requires and subtract what you know. To combine groups, add their totals and divide by the combined count; never average the two means unless the groups are the same size.

  2. 2

    Let the shape tell you which is bigger, mean or median

    In a distribution with a long tail to the right, a few large values pull the mean above the median. A long tail to the left pulls the mean below it. Symmetric distributions have them roughly equal. So when a histogram or dot plot is shown and the question asks which statement about mean and median is true, look at the tail first and compute nothing.

  3. 3

    Test differences for linear and ratios for exponential

    Down a table of equally spaced inputs, subtract consecutive outputs; a constant difference is linear with that slope. If the differences grow, divide consecutive outputs instead; a constant ratio is exponential with that base. The SAT usually gives three or four rows, and it usually offers one linear and one exponential model with the right numbers and two with the wrong ones.

  4. 4

    Read the study design before you read the conclusion

    Was the sample chosen at random from a population? Then results generalize to that population and no further. Were participants randomly assigned to treatments? Then a difference between treatments can be called a cause. If neither, the study shows an association among the people studied, nothing more. The correct answer choice is the one that claims exactly what the design supports, and the wrong ones claim more, or claim it about a wider group.

  5. 5

    Margin of error shrinks with the square root of the sample size

    Quadrupling a sample halves the margin of error; multiplying it by nine cuts the margin to a third. The population size does not matter. When asked which change would most reduce the margin of error, the answer is a larger random sample, and when asked to interpret one, the answer is a range of plausible values for the population parameter, never a statement about the sample itself.

Theory

Because the mean is the sum divided by the count, the sum is the mean times the count, and that turns every mean question into an addition problem. If six numbers have mean 15, their sum is 90; if a seventh is added and the mean becomes 17, the new sum is 119, so the seventh number is 29. If a class of 20 has mean 78 and one score is removed leaving a mean of 80 for the remaining 19, the removed score is 1,560 - 1,520 = 40. Two groups combine by total, not by averaging the means: 25 boys with mean 70 and 15 girls with mean 78 have a combined total of 1,750 + 1,170 = 2,920 over 40 students, a mean of 73, which is a weighted mean pulled toward the larger group. Only when the groups are equal in size does the combined mean equal the average of the means. The same reasoning answers questions about how a mean changes when every value is increased by a constant (the mean rises by that constant) or multiplied by one (the mean is multiplied).

The median depends only on order and the mean on every value, and the difference between them is a statement about shape. In a distribution skewed to the right, with a tail of unusually large values, the mean is greater than the median; skewed to the left, the mean is less. Removing an outlier moves the mean substantially and the median little; the range collapses; the standard deviation falls, because standard deviation measures spread around the mean and the outlier was the farthest point. Comparing two distributions, the one with values more concentrated around its center has the smaller standard deviation regardless of where that center is, and two distributions with identical shapes shifted along the axis have the same standard deviation and different means. A question that says "which measure is most affected by removing the value 100" from a set of otherwise small values is asking you to name the mean, or the range, not the median.

A table of data with equally spaced inputs is linear if the outputs change by a constant difference and exponential if they change by a constant ratio. Outputs 300, 360, 432 at inputs 0, 1, 2 have differences 60 and 72, not constant, but ratios 1.2 and 1.2, so the model is y = 300(1.2)^x, a 20 percent increase per step. Percent problems with an unknown original amount are solved the same way in reverse: if a price rose 20 percent and then fell 25 percent to end at 72 dollars, the original was 72 / (1.20 x 0.75) = 80. A chain of percent changes is one multiplier, and the original is the final value divided by it.

Conditional probability with percents is a fraction of a fraction. If 40 percent of a group own a dog and 25 percent of the group own both a dog and a cat, then among dog owners the fraction owning a cat is 0.25 / 0.40 = 0.625. Formally, the probability of A given B is the probability of both divided by the probability of B. Two events are independent when knowing one does not change the other: the probability of A given B equals the probability of A. In a two-way table, independence shows up as every row having the same proportions; a question may ask whether the data suggest the two variables are associated, and the answer is yes when the conditional proportions differ across rows.

Statistical conclusions are limited by how the data were collected. Random sampling from a population is what allows a sample result to be generalized to that population: a random sample of a city's adults tells you about that city's adults, and a survey of people at a gym tells you about gym-goers only. Random assignment of participants to treatments is what allows a cause-and-effect conclusion, because it makes the treatment groups alike in everything but the treatment. An observational study, where people chose their own group, can show an association but not causation, since the groups may differ in other ways. So a study that randomly samples a school's students and randomly assigns them to two teaching methods can conclude that the method caused the difference for students at that school; the same study without random assignment can conclude only that the method and the outcome are associated; and with neither, only that they were associated among the participants.

A sample statistic estimates a population parameter with a margin of error, and the true value is plausibly anywhere within the margin on either side of the estimate. The margin depends on the sample size, shrinking in proportion to the square root of the size: a sample four times as large has half the margin, a sample nine times as large has a third. It does not depend on the size of the population. The margin says nothing about the individuals sampled, whose values are known exactly, and it does not guarantee that the true value lies in the interval; it identifies the range of values consistent with the sample. Answer choices that say the sample's own value lies in the interval, that the population value is certainly in it, or that a smaller sample would be more precise are the standard wrong answers.

Rates with compound units need every unit converted, and squared or cubed units convert by the square or cube of the factor. One meter is 100 centimeters, so one square meter is 100² = 10,000 square centimeters and one cubic meter is 1,000,000 cubic centimeters. Density is mass per unit volume: an object of volume 4,000 cubic centimeters and density 2.5 grams per cubic centimeter has mass 10,000 grams, which is 10 kilograms. Write every quantity with its units and cancel; if the leftover unit is not the one the question asks for, the conversion is incomplete.

Worked figures

Every mean question is a sum question

Mean times count is the total. Missing values, removed values and combined groups are all found by working with totals, never by averaging means.

Table 1
SituationTotalsAnswer
6 numbers, mean 15; add one, mean becomes 176 x 15 = 90 before, 7 x 17 = 119 afteradded number: 119 - 90 = 29
20 scores, mean 78; remove one, mean of the rest is 8020 x 78 = 1560, 19 x 80 = 1520removed score: 40
25 boys mean 70, 15 girls mean 781750 + 1170 = 2920 over 40combined mean: 73, not 74
Every value increased by 5total rises by 5 x countmean rises by 5; spread unchanged
Every value doubledtotal doublesmean and standard deviation both double

Skew decides which is larger, mean or median

Commute times for 24 workers. Most are short, a few are very long, so the distribution has a right tail: the mean (about 27 minutes) is pulled above the median (about 20). Removing the two longest commutes would drop the mean by several minutes and the median by almost nothing.

Figure 1
024680-1010-2020-3030-4040-5050-6060-7070-80Commute time (minutes)Number of workers
  • Workers

Linear or exponential from a table

Equally spaced inputs. The first output column has constant differences, so it is linear: y = 60x + 300. The second has constant ratios, so it is exponential: y = 300(1.2)^x. Both start at 300 and agree at x = 1; by x = 5 they are 600 and about 746.

Table 2
xLinear yDifferenceExponential yRatio
0300300
1360+60360x1.2
2420+60432x1.2
3480+60518.4x1.2
5600+60 each step746.5x1.2 each step

What a study design lets you conclude

Random selection controls who is in the study and so what population the result describes; random assignment controls who gets which treatment and so whether a difference can be called a cause. The two are independent choices.

Table 3
Random sample from population?Random assignment to treatments?Can generalize to the population?Can conclude cause and effect?
YesYesYesYes
YesNo (observational)Yes: an association holds in the populationNo
NoYesNo: only for the participantsYes, for the participants
NoNoNoNo: an association among the participants only

Margin of error and sample size

Same population, same method, different sample sizes. The margin falls with the square root of the sample size: four times the sample, half the margin. The population's size never enters.

Figure 2
02468101004001,6006,400Random sample sizeMargin of error (percentage points)
  • Margin of error

Worked examples

Try each one before opening the solution.

Example 1: Find a removed value from two means

The mean score of a class of 20 students on a test was 78. After one student's score was removed, the mean of the remaining 19 scores was 80. What was the removed score?

Show solution
  1. Total before removal: 20 x 78 = 1,560.
  2. Total after removal: 19 x 80 = 1,520.
  3. The removed score is the difference: 1,560 - 1,520 = 40.

    The mean went up when the score left, so the removed score had to be below the mean; 40 passes that check.

Answer: 40

Example 2: Combine two groups into one mean

In a class, the 25 boys have a mean height of 70 inches and the 15 girls have a mean height of 78 inches. What is the mean height of the whole class?

Show solution
  1. Totals: boys 25 x 70 = 1,750 inches; girls 15 x 78 = 1,170 inches.
  2. Combined total 2,920 inches over 25 + 15 = 40 students.
  3. Mean: 2,920 / 40 = 73 inches.

    Averaging 70 and 78 gives 74, which is wrong because the groups are not the same size. The combined mean sits closer to the larger group's mean.

Answer: 73 inches

Example 3: Identify an exponential model from a table

A quantity is 300 at x = 0, 360 at x = 1, and 432 at x = 2. Which model fits, and what is the quantity at x = 3?

Show solution
  1. Differences: 360 - 300 = 60 and 432 - 360 = 72. Not constant, so not linear.
  2. Ratios: 360/300 = 1.2 and 432/360 = 1.2. Constant, so exponential with base 1.2.

    A base of 1.2 is a 20 percent increase per step.

  3. Model: y = 300(1.2)^x.
  4. At x = 3: 432 x 1.2 = 518.4.

Answer: y = 300(1.2)^x; at x = 3 the quantity is 518.4.

Example 4: Decide what a study can conclude

Researchers selected 200 students at random from a large high school and randomly assigned half to study with flashcards and half to study by rereading. The flashcard group scored higher on a test. Which conclusion is supported?

Show solution
  1. Random selection from the school's students: the result generalizes to students at that school, not to all high school students everywhere.
  2. Random assignment to the two methods: the difference in scores can be attributed to the method, so cause and effect is supported.

    Without random assignment, students who chose flashcards might differ in motivation or ability, and only an association could be claimed.

  3. Combine: studying with flashcards is likely to cause higher scores than rereading for students at this high school.

Answer: Flashcards likely cause higher scores than rereading among students at this school; the claim cannot be extended to students at other schools.

Practice

Check your understanding 1

The mean of six numbers is 15. When a seventh number is added, the mean of all seven numbers is 17. What is the seventh number?

Check your understanding 2

A data set consists of the values 12, 15, 15, 18, 20 and 100. If the value 100 is removed, which of the following is true?

Check your understanding 3

In a class, the 25 boys have a mean test score of 70 and the 15 girls have a mean test score of 78. What is the mean score of the 40 students in the class?

Check your understanding 4

A quantity Q is measured at equally spaced times: 300 at t = 0, 360 at t = 1, and 432 at t = 2. Which function best models Q as a function of t?

Check your understanding 5

A researcher selected 500 adults at random from a city and asked about their exercise habits and sleep. Adults who exercised regularly reported sleeping more hours per night on average than those who did not. Which conclusion is best supported?

Check your understanding 6

A poll of 400 randomly selected voters estimated support for a candidate with a margin of error of 5 percentage points. If the poll were repeated with 1,600 randomly selected voters, the margin of error would be closest to which value?

Check your understanding 7

In a survey, 40 percent of respondents own a dog and 25 percent of respondents own both a dog and a cat. If a respondent who owns a dog is selected at random, what is the probability that the respondent also owns a cat?

Check your understanding 8

A metal block has a volume of 4,000 cubic centimeters and a density of 2.5 grams per cubic centimeter. What is the mass of the block, in kilograms?

Common mistakes

Averaging two group means when the groups differ in size. Combine totals, then divide by the combined count.

Assuming a value added to a data set equals the new mean, or that removing a value changes the mean by the amount of the value.

Calling a right-skewed distribution's median larger than its mean. The tail pulls the mean toward itself.

Checking only the first two rows of a table for a constant difference and declaring the model linear.

Claiming cause and effect from an observational study, or generalizing a random sample of one population to a different population.

Believing a larger population needs a larger sample for the same margin of error, or that a smaller sample gives a smaller margin.

Converting a squared or cubed unit by the plain factor: a square meter is 10,000 square centimeters, not 100.

Stopping in the wrong unit: 10,000 grams is the right number and the wrong answer when kilograms were asked for.

Practise this: drills for this topic in the question bank.