SAT
mathmedium~35 min

Medium: Problem Solving and Data Analysis

Percent change and successive percent changes with multipliers, multi-step rates and unit conversions, scatterplots and the line of best fit (slope meaning, predicted versus actual), conditional probability from a two-way table, comparing spread with range and standard deviation, and what a random sample and a margin of error let you conclude.

Introduction

At the Medium level the data questions stop being one-step. A percent is applied and then another percent; a rate must be converted before it is used; a scatterplot comes with a line of best fit and asks about the gap between a point and the line; a two-way table asks a probability about a subgroup, not the whole; two distributions must be compared by spread, not center; a survey result comes with a margin of error and asks what it means.

Each of these has a standard move. Percent changes are multipliers you can chain. Rates convert by cancelling units. A line of best fit is read like any line, and the distance from a data point to it is actual minus predicted. Conditional probability shrinks the denominator to the given group. Standard deviation is spread around the mean, larger when values sit farther from it. A margin of error turns one sample statistic into a plausible range for the population.

The Advanced lesson pushes each of these further: weighted means, the effect of changing one value on mean and median, exponential versus linear fits, and the difference between what an experiment and an observational study can prove.

Game plan

How to attack these questions on test day.
  1. 1

    Convert every percent change into a multiplier, then multiply the multipliers

    An increase of r percent is a multiplier of 1 + r/100, a decrease is 1 - r/100. Two successive changes are one multiplication: up 20 percent then down 25 percent is 1.20 x 0.75 = 0.90, a net 10 percent decrease. This is faster than computing each price in turn and it exposes the trap that up 10 and down 10 is 0.99, not 1.00.

  2. 2

    Compute percent change from the original, and name the original first

    Percent change is (new - original) / original, times 100. The denominator is the value before the change, always. From 40,000 to 46,000 is 6,000 / 40,000 = 15 percent, not 6,000 / 46,000. When a question says "what percent greater is A than B", B is the original.

  3. 3

    Read a line of best fit as a line, and a residual as actual minus predicted

    The slope of a line of best fit is the predicted change in y for one more unit of x; the intercept is the predicted y when x is zero, which may be meaningless in context. For a specific x, the line gives the predicted y and the dot gives the actual y. The question "how much greater is the actual value than the predicted value" is dot minus line, read at the same x. A point above the line has a positive difference, below the line a negative one.

  4. 4

    For "given that" or "of those who", shrink the table to one row or column

    Conditional probability restricts the group being chosen from. The probability that a randomly chosen student with a job is a junior uses the has-a-job column as the whole: juniors with jobs divided by all students with jobs. Cover the rest of the table with your hand and read the two numbers from the row or column that remains.

  5. 5

    Judge spread by distance from the mean, not by the number of values

    Standard deviation measures how far values typically sit from their mean. Two data sets with the same mean and the same count can have very different standard deviations; the one whose values are spread over a wider range around the mean has the larger one. The SAT never asks you to compute a standard deviation, only to compare two or to say what a larger one means, so reason about how bunched or scattered the values are.

Theory

A percent change is a multiplication. Increasing a quantity by r percent multiplies it by 1 + r/100; decreasing it multiplies by 1 - r/100. Because they are multiplications, successive changes combine by multiplying: a 20 percent increase followed by a 25 percent decrease is 1.20 x 0.75 = 0.90, a net decrease of 10 percent, regardless of the starting value. The reverse question, finding the original from the changed value, is a division by the multiplier: a jacket that costs 68 dollars after a 15 percent discount cost 68 / 0.85 = 80 dollars. The percent change between two values is the difference divided by the original value: from 40,000 to 46,000 is an increase of 6,000 / 40,000 = 0.15 = 15 percent. A percent of a percent is common in context: if 40 percent of a school's students play a sport and 25 percent of those play soccer, then 0.25 x 0.40 = 0.10, so 10 percent of the school plays soccer.

Rates and unit conversions at this level take more than one step. Speed in kilometers per hour becomes meters per second by converting each unit in turn: 72 km/h = 72,000 m / 3,600 s = 20 m/s. Density is mass per volume, so mass is density times volume and volume is mass over density. A question that gives a rate and asks for a total over a time, or gives two rates and asks which is faster, is solved by writing each rate with units and cancelling. The units are the method: if the unit you want is left after cancelling, the arithmetic was set up correctly.

A scatterplot shows paired data, one dot per observation. The association is positive if the dots rise left to right, negative if they fall, and its strength is how tightly they cluster around a line or curve. A line of best fit is the line that summarizes a linear trend; the SAT gives it as a drawn line or an equation such as y = 3.2x + 15. Its slope is the predicted change in y per unit of x, in the units of the axes; its y-intercept is the predicted value at x = 0, which may or may not be meaningful. For a particular data point, the predicted value is the line's y at that x and the actual value is the dot's y; the difference actual minus predicted is the residual, positive for points above the line. Predictions inside the range of the data are interpolation and are usually reasonable; predictions far outside it are extrapolation and are not to be trusted, a fact the test sometimes asks about directly.

A two-way table sorts a group by two categorical variables. The probability that a randomly chosen member has property A is the count with A divided by the grand total. A conditional probability changes the group: the probability of A given B is the count with both A and B divided by the count with B. In a table of 150 students where 24 of 60 juniors and 45 of 90 seniors have jobs, the probability a random student is a junior with a job is 24/150, but the probability a randomly chosen student who has a job is a junior is 24/69, because 69 students have jobs. Relative frequencies are these same fractions expressed as decimals or percents, and a table may give them instead of counts.

Measures of spread describe how far the data extend. The range is the largest minus the smallest value and is sensitive to a single outlier. The standard deviation measures the typical distance of values from the mean: data bunched near the mean have a small standard deviation, data spread far from it have a large one. Adding the same constant to every value shifts the mean and median but leaves the range and standard deviation unchanged; multiplying every value by a constant scales all of them. Comparing two dot plots or histograms, the one whose values are more concentrated near its center has the smaller standard deviation, whatever its mean. An outlier pulls the mean toward it and inflates the range and standard deviation while barely moving the median, so a distribution with a long tail has its mean pulled toward the tail.

Statistics from a sample are used to estimate a population. The estimate is trustworthy only if the sample was selected at random from the population in question; a survey of people leaving a gym says something about gym-goers, not about all residents of the city. A margin of error gives the uncertainty: if a random sample of 500 residents finds 62 percent support for a proposal with a margin of error of 4 percentage points, the plausible values for the whole population's support run from 58 to 66 percent. The margin of error shrinks as the sample grows and depends on the sample, not on the population size. It says nothing about the individuals sampled, only about the estimate for the population, and it does not mean the true value is certainly inside the interval.

Worked figures

Percent changes as multipliers

Every percent change is one multiplication; successive changes multiply together. The last two rows are the trap: up and down by the same percent does not return to the start.

Table 1
ChangeMultiplierOn 80Net effect
Increase 25%1.25100+25%
Decrease 15%0.8568-15%
Up 20%, then down 25%1.20 x 0.75 = 0.9096, then 72-10%
Up 10%, then down 10%1.10 x 0.90 = 0.9988, then 79.2-1%
Up 50%, then up 50%1.5 x 1.5 = 2.25120, then 180+125%, not +100%

A scatterplot with its line of best fit

Hours studied against test score for eight students, with the line of best fit y = 3.2x + 15. The slope says each extra hour predicts 3.2 more points. At x = 5 the line predicts 31; the student who studied 5 hours scored 35, so the actual value exceeds the predicted value by 4.

Figure 1
0102030405060024681012(5, 35): actualpredicted 31Hours studiedTest score
  • y = 3.2x + 15

Conditional probability from a two-way table

150 students by grade and job status. P(junior) = 60/150. P(junior and has job) = 24/150. P(junior given has a job) = 24/69: the group is only the 69 students in the has-a-job column. P(has a job given junior) = 24/60 = 0.4: now the group is the junior row.

Table 2
Has a jobNo jobTotal
Juniors243660
Seniors454590
Total6981150

Same mean, different spread

Both sets have mean 50 and five values. Set B's values sit far from 50, so its standard deviation is larger. Adding 10 to every value in either set would move its mean to 60 and leave its spread unchanged.

Table 3
Data setValuesMeanRangeDistances from the meanStandard deviation
A48, 49, 50, 51, 525042, 1, 0, 1, 2small, about 1.4
B30, 40, 50, 60, 70504020, 10, 0, 10, 20large, about 14
A + 1058, 59, 60, 61, 626042, 1, 0, 1, 2same as A

Reading a margin of error

Random sample of 500 residents: 62 percent support, margin of error 4 points. The interval 58 to 66 percent is the range of plausible values for support among ALL residents. It is not a statement about the 500 people asked, and a larger sample would narrow it.

Figure 2
0102030405060Lower plausible (58%)Sample estimate (62%)Upper plausible (66%)EstimatePercent supporting
  • Percent

Worked examples

Try each one before opening the solution.

Example 1: Chain two percent changes

A shirt's price is increased by 20 percent and, a month later, the new price is decreased by 25 percent. The final price is what percent of the original price?

Show solution
  1. Write each change as a multiplier: up 20 percent is 1.20, down 25 percent is 0.75.
  2. Multiply the multipliers: 1.20 x 0.75 = 0.90.

    The order does not matter for multiplication; the result would be the same if the decrease came first.

  3. Interpret: the final price is 90 percent of the original, a net decrease of 10 percent.
  4. Check with a number: 80 dollars becomes 96 after the increase and 72 after the decrease, and 72/80 = 0.90. Correct.

Answer: 90 percent of the original price

Example 2: Convert a rate in two steps

A train travels at 72 kilometers per hour. What is its speed in meters per second?

Show solution
  1. Write the rate with units: 72 km / 1 h.
  2. Convert kilometers to meters: 72 km x (1000 m / 1 km) = 72,000 m per hour.
  3. Convert hours to seconds: 1 h = 3600 s, so 72,000 m / 3600 s.

    Both conversions are multiplications by a fraction equal to 1, arranged so the old unit cancels.

  4. Divide: 72,000 / 3,600 = 20 m/s.

Answer: 20 meters per second

Example 3: Predicted, actual and the difference on a scatterplot

A line of best fit for a scatterplot of hours studied (x) and test score (y) is y = 3.2x + 15. One student studied 5 hours and scored 35. By how much does that student's actual score exceed the score predicted by the line, and what does the slope 3.2 mean?

Show solution
  1. Predicted score at x = 5: 3.2(5) + 15 = 16 + 15 = 31.
  2. Actual score: 35, read from the data point.
  3. Difference, actual minus predicted: 35 - 31 = 4. The point sits 4 above the line.
  4. Interpret the slope: for each additional hour studied, the predicted score increases by 3.2 points.

    Slope is a predicted change per unit of x, in y's units. It is not the score of a student who studied one hour; that would be 18.2.

Answer: The actual score exceeds the predicted score by 4 points; the slope means each extra hour of study predicts 3.2 more points.

Example 4: Conditional probability from a two-way table

Of 150 students, 60 are juniors and 90 are seniors. 24 of the juniors and 45 of the seniors have part-time jobs. If a student with a part-time job is chosen at random, what is the probability that the student is a junior?

Show solution
  1. Identify the group being chosen from: students with jobs. Count them: 24 + 45 = 69.

    "A student with a part-time job is chosen" makes the has-a-job column the whole. The 150 is irrelevant.

  2. Count the favorable members of that group: juniors with jobs, 24.
  3. Divide: 24/69 = 8/23, about 0.348.
  4. Contrast: the probability that a random student is a junior with a job would be 24/150 = 0.16; the probability that a random junior has a job would be 24/60 = 0.4. Three different questions, three different denominators.

Answer: 24/69, which simplifies to 8/23

Practice

Check your understanding 1

A town's population grew from 40,000 to 46,000 over a decade. By what percent did the population increase?

Check your understanding 2

The price of an item is increased by 20 percent, and then the new price is decreased by 25 percent. The final price is what percent less than the original price?

Check your understanding 3

A scatterplot relating hours of practice x to points scored y has a line of best fit with equation y = 3.2x + 15. A player who practiced 5 hours scored 35 points. How much greater is the player's actual score than the score predicted by the line?

Check your understanding 4

A two-way table shows that of 150 students, 24 juniors and 45 seniors have part-time jobs, and 36 juniors and 45 seniors do not. If a student who has a part-time job is selected at random, what is the probability that the student is a junior?

Check your understanding 5

Data set A is 48, 49, 50, 51, 52. Data set B is 30, 40, 50, 60, 70. Which statement is true?

Check your understanding 6

A random sample of 500 residents of a city was surveyed, and 62 percent said they support a new park, with a margin of error of 4 percentage points. Which conclusion is best supported?

Check your understanding 7

A car is traveling at 72 kilometers per hour. What is the car's speed in meters per second?

Common mistakes

Adding or subtracting successive percents instead of multiplying their multipliers: up 20 then down 25 is a 10 percent decrease, not a 5 percent one.

Dividing a change by the new value instead of the original when computing percent change.

Reading the slope of a line of best fit as a value of y rather than a rate of change of y per unit of x.

Computing predicted minus actual when the question asks how much the actual exceeds the predicted, and getting the sign backwards.

Using the grand total as the denominator of a conditional probability. "Of those who" and "given that" name the denominator.

Judging standard deviation by the number of data points or by the mean. It is about distance from the mean.

Applying a margin of error to the sample instead of the population, or believing a larger sample gives a larger margin of error.

Generalizing from a sample to a population the sample was not drawn from.

Practise this: drills for this topic in the question bank.