Skip to main content

AP® · Full Course

AP® Statistics

All 5 units of the Fall 2026 course: exploring and collecting data, probability and distributions, and inference for proportions, means, and regression.

Updated for Fall 2026

Start Unit 1 free. Unit 1: Exploring One-Variable Data is open to everyone, no account needed. Other topics are locked.

Unit 1: Exploring One-Variable Data

THE BIG PICTURE. In the Fall 2026 framework, Unit 1: Exploring One-Variable Data and Collecting Data is 20–30% of the AP exam, the heaviest unit. This first section covers DESCRIPTIVE STATISTICS for a single variable: types of data, displays, summary measures, and the normal distribution. Collecting Data (the rest of Unit 1) is the next section. Mastery of xˉ\bar{x}, ss, zz-scores, the normal distribution, and the CUSS framework for describing distributions is essential for every later unit.

TYPES OF DATA

CATEGORICAL (qualitative): values are labels/categories. Examples: eye color, favorite sport, blood type, yes/no responses.

QUANTITATIVE (numerical): values are numbers with meaning. Two subtypes

  • DISCRETE: countable; usually integers (number of siblings, dice roll).
  • CONTINUOUS: measurable; can take any value in a range (height, weight, temperature).

The DATA TYPE determines which displays and analyses are appropriate.

DISPLAYS FOR CATEGORICAL DATA

  • FREQUENCY / RELATIVE FREQUENCY TABLE: counts and percents per category.
  • BAR CHART: bars for each category; height = frequency or percent. Bars do NOT touch (contrast with histogram).
  • PIE CHART: circular; slices proportional to category percent. Less precise than bar charts.

DISPLAYS FOR QUANTITATIVE DATA

Histogram of Wind Speeds (Lee Ranch) A histogram displays the frequency distribution of a quantitative variable using adjacent bars. Bin width affects appearance: too wide hides shape, too narrow exaggerates noise. Always describe in context using CUSS: Center, Unusual features, Shape, Spread.

U.S. Department of Energy / Wikimedia Commons contributors (opens in new tab), Public domain (U.S. Government)

Elements of a Box Plot A box plot displays the five-number summary: min, Q1, median, Q3, max: with the box spanning the IQR. Whiskers extend to the most extreme non-outlier data points; observations beyond 1.5 × IQR from the box are flagged as outliers. Best for comparing distributions across groups.

KStrileckis / Wikimedia Commons (opens in new tab), CC0
  • DOTPLOT: each value plotted above a number line; good for small datasets.
  • STEMPLOT: splits each number into stem (left digits) and leaf (right digit). Preserves all data values.
  • HISTOGRAM: bars touch; bins of equal width; height = frequency or relative frequency. Most common quantitative display.
  • BOXPLOT (BOX-AND-WHISKER): shows 5-number summary (min, Q1, median, Q3, max). Outliers shown as separate points beyond whiskers.
  • CUMULATIVE FREQUENCY (OGIVE): running total of frequencies; useful for percentile readings. (Beyond the 2026 exam: the CED lists only dotplots, stemplots, and histograms as displays.)

One data set, three displays The dotplot and stemplot keep every value; the histogram groups values into bins of width 10. All three show a right-skewed distribution with an outlier at 52 minutes.

WORKED EXAMPLE: THREE DISPLAYS OF ONE DATA SET

A random sample of 20 students at School A reported their commute times in minutes: 5, 8, 10, 12, 12, 14, 15, 15, 16, 18, 18, 20, 20, 22, 24, 25, 28, 30, 35, 52

  • Dotplot: one dot per student above a number line; repeated values stack. Every value is visible, so the lone dot at 52 stands out at once.
  • Stemplot: the stem is the tens digit and the leaf is the ones digit (key: 1 | 4 = 14 minutes). It keeps every value and shows shape like a sideways histogram. Always include a key.
  • Histogram: with bins of width 10 the counts are 2, 9, 6, 2, 0, 1. Values on a boundary go in the bin that starts there (20 goes in 20 to 30). A different bin width can change the picture, so read shape with care.

All three show the same story: most commutes fall between 10 and 30 minutes, the distribution is skewed right, and one student reports an unusually long 52-minute commute.

DESCRIBING DISTRIBUTIONS: CUSS

Always describe quantitative distributions with:

  • C: CENTER: mean, median, mode.
  • U: UNUSUAL FEATURES: outliers, gaps, clusters.
  • S: SHAPE: symmetric, skewed (left/right), unimodal/bimodal/uniform.
  • S: SPREAD: range, IQR, standard deviation.

ALWAYS in CONTEXT: name the variable and units.

Shape and the mean versus the median The mean is pulled toward the long tail; the median is not. In a symmetric distribution the two are about equal.

WORKED EXAMPLE: DESCRIBING A DISTRIBUTION IN CONTEXT

Describe the distribution of commute times for the 20 School A students above.

Summary statistics (TI-84 1-Var Stats): xˉ=19.95\bar{x} = 19.95, s≈10.67s \approx 10.67, min = 5, Q1=13Q_1 = 13, median = 18, Q3=24.5Q_3 = 24.5, max = 52.

  • Shape: the distribution of commute times is skewed to the right and unimodal, with a peak between 10 and 20 minutes.
  • Center: because of the skew, report the median: 18 minutes. (The mean, 19.95 minutes, is pulled up toward the long right tail.)
  • Variability: the IQR is 24.5−13=11.524.5 - 13 = 11.5 minutes, so the middle half of commutes spans 11.5 minutes. The range is 52−5=4752 - 5 = 47 minutes.
  • Unusual features: the 52-minute commute is an outlier (checked with the 1.5 × IQR rule below); there is a gap between 35 and 52 minutes.

Every sentence names the variable (commute time) and the units (minutes). A description that says only "skewed right, center 18, spread 11.5" loses credit on the AP exam because it has no context.

PRACTICE: WHICH SUMMARY SHOULD YOU REPORT?

SituationCenterVariabilityWhy
Roughly symmetric, no outliersMeanStandard deviationMean and SD use every value and describe symmetric data well
Strongly skewed (home prices, incomes)MedianIQRMedian and IQR are resistant to the long tail
Symmetric except one extreme outlierMedian (or mean without the outlier, clearly stated)IQRThe outlier inflates the mean and SD
Comparing two skewed groupsMediansIQRsUse the same resistant measures for both groups
Data will be used in a normal model laterMeanStandard deviationNormal models are defined by μ\mu and σ\sigma

CENTER (Measures of Central Tendency)

  • MEAN: xˉ=∑xn\bar{x} = \dfrac{\sum x}{n} (sample mean) or μ\mu (population mean). Sensitive to outliers.
  • MEDIAN: middle value when data sorted. Resistant to outliers. For even nn, average of two middle values.
  • MODE: most common value(s). Can be 0, 1, or many.

SHAPE INFLUENCE on center:

  • SYMMETRIC: mean ≈ median.
  • SKEWED RIGHT (long tail to right): mean > median (mean pulled toward tail).
  • SKEWED LEFT (long tail to left): mean < median.

SPREAD (Measures of Variability)

  • RANGE = max − min. Very sensitive to outliers.
  • INTERQUARTILE RANGE (IQR) = Q3 − Q1 (middle 50%). Resistant.
  • VARIANCE = average squared deviation. Population: σ2=∑(x−μ)2N\sigma^2 = \dfrac{\sum (x - \mu)^2}{N}. Sample: s2=∑(x−xˉ)2n−1s^2 = \dfrac{\sum (x - \bar{x})^2}{n-1} (note n−1n-1 for sample: Bessel's correction).
  • STANDARD DEVIATION = square root of variance: σ\sigma or s=∑(x−xˉ)2n−1s = \sqrt{\dfrac{\sum (x - \bar{x})^2}{n-1}}. Same units as data.

OUTLIER RULE (1.5 × IQR)

An observation is an outlier if:

  • Below Q1−1.5×IQRQ_1 - 1.5 \times IQR, OR
  • Above Q3+1.5×IQRQ_3 + 1.5 \times IQR.

WORKED EXAMPLE: THE 1.5 × IQR RULE, STEP BY STEP

For the School A commute times: Q1=13Q_1 = 13 and Q3=24.5Q_3 = 24.5, so IQR=11.5IQR = 11.5 minutes.

  • Lower fence: 13−1.5(11.5)=13−17.25=−4.2513 - 1.5(11.5) = 13 - 17.25 = -4.25. No commute can be below this, so there are no low outliers.
  • Upper fence: 24.5+1.5(11.5)=24.5+17.25=41.7524.5 + 1.5(11.5) = 24.5 + 17.25 = 41.75 minutes.
  • Conclusion: 52 > 41.75, so the 52-minute commute is an outlier. The next largest value, 35, is not.

Resistance check: removing the 52 changes the median not at all (still 18) but drops the mean from 19.95 to about 18.26 and the standard deviation from about 10.67 to about 7.76 minutes. That is what "resistant" and "nonresistant" mean. The CED also accepts a second rule, more than 2 standard deviations from the mean: here 19.95+2(10.67)≈41.319.95 + 2(10.67) \approx 41.3, so 52 is flagged by that rule too. State which rule you use.

Comparing distributions with boxplots School B has the higher median; the IQRs are similar. School A's 52-minute commute lies beyond the upper fence, Q3 + 1.5(IQR) = 41.75, so it is plotted as an outlier and the whisker stops at 35.

WORKED EXAMPLE: COMPARING TWO DISTRIBUTIONS

A random sample of 20 students at School B gives commute times with five-number summary 12, 21.5, 26.5, 31.5, 42 minutes (mean 26.7, SD about 7.62). Compare the two schools.

  • Center: the median commute at School B (26.5 minutes) is greater than the median at School A (18 minutes).
  • Variability: the IQRs are similar (10 minutes at B versus 11.5 at A), but School A's range (47 minutes) is much larger than School B's (30 minutes) because of its outlier.
  • Shape: School A is skewed to the right; School B is roughly symmetric.
  • Unusual features: School A has a high outlier at 52 minutes (fence 41.75). School B has no outliers (fences 6.5 and 46.5).

Use comparative words ("greater than", "less than", "similar to"). Listing two sets of numbers side by side without comparing them does not earn the comparison point.

POSITION

  • PERCENTILE: value below which a given percent of observations fall. The 75th percentile = Q3Q_3.
  • z-SCORE: z=x−μσz = \dfrac{x - \mu}{\sigma}. Number of standard deviations from the mean. Allows comparison across distributions. Negative = below mean; positive = above.
  • QUARTILES: Q1Q_1 (25th percentile), Q2Q_2 (median, 50th), Q3Q_3 (75th).

WORKED EXAMPLE: Z-SCORES ACROSS DIFFERENT SCALES

Suppose one student scores 1310 on a test with mean 1050 and SD 210, and another scores 29 on a test with mean 20.8 and SD 5.8. Who did better relative to other test takers?

  • First student: z=1310−1050210≈1.24z = \dfrac{1310 - 1050}{210} \approx 1.24.
  • Second student: z=29−20.85.8≈1.41z = \dfrac{29 - 20.8}{5.8} \approx 1.41.
  • The second student's score is 1.41 standard deviations above the mean, compared with 1.24, so the second student did better relative to the other test takers.

A z-score has no units, which is exactly why it allows comparisons across different scales.

PRACTICE: WHAT HAPPENS WHEN YOU CHANGE UNITS OR SHIFT THE DATA?

Change to every valueMean and medianSD and IQRShape
Add 5 minutes (a detour for everyone)Increase by 5 (19.95 becomes 24.95)UnchangedUnchanged
Convert minutes to hours (divide by 60)Divide by 60 (19.95 becomes about 0.33 hours)Divide by 60 (10.67 becomes about 0.18 hours)Unchanged
Multiply by 2DoubleDoubleUnchanged
Subtract the mean and divide by the SDMean becomes 0SD becomes 1Unchanged: z-scores do not make data normal

NORMAL DISTRIBUTION

Symmetric, bell-shaped. Defined by μ\mu (center) and σ\sigma (spread). Notation: X∼N(μ,σ)X \sim N(\mu, \sigma).

EMPIRICAL RULE (68-95-99.7)

  • ~68% of data within ±1σ\sigma of mean.
  • ~95% within ±2σ\sigma.
  • ~99.7% within ±3σ\sigma.

STANDARD NORMAL Z∼N(0,1)Z \sim N(0, 1): mean 0, SD 1. Convert any normal XX to ZZ via z=x−μσz = \dfrac{x - \mu}{\sigma}.

CALCULATOR FUNCTIONS (TI-83/84)

  • P(a<X<b)P(a < X < b): normalcdf(a, b, μ, σ).
  • P(X<P(X < value): normalcdf(-1E99, value, μ, σ).
  • Find percentile value: invNorm(percentile, μ, σ).

ASSESSING NORMALITY: histogram should be roughly bell-shaped; boxplot symmetric; NORMAL PROBABILITY PLOT (NPP) should be roughly linear.

DENSITY CURVES

describe continuous distributions. Total area under curve = 1. Probability of value in an interval = area under curve over that interval.

CASE STUDY: WHY THE CENSUS REPORTS MEDIAN INCOME

The U.S. Census Bureau's report Income in the United States: 2024 (September 2025) put median household income at about $83,730, while mean household income was about $121,000, roughly $37,000 higher. The previous year's report showed the same pattern (median $80,610, mean $114,500 for 2023). A relatively small number of very high-income households pulls the mean up, so the Census Bureau's headline figure is the median.

  • Concept: a strongly right-skewed distribution; mean greater than median.
  • Graph: a histogram of household incomes has a long right tail, like the "skewed right" panel of the shapes figure.
  • Lesson: for skewed data, report the median and IQR; quoting only the mean describes a "typical" household that earns more than most households do.

CASE STUDY: OLD FAITHFUL'S TWO PEAKS

A widely used data set of 272 eruptions of the Old Faithful geyser in Yellowstone National Park (the "faithful" data in R, from Härdle, 1991; see also Azzalini and Bowman, 1990) records how long each eruption lasted and the wait until the next one. Waiting times range from 43 to 96 minutes with a mean of about 70.9 and a median of 76 minutes, and a histogram shows two clear peaks: 99 waits are shorter than 67 minutes and 173 are 67 minutes or longer. The National Park Service uses the same pattern to predict eruptions: a short eruption (under about 3 minutes) tends to be followed by a shorter wait than a long one.

  • Concept: a bimodal distribution, often a sign that two groups are mixed together.
  • Graph: a histogram with two peaks and a dip between them; a single mean (70.9 minutes) falls in the dip, where few waits actually occur.
  • Lesson: always look at the shape before summarizing; a single center can describe almost nobody in a bimodal distribution.

EXAM CONNECTIONS. Always describe distributions with CUSS in context. Distinguish mean vs median based on skewness (mean follows the tail). For position questions, calculate z-scores and use them for cross-distribution comparison ("Anna scored 1.5 SD above her class mean; Ben scored 1.2 SD above his: Anna ranked higher within her class"). Apply the empirical rule for quick normal estimates (e.g., between μ−2σ\mu - 2\sigma and μ+2σ\mu + 2\sigma holds ~95%). For outliers, use the 1.5 × IQR rule unless told otherwise (the CED also accepts the 2-standard-deviation rule). Be ready to calculate ALL summary statistics by hand AND with calculator.

Key Terms

Categorical vs Quantitative Variables

Categorical = labels (eye color). Quantitative = numbers with meaning. Quantitative subtypes: discrete (countable, usually integer) vs continuous (measurable, any value in range).

Mean and Standard Deviation

Mean: xˉ=∑xn\bar{x} = \dfrac{\sum x}{n} (sample) or μ\mu (population). Standard deviation: s=∑(x−xˉ)2n−1s = \sqrt{\dfrac{\sum (x - \bar{x})^2}{n-1}} (sample) or σ\sigma (population). SD has same units as data.

Median

Middle value when data sorted. For even nn, average of two middle values. RESISTANT to outliers (unlike mean). Use median when data are skewed.

Five-Number Summary

Min, Q1Q_1, median, Q3Q_3, max. Visualized in boxplot. IQR=Q3−Q1\text{IQR} = Q_3 - Q_1 measures middle-50% spread.

Outlier (1.5 × IQR Rule)

A value is an outlier if it falls below Q1−1.5×IQRQ_1 - 1.5 \times IQR or above Q3+1.5×IQRQ_3 + 1.5 \times IQR. Standard rule for boxplots.

CUSS Framework

Center, Unusual features, Shape, Spread: what to mention when describing any quantitative distribution. Always IN CONTEXT (variable + units).

Skewness

Symmetric: mean ≈ median. Skewed RIGHT (positive skew): long right tail; mean > median. Skewed LEFT: long left tail; mean < median. Mean is pulled toward the tail.

z-score

z=x−μσz = \dfrac{x - \mu}{\sigma}. Number of standard deviations a value is from the mean. Allows comparison across distributions. Negative = below mean; positive = above.

Empirical Rule (68-95-99.7)

For approximately normal distribution: 68% within ±1σ\sigma of mean; 95% within ±2σ\sigma; 99.7% within ±3σ\sigma.

Standard Normal Distribution

Z∼N(0,1)Z \sim N(0, 1): mean 0, SD 1. Any normal distribution can be converted to standard normal using z-scores. Allows use of standard tables/calculator.

normalcdf and invNorm

TI calculator functions. normalcdf(low, high, μ, σ) gives probability between bounds. invNorm(percentile, μ, σ) gives the value at a given percentile.

Density Curve

Continuous probability distribution. Total area under curve = 1. Probability of falling in an interval = area under curve over that interval.

Exam Tips

  • ALWAYS describe distributions in CONTEXT (variable + units), and ALWAYS use CUSS: Center, Unusual features, Shape, Spread.
  • For SKEWED data, report MEDIAN + IQR (resistant). For SYMMETRIC, report MEAN + SD.
  • Use z-scores to compare values from different distributions.
  • Memorize the 68-95-99.7 rule for normal distributions cold: questions test it without giving the rule.
  • For outliers, default to the 1.5 × IQR rule unless told otherwise; the CED also accepts values more than 2 standard deviations from the mean.

Now practice this topic

Test yourself on what you just read.

End of the free unit

Ready to keep going?

That was the free first unit of the AP® Statistics Course Companion. Unlock the complete guide to continue with the rest of the course.

  • All 10 topics with full study notes
  • 225 flashcards
  • 170 practice questions with explanations
  • 38 free-response problems with worked solutions
  • Term-matching games
  • 120 key terms

$24.99 one-time, 12 months of access.

Taking more than one course? Student Pass: up to 3 guides, $49.99 · Platinum Unlimited: every guide, $74.99.

Already bought it? Sign in

The Prep Den AP® Statistics study guide is a complete companion for revising the course: topic-by-topic study notes, interactive flashcards, practice quizzes with worked explanations, a key-terms bank, and exam-technique tips. The first unit above is free to use; the full guide unlocks every topic.

Also searched as: AP Statistics study guide, AP Stats study guide, AP Statistics quiz, AP Stats quiz, AP Statistics practice quiz, AP Stats practice quiz, AP Statistics practice questions, AP Stats practice questions, AP Statistics flashcards, AP Stats flashcards, AP Statistics notes, AP Stats notes, AP Statistics review, AP Stats review.