0% found this document useful (0 votes)
9 views26 pages

AP Statistics Study Guide Overview

Uploaded by

Qamar Jan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views26 pages

AP Statistics Study Guide Overview

Uploaded by

Qamar Jan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

ap stats study guide

ap stats study guide


general tips
FRQs
● ALWAYS include context (variable & units)
MCQs
strategies: process of elimination, read actively & carefully, underline important info, anticipate
answers

chapter 1: one variable data


content overview
introduction: statistics
statistics: the science & art of collecting, analysing, and drawing conclusions from data
types of variables

categorical quantitative
● nominal: certain order ● discrete: Fxed set of possible values
● ordinal: no order with gaps between them
○ whole numbers or deFned
*could be numbers if they don’t intervals
measure anything (eg. cell phone ○ countable or countably inFnite
digits) ● continuous: inFnite possibilities
○ decimals / fractions
○ any value in an interval on the
number line

the main difference: whether they measure something

basic statistics vocab


● individual: object described in a set of data
○ also known as cases / observational units
● variable: aspect that can take diff values for diff individuals
● distribution: pattern of variation of a variable
○ shows what values the variable takes & how often it takes them
types of statistics
● descriptive statistics: analysing data (units 1-7)
● inferential statistics: making inferences / drawing conclusions from data (units 8-12)
:
1.1 analysing categorical data

one-variable categorical data


tables graphs
frequency table bar graph
relative frequency table tips for bar graphs
● equal width bars
*these can also be used for quantitative data but it’s far ● leave gaps in between bars
less common bc it’s not usually very convenient ● label & scale axes
● can use either frequencies or relative
frequencies (but need to indicate which one!)

pie chart
tips for pie charts
● good for: comparing categories to the whole
● areas of slices proportional to frequencies /
relative frequencies
● must include all possible categories in the
whole (add an “other” option if necessary)
● include a legend key

two-variable categorical data


tables graphs
two-way tables side-by-side bar graph
summarises data on relationship between two bar graphs showing the distribution of a categorical
categorical variables for a group of individuals variable for each value of another categorical variable
(grouped side-by-side)
IMPORTANT: FREQUENCIES (see chart below)
● marginal relative frequency: B/C tip: it is acceptable for the bars to touch here w/in each
● joint relative frequency: A/C of the values for the categorical variable! but leave
● conditional relative frequency: A/B spaces in between the distributions for each value.

var 1 total
segmented bar graph
distribution of a categorical variable as segments of a
whole (bars stacked on top of each other & proportional
var 2 A B
to relative frequencies)

uses relative frequencies!


total C
tip: the bars do not touch here!

mosaic plot
segmented except the width of the bars proportional to
number of individuals in that category

tip: the bars do touch here!


*note: tables are not data, they are summaries of data!
avoid these bad statistics practices: truncating axes & using pictograms
association: knowing the value of one variable allows you to predict value of the other
● association does NOT equal causation!
● to determine: calculate the distribution of the variable you want to predict for each
value of the other variable

1.2 displaying quantitative data with graphs


types of graphs for quantitative data: dotplot, stemplot, histogram, & boxplot
:
dotplot stemplot histogram boxplot
tips always add a key bins corresponds to
equal size bins the Fve-number
stems: all but Fnal digit summary
try to go with a minimum of 5
leaves: Fnal digit bins

split stems to better see the bin size will affect the
distribution if needed (try to appearance of the distribution
get at least 5 stems). make (more bins -> more detail but
sure each stem has an less clear overall pattern)
equal number of possible
leaf digits. edges of bins either inclusive or
noninclusive (typically the right
back-to-back edge of a bin is noninclusive)
stemplot: quantitative data
that’s split into two groups bars

bars touch

either frequencies or relative


frequencies (but need to
indicate which one!)
*use relative frequencies if
you’re using the histogram to
compare distributions with diff
numbers of observations*

pros can see every can see every individual easier to make for large data nonresistant
individual value value sets
easy to make
easy to see easy to see shape easy to see shape (esp for large for large data
shape data sets–simpliFes the overall sets
pattern)
shows Fve-
number
summaries

splits data into


quartiles

cons didcult to make didcult to make with large doesn’t show every individual doesn’t show
with large data data sets value individual
sets values, slight
skewness,
gaps, clusters,
or peaks

describing quantitative distributions: SOCV


● shape
○ skewness & modality
○ and any clusters / gaps
● outliers (1.5 x IQR)
● center (measures of the typical value: mean / median / mode)
● variability (range / IQR / st dev)
● and ALWAYS ADD CONTEXT! (variables & units)

1.3 describing quantitative data with numbers


:
measures of center
● mean: average (x̅ or µ)
● median: middle value
● mode: most common value
comparing mean & median: mean nonresistant & median resistant
resistant: not sensitive to skewness / outliers
measures of variability
● range
○ pros: easy to calculate
○ cons: nonresistant & doesn’t express variability from the center
● IQR: interquartile range (middle 50% of values)
○ pros: resistant
○ calculating quartiles: split each half (leave out median!)
● standard deviation: typical distance from mean (sx or σ)
○ to calculate: calculate all the deviations (value – mean), square each, add up,
divide by n-1, take square root
○ properties
■ only use in tandem with the mean
■ sx always greater than or equal to σ
○ cons: nonresistant
outliers
● 1.5 x IQR (above Q3 or below Q1)
Gve number summary: minimum, Q1, median, Q3, maximum
also good to know: upper & lower bounds (1.5 x IQR) (might not be actual data values)
*decide which measures to use based on whether resistancy is a concern for the distribution
additional vocab
● statistic: a value describes a characteristic of a sample
● parameter: a value that describes a characteristic of a population

types of mcqs (or frqs)


● compare mean & median (use knowledge of resistancy & skewness to answer)
● interpret graphical or tabular representation of data
○ (using it to answer questions about the variable / distribution of the variable)
● interpret summary statistics / a value abt the distribution
○ put it in context (variable & units!)
○ especially interpreting standard deviation: “the typical distance from the mean”
● match data with its graphical representation OR match graphical representations of the
same data
● determine whether there is an association between two categorical variables given
data: graphical / tabular representations OR summary statistics
○ calculate distribution of one categorical variable for each value of the other
(basically a whole bunch of conditional relative frequencies)
○ and then Fnd whether knowing the value of one variable allows you to predict
the value of another (eg. are those distributions you calculated the same or not)

types of frqs
describe a distribution (quantitative)
● SOCV (shape, outliers, center, variability) (see content overview for more information)
● context (variables & units of the distribution)
compare distributions (quantitative)
● SOCV for both distributions
● use explicitly comparative language that relates the two distributions
make a claim / argument based on a distribution (quantitative)
● refer to speciFc characteristics (eg. SOCV) of the distribution in your answer
:
○ give speciFc numbers as much as possible
● context (variables & units of the distribution)
● then, explain why those characteristics support your claim / argument
describe a distribution / compare distributions / make a claim based on a distribution
(categorical)
construct a certain type of graph for given data
● follow appropriate guidelines for the type of graph (see content summaries for tips)
● always label & scale axes appropriately! with units!
● add a title with context
& others
● what is apparent from the histogram but not from the boxplot?
● misrepresenting / manipulating data
○ why would it be misleading to only report [insert statistic / parameter here]?
○ what would you want to report in order to [achieve speciFed goal]?
● why does one method for determining outliers give you more outliers than the other?

language & wording / general common mistakes


language & wording
● ALWAYS INCLUDE CONTEXT
○ distribution & variables & units
● describing distributions: “appears to be” / “approximately” (bc you cannot be sure)
○ ex. “approximately normal” & “roughly symmetric” (this is a very important one!)
● be VERY careful with relative frequencies vs. frequencies / raw counts! (this is a very
important one!)
○ use relative frequencies with groups of different sizes!
○ & say “a greater percentage” not “more”
○ plurality vs majority
○ always indicate which one you’re using
● for histograms & boxplots: keep in mind that you can’t conclusively determine what the
values are
○ esp for histograms: need to say that a value is in a certain bin (“between [value]
and [value]”)
common mistakes
● range, IQR, and st. dev. are single values! not a range of values
● avoid bad statistics: truncated axes & pictograms
● association is not causation!

chapter 2: modeling distributions of


quantitative data
content overview
2.1 describing location in a distribution

percentile: pth percentile is value with p% observations less than or equal to it


● a way to describe location in a distribution
● allows for a standard scale to compare values from different distributions
cumulative relative frequency graphs / ogives: plots points corresponding to the
percentile of a value in the distribution & points connected with line segments to create
the graph
● works well with a frequency table of quantitative data
:
standardized scores (z-scores): how many standard deviations from the mean a value is (& what
direction)
● another way to describe location in a distribution
● (value – mean) / st dev
● allows for a standard scale to compare values from different distributions
transforming data
adding / subtracting constant: affects measure of center / location (not shape / variability)
multiplying / dividing by a constant: affects measures of center, location, variability (not shape)
multiple transformations: follow order of operations
transformations related to z-scores: in a distribution of z-scores, shape remains the same as
original distribution, mean always 0, standard deviation always 1

2.2 density curves & normal distributions


density curve: simpliFed model of a distribution of a quantitative variable
● always on or above horizontal axis, has an area of exactly 1 underneath it
● always an approximation of data (not an exact model)
● models continuous data but often used to approximate discrete distributions as well
describing density curves
● shape: same ways as usual
● center: mean (balance point) (µ) & median (divides area of curve in half)
○ if symmetric, they’re the same
● variability: same measures as usual (σ)
normal distributions: bell shaped & symmetric & unimodal distribution
● approximated with a normal curve (density curve)
● can fully be described by mean (same as median) (µ) and standard deviation (σ)
● useful for: real data, chance outcomes, inference methods
empirical rule: 68% (within 1σ of µ) - 95% (within 2σ of µ) - 99.7% (within 3σ of µ) (for normal
distributions)
Gnding areas in normal distributions
● empirical rule (when applicable)
● Fnd z-score & use table a to look up p-value (percent of values to left of z-score)
○ table a connects z-scores to percentiles in a normal distribution
○ standard normal distribution: distribution of z-scores (mean 0, st dev 1)
● technology (see calculator functions)
types of problems: area to left, area to right, area between, working backwards (z-score given
area)
assessing normality
● plot data (see if it looks normal)
● check against empirical rule
○ check amount of data within 1, 2, and 3 st dev from mean (w/in 3-5% is pretty
good!)
● normal probability / normal quantile plot
○ plots actual z-score (x) vs predicted z-score if it was normal (y)
○ look for a straight-ish line on the normal probability plot
*ideally use all three methods to check!

types of mcqs (or frqs)


● interpret percentile or z-score (see content summary for info)
○ provide speciFc values for the percentages / mean & st dev
:
● use cumulative relative frequency graphs to determine percentile
● describe how distribution of data will change with a given type of transformation
● Fnd area in a normal distribution (see content summary for info on how)

types of frqs
● use percentiles or z-scores to evaluate claims about data
○ Fnd & interpret percentile / z-score
■ percentile: value with p% of observations less than or equal to it
■ z-score: abt that many standard deviations above / below the mean
○ draw conclusion using percentile / z-score
● normal distribution questions
○ picture
■ draw normal curve
■ label speciFc distribution (context & mean / st dev)
■ label boundary values & shade area of interest
○ calculate z-score(s)
○ calculate p-values using table a or calculator
■ if you use calculator, ALWAYS LABEL VALUES
○ state answer in context
● is it extrapolation / is the answer reasonable?
● predict value & comment on whether it’s reliable
○ E: correct prediction, plugged into formula for LSRL to get predicted value, say
whether it is reliable or not (extrapolation)

language & wording / general common mistakes


● ALWAYS ADD CONTEXT esp when describing location in a distribution
● percentile: “at” a percentile NOT “in” a percentile (bc it’s a location!)
● z-scores: always provide the direction! not just “away” from the mean
○ and provide context (units & distribution / variables)!
● normal distributions
○ be careful with direction & tails (two-sided vs one-sided)
○ distributions of real world data always “approximately normal” (never perfect)

overall: units 1 & 2 & 3


quantitative data
● discrete & continuous
● one-var
○ tabular: frequency table, relative / cumulative frequency table
○ graphical: dotplot, stemplot, boxplot, histogram, cumulative relative frequency
graph
○ numerical: 5-number summary, center, variability, percentile / z-score
● two-var
○ graphical: scatterplot
○ numerical: r, r2, LSRL, s
simpliFed model of data: density curves
categorical data
● nominal & ordinal
● one-var
○ tabular: frequency table, relative frequency table
○ graphical: bar chart, pie chart
○ numerical: proportions, etc
● two-var
○ tabular: two-way table (frequency OR relative frequency)
○ graphical: side-by-side bar chart, segmented bar chart, mosaic plot
○ numerical: proportions, association, etc
:
chapter 3: exploring two-variable
quantitative data
3.1 scatterplots & correlation
explanatory & response variables
● not necessarily causation (though it could be), just which helps to explain the other
scatterplots
tips:
● explanatory x-axis, response y-axis
● label & scale axes (you CAN truncate the axes here)
describing scatterplots: CDOFS
● context: state variables & units
● direction: pos / neg / no correlation
● outliers / unusual features: outliers & points outside the general pattern / clusters
● form: linear / nonlinear
● strength: r (correlation coedcient) / r2 value (measures whether LSRL is a good Ft)
r value: measures strength & direction (ONLY for linear models)
● cautions
○ r is nonresistant
○ only for linear
○ correlation is not causation!
○ no units
○ unaffected by changing units / changing explanatory & response variables
● | r | less than 0.5: weak, | r | between 0.5 and 0.75: moderate, | r | greater than 0.75:
strong
extrapolation: using a regression line to make predictions way outside of the interval of x-values
used to generate the line (beyond the scope of your data)
● won’t be accurate bc it might not remain linear at such extreme points

3.2 linear regression


regression line: model of how response variable (y) changes as explanatory variable (x) changes
ŷ = a + bx (y-hat is predicted y-value for a given x-value)
predicting so it’s okay if y-hat is not an integer for a real world situation (think of it as an
average)
residuals: actual value – predicted value (based on line)
● a good linear regression line minimizes the residuals
least-squares regression line: line that minimises sum of squared residuals
● sum of residuals on an LSRL is always 0
residual plots: scatterplot that plots residuals against explanatory variable
● determines whether a linear model is appropriate (check for random scatter & no
leftover curved pattern)
● how well does the line work? -> how good will predictions be?
s: standard deviation of residuals
● measures typical residual (distance between predicted & actual)
● calculate in the same way as st dev but divide by n-2

r2: coedcient of determination (value between 0 and 1, usually expressed as a percentage)


● square of correlation r
:
○ when Fnding r from r2, make sure to consider direction of correlation!
● how well the LSRL Fts the data: percent reduction in sum of squared residuals when
using LSRL instead of mean to make predictions
○ what percent of the variability in the response variable that can be explained by
the linear association
regression to the mean (to calculate LSRL):
slope: b = r (sy / sx )
y-int: a = mean y - b(mean x)
since LSRL passes through (mean x, mean y)
correlation & regression wisdom
● correlation and LSRLs only describe linear relationships
● r, s, r2, and LSRL are non-resistant (see intuential points)
inOuential points
points that, if removed, substantially change the slope, y-int, r, r2, or s
outliers: doesn’t follow pattern of data and has a large residual
high leverage: much larger / smaller values than other values in data set
*these are very often intuential (but not automatically guaranteed to be)
*can do regression calculations with & without the points to see how much intuence they have
removing high-leverage points
● lower than line & slope negative: slope closer to 0 & y-int lower
● lower than line & slope positive: slope steeper & y-int lower
● higher than line & slope negative: slope steeper & y-int higher
● higher than line & slope positive: slope closer to 0 & y-int higher
● and sometimes effects on r, r2, or s (use the point to evaluate this)
removing outliers
● impacts r, r2, and s values heavily
● usually makes them go up because strength of association & Ft of LSRL way higher
● doesn’t generally impact LSRL though

3.3 transforming to achieve linearity


applying a function to a quantitative variable (changes the scale of measurement) in order to
make the scatterplot more approximately linear (in order to use linear regression methods)

transforming with powers & roots (for power models: y = axp)


option 1: raise values of x to power p
graph (xp, y) (it will be linear)
option 2: pth root of values of y
graph (x, p√9) (it will be linear)
when p is known: use the above methods
when p is unknown: guess & check OR use log (more universal & works for unknown power
models)
transforming with logs (for power OR exponential models)
apply log transformation (log10 or ln)

for power models (y = axp): use a log-log (both variables)

for exponential models (y = abx): take log of y-var


:
to choose a model: most random scatter (tiebreaker highest r2 value)

types of mcqs (or frqs)


● effect of outliers / high leverage points on measures of strength or the LSRL
○ what would happen when they are removed?
● can you infer causation based on correlation? (no) (might not be worded like this
directly though)
● which is explanatory & which is response?

types of frqs
● interpret a feature of the association / regression line
○ slope: for every [increase / decrease] in one [unit of x], there is a predicted
[increase / decrease] in [units of y]
○ y-int: when the [context of x] is 0 [units], the predicted value of [context of y]
would be [y-int]
○ r-value / correlation coedcient: the correlation coedcient of [r] indicates that
there is a [strong / moderate / weak], [positive / negative] correlation between
[context of x] and [context of y]
○ s (standard deviation of residuals): on average, the model mispredicts [context
of y] by [s units] using the LSRL
○ r2 value: [r2] percent of the variability in [context of y] can be explained by the
linear association with [context of x]
○ residual plot: the residual plot [is randomly scattered / has a pattern], indicating
that a linear model [is / is not] appropriate
○ LSRL in general: for every increase in [1 unit of x], there is a predicted increase in
[b units of y] above / below a [unit of y] of [a – the y-int]

tips / common errors


● be very careful with predicted vs actual values
○ remember to add a hat on top of any predicted value!
○ and ALWAYS SAY THEY’RE “PREDICTED”
● when deFning LSRL: always deFne variables & add units for x & y
○ particularly important when data has been transformed to get the LSRL
● with transformed data: be mindful of units & always convert back to “regular” units
where appropriate / needed!
● with residuals: pay attention to whether the question is asking for predicted residual
(from LSRL / LSRL equation) or actual residual (from residual plot / graph)!
● can’t go backwards with LSRL & predict x value given a deFnite y value (can Fnd x value
from y-hat though)

how to tell if linear model is a good Ft: high r2 value, s is small relative to the data

chapter 4: collecting data


content overview
4.1 sampling & surveys
sampling: selecting a random group of people out of a whole population (that’s representative of
the population)
● uses resources more edciently than a census
sampling frame: the group of members from the population from which we select our sample
sampling survey: collects data from the individuals in the sample (to learn about the population)
:
types of sampling
random sampling: involves a chance process to determine which individuals are in the sample
● SRS: every group of n individuals has an equal chance of being selected
○ label individuals with numbers, random number generator, select individuals that
correspond
○ *sample without replacement! don’t include repeated numbers
○ calculator: math -> prob -> randintnorep(1, N) OR use table d
● stratiGed: SRS selected from each strata
○ strata: group w similar characteristics assumed to be associated with the
variables being measured
○ ensures you get some from each strata (more precise & accurate estimates)
● clustered: randomly selecting entire clusters
○ clusters: diff responses between (hopefully representative of population)
○ no statistical advantage but resource-edcient
● systematic: randomly select starting point & select every kth individual after
○ make sure no patterns coincide with your systematic pattern
○ allows you not to have to have identiFers (eg. names) for all the individuals in
the population (useful w unknown population size)
*decide on sampling type based on population & variable & resources available to you
bad sampling
● convenience sampling: individuals who are easy to reach
● voluntary response sampling: allows individuals to choose to be in sample
○ leads to voluntary response bias
○ individuals who feel strongly / have similar opinions more likely to respond
bias: likely to systematically overestimate or underestimate the value
● undercoverage: certain individuals less likely / cannot be chosen in a sample
● nonresponse: individual chosen for sample can’t be contacted / doesn’t participate
○ *diff from voluntary response bias bc this is only after sample has been
selected
● response: systematic pattern of inaccurate answers to a survey question
○ ex. wording or order of questions
selection biases: voluntary, convenience, undercoverage
the other two: nonresponse, response

4.2 experiments

observational studies: observes individuals & measures variables of interest (does not intuence
responses)
● retrospective: existing data
● prospective: tracks into future
● pros: ethics
● cons: confounding variables (cannot determine causality with observational studies)
experiments: imposes a treatment on individuals & measures their responses
experiment vocab
● placebo: no active ingredient
● treatment: condition imposed on individuals
● experimental unit: individual to which treatment applied
○ subject: human experimental unit
● factor: explanatory var that’s manipulated (may cause change in response var)
● levels: diff possible values of a factor
○ treatments formed using levels of each of the factors (in a multifactor
experiment)
● pros: can establish causality bc avoids confounding
:
● cons: resources, ethics
important vocab:
confounding: when variables are associated so that their effects on a response variable can’t be
distinguished from one another

designing experiments
required in a good experiment
● control group: provides a baseline for comparison
○ not always required but there does need to be some sort of comparison
○ avoids confounding variables
● random assignment of treatments
○ avoids confounding variables
● control: all other variables constant
○ avoids confounding variables & reduces variation
● replication: use enough subjects (diff in effects can be distinguished from chance
variation)
optional: placebo effect
● double blind: neither subjects nor the ppl measuring know the treatment
● single-blind: only one of the groups (above) knows
● triple blind: statistician doesn’t know either
experiment designs
completely randomized design: experimental units assigned to treatments completely at random
randomized block design: random assignment within each block
● block: group of experimental units known to be similar in some way that could affect
their response to the treatments
○ form blocks based on: confounding variable / the variable that’s the best
predictor of the response variable
matched pairs design: a type of RBD where blocks are pairs
● pairs of similar experimental units
○ either true pairs & randomize treatment
○ or each individual receives both treatments in a random order

4.3 using studies wisely


statistical signiGcance: observed diff is larger than can be attributed to chance alone
● ensured by random assignment of treatments
● can establish causality
statistical inference: generalising results to population
● assuming sample is representative of population (ensured by random sample)
ethical data gathering: informed consent, beneFt
more vocab
● sampling variability: diff random samples (same size, same population) produce diff
estimates
○ larger sample sizes produce more accurate estimates (closer to true value)

types of mcqs (or frqs)


● is it an observational study or an experiment?
:
types of frqs
● describe the type of bias present
○ describe how members respond differently
○ then describe how this leads to overestimated / underestimated values
● describe a confounding variable
○ describe that variable’s association with both explanatory AND response vars
● describe an appropriate experiment design
○ describe how to randomly assign treatments
■ create groups AND deFne which group gets which treatment
○ describe why certain experiment designs would be preferable
■ randomized block design: it accounts for variability in [response var]
created by [block var]
○ draw a diagram to explain the experiment design

tips / common errors


● describing how to use a random number generator / table: ALWAYS account for
repeated numbers
● if you tip coins: make sure everyone tips a coin (not just until you have one group &
then put the rest in another. this would not be random sampling)
● be really careful not to mix up language for experiments & language for observational
studies!

chapter 5: probability
content overview
5.1 randomness, probability, and simulation
random process: generates outcomes purely by chance
unpredictable in short-term but predictable in the long run
probability: likelihood of an event to happen
proportion of times an outcome would occur in the long run
law of large numbers: more trials means proportion approaches true probability (more
accurate)
simulation: imitates random process such that simulated outcomes are consistent with real-
world outcomes
options: random number generator, random number table, tipping coin, draw cards, etc

5.2 probability rules


probability model: description of a random process that includes a list of all possible outcomes
& the probability for each outcome
to determine theoretical probability
sample space: list of all outcomes
options: chart, table, venn diagram, probability tree
sum of all probabilities is 1 (or 100%), each probability is between 0 & 1 (or 0% & 100%)
notation & vocab:
event: any collection of outcomes from a random process
P(A) = (number of outcomes in event A) / (total number of outcomes in sample space)
complement: the probability that an event does not occur
P(AC) = 1 – P(A)
:
intersection: P(A and B) = A ∩ B (both A and B must be true)
joint probability
union: P(A or B) = A ⋃ B (at least one–either A or B, or both–must be true)
mutually exclusive events: cannot occur simultaneously (no outcomes in common)
(also known as disjoint)
if P(A and B) = 0
addition rule: P(A or B) = P(A) + P(B)
non-mutually exclusive events: can occur simultaneously
use a two-way table!
general addition rule: P(A or B) = P(A) + P(B) – P(A and B)

5.3 conditional probability & independence


conditional probability: probability that an event happens given that another event is known to
have happened: P(A | B)
formula: P(A ∩ B) / P(B)
tree diagrams are useful with conditional probability
independent events: knowing whether or not one event has occurred does not change the
probability that the other event will happen
P(A | B) = P(A | BC) = P(A) OR P(B | A) = P(B | AC) = P(B)
Gnding intersection (general multiplication rule): P(A and B) = P(A ∩ B) = P(A) * P(B | A)
multiplication rule for independent events: P(A) * P(B)
mutually exclusive vs independent: mutually exclusive are automatically not independent (if one
happens, the other is guaranteed not to happen)

recap: probability strategies

strategy when it’s useful what it looks like

simulation when they tell you to use it spinners, dice, beans,


cards, random number
generator

sample space cards / coins / etc. table of outcomes


list of outcomes

venn diagram consolidating two pieces of info (esp with


intersections)

two-way table consolidating two pieces of info (esp with note: usually given to you
intersections) & you have to interpret it

questions about independence

tree diagram successive events

formulas not given that much info

info in a format that’s easy to plug in


:
types of mcqs (or frqs)
● interpret probability
○ the proportion of times an outcome occurs in the long run
● given certain past outcomes, what is the probability of the next one being [outcome]?
○ the probability doesn’t change from trial to trial
● calculate probability
○ write out all your work! don’t just do it on your calculator

types of frqs
● explain how to perform a simulation
1) clear explanation of how to use a random process to perform one trial
(including what will you record at the end of the trial)
2) perform many trials
3) use results to answer question of interest

common errors
● describing simulation: remember to say that repeated numbers will be ignored!
○ for random number table: remember that numbers need to be the same length
digit-wise
● concluding that events are independent / mutually exclusive / there’s an association
○ specify that it’s based on the sample!!!
● pay attention to whether you are sampling with replacement or without replacement!
○ if without replacement & calculating probabilities: remember to adjust all
numbers accordingly
● not independent means associated (BUT DON’T USE THE WORD DEPENDENT)
chapter 6: random variables & probability
distributions
6.1 discrete & continuous random variables
random variable: takes numerical values that describe the outcomes of a random process
capital letters: random variables (C), lowercase letters: speciFc values of the variables
(ci which corresponds to probability pi)

discrete: Fxed set of values with gaps between them


can be described using probability distributions & histograms (each bar a value)
histogram models population distribution of the quantitative variable
● shape (using histogram)
● center
○ usually use mean (expected value / µx)
■ multiply each value by its probability & add up
■ interpretation: average amount (does not need to be an integer)
○ but to Fnd median: smallest value greater than or equal to 50th
percentile
● variability

○ standard deviation (σx):


■ interpretation: for all [population in context] OR if many
[individuals in context] are randomly selected: [variable in context]
typically varies from [σx] from the mean of [µx]

continuous: any value in an interval on the number line


probability distribution: density curve
for continuous: intervals of outcomes have probability while individual outcomes (exact
values) have no probability

6.2 transforming & combining random variables


transforming (Y = a + bX)
:
adding & subtracting constants: same as regular rules for quantitative variables
multiplying & dividing by a constant: same as regular rules for quantitative variables
combining (X + Y or X – Y)
● means: add / subtract
● variances: add ONLY IF independent random variables
independent random variables: knowing the value of one variable doesn’t change the probability
distribution of the other
putting it together: aX + bY (linear combination of random variables)
● mean: aµx + bµy

● standard deviation (if independent):


● any linear combination of normal random variables is also normally distributed
use transforming and combining appropriately depending on the context!
X1 + X2 is not the the same as 2X

6.3 binomial & geometric random variables

binomial random variables


count of successes in a binomial setting
binomial distribution: probability distribution of a binomial random variable
use acronym BINS to check for binomial setting
● binary: outcomes either success or failure
● independent: independent trials
○ 10%: can treat individual observations as independent as long as n < 0.10N
(sample size less than 10% of population size)
● number: Fxed number of trials in advance
● same probability: same probability of success p on each trial
mean: np
standard deviation:
calculating binomial probabilities: use formula sheet OR calculator functions
binompdf(n, p, x): exactly x
binomcdf(n, p, x): less than or equal to x
large counts condition: expected counts of successes & failures at least 10
np ≥ 10 and n(1-p) ≥ 10
probability distribution of binomial random variable X ~Normal
more normal as n increases and p is closer to 0.5
geometric random variables
number of trials it takes to get a success in a geometric setting
geometric distribution: probability distribution of a geometric random variable
geometric setting: number of trials it takes to get a success
no Fxed number of trials
but trials must be independent & probability p of success must remain the same
calculating geometric probabilities: use formula sheet OR calculator functions
geometpdf(p, x): exactly x
geometcdf(p, x): less than or equal to x
describing a geometric distribution
● shape: always unimodal & skewed right with peak at 1 (bc that’s the most likely value)
○ probability of each successive success decreases by a factor of (1–p)
:
● center: µx = 1/p

● variability: σx =
types of frqs
● calculate mean or standard deviation of a random variable
○ write out formula (use ellipses)
○ then can use calculator (L1 values, L2 probabilities) & 1-var stats to calculate it
● calculate binomial or geometric probabilities
○ see content overview
● describe a binomial or geometric distribution
○ use SOCV and the formulas for mean & sd

chapter 7: sampling distributions


content overview
7.1 sampling distributions
parameters & statistics
● unbiased estimator vs biased estimator
○ mean of sampling distribution of a statistic equal to true value of parameter
population distribution: distribution of the values of the variable in the population
sampling distribution: distribution of a statistic in all possible samples of the same size from the
population
● vs. distribution of sample data: values of variable for all individuals in a sample
● center (mean) & variability (smaller samples mean higher variability)
bias: statistics consistently do not march parameters
same as accuracy
check center (unbiased estimator)
variability: widely scattered statistics
same as precision
check variability
choose an estimator with low bias & low variability

7.2 sample proportions


sampling distribution of the sample proportion
● how statistic p-hat varies in all possible samples of the same size from the population
● closely related to binomial count X
conditions: large counts (~normal) & 10% (independence)
describing: SOCV
● shape: ~normal if large counts
○ can use normal calculations to solve problems
● center: µp-hat = p (unbiased estimator)

● variability: (if 10% satisFed for both populations)


○ decreases as n gets larger & increases as p approaches 0.5
○ typical distance between sample proportion (p-hat) & population proportion (p)
sampling distribution of a difference between two sample proprortions
for two different populations (independent SRSs)
● shape: ~normal if large counts for both populations
● center: µp-hat 1 – p-hat 2 = p-hat 1 – p-hat 2
:
● sd: add variances & take sqrt (if 10% satisFed)
○ “the difference (p-hat 1 – p-hat 2) in sample proportions of [successes] in
[population] typically varies by about [sd] from the difference in true proportions of
[mean].”

7.3 sample means


sampling distribution of the sample mean (x-bar)
conditions: large counts (~normal) & 10% (independence)
describing:
● shape: ~normal if normality condition met
○ two ways: population distribution is normal OR CLT met (n ≥ 30)
● mean: µx-bar = µ (unbiased estimator)

● sd: σx-bar = (if 10% satisFed)


○ decreases as n gets larger
sampling distribution of a difference between two sample means
for two different populations (independent SRSs)
● shape: ~normal if both populations satisfy one of the conditions
● mean: subtract
● sd: add variances & take sqrt (if 10% satisFed)

types of frqs
● deFne population parameter
○ make sure to deFne it accurately! use the words “all” and “true” and might need
to describe how the individuals “would” respond (not “did” respond)
● describe sampling distribution
● use sampling distribution to evaluate a claim
○ evidence that claim is true (provide statistic), possibilities (claim is true or
false), would it be surprising to get that statistic value based on the sampling
distribution?

tips
● be very careful with the language you use!
○ clearly specify which distribution (population, sample data, sampling
distribution)
● sampling distribution of the SAMPLE [statistic name]
● if sampling without replacement: actual sd smaller than formula suggests
● true difference in proportions NOT difference in true proportions

chapter 8: confidence intervals for


proportions
content overview
8.1 conFdence intervals: the basics
conGdence interval: an interval of plausible values for an unknown population parameter based
on sample data
● accounts for sampling variability & increases conFdence that our parameter value is
correct
● but trade off between the info in the interval & the amount of conFdence you have
point estimate ± margin of error
conGdence level: success rate / capture rate of the method that produces the interval
:
percent of intervals that will capture the true parameter value / intervals will capture true
parameter value in C% of all possible samples
to decrease moe:
● decrease conFdence level
● increase sample size n
moe only accounts for chance variation, not other sources of error (eg. bias)

8.2 estimating a population proportion & 8.3 estimating a difference in proportions


see below for the four-step process (choose, conditions, calculate, conclude)
conditions: random (generalize to population & independence), 10% (independence), LLC
(~normality)
when violated: random (no point), LLC (too narrow), 10% (too large)
critical value: table A or invNorm
REMEMBER: based on the tail area, so half it (eg. for 95% conFdence, look for the value
with 0.025 area to its left)
note: standard error is when you use p-hat instead of p (based on the sample, not the population)
for experiments, conditions change slightly
● random: can’t generalize to larger population (only population like the one in the study)
○ usually volunteers
● 10%: don’t need to check unless randomly selected without replacement from the
population
● LLC: no change

types of mcqs (or frqs)


● given moe, estimate required sample size
○ plug into moe equation & solve for n (always round up)
○ if no given p-value, use best guess or 0.5 (maximum so most conservative)

frq answer formulas


conGdence levels
In C% of all possible samples, the interval computed from the sample data will capture the true
[parameter in context].
conGdence intervals
We are C% conFdent that the interval from ___ to ___ capures the true population [parameter in
context].
assessing claims about a parameter (using a conFdence interval)
Since there are plausible values of p [less than / greater than] [value] in the conFdence interval,
___ [does / does not] provide convincing evidence that [claim about parameter in context].
margin of error
In a C% conFdence interval, the distance between the point estimate & the margin of error will be
less than the margin of error in C% of all possible samples.

tips
● Always interpret parameters correctly! It is often conditional tense, NOT past tense.
● Types of problems: conFdence intervals, assess claims about a parameter using a
conFdence interval, given moe Fnd the minimum required sample size)
● APPROXIMATELY NORMAL (when describing the function of the LLC)
● for two-sample: deFne which mean is which & order of subtraction AND “difference in
true proportions” (not “true difference in proportions”)

solving problems
choose: inference procedure, parameter, conFdence level
:
conditions: LLC, 10%, random sample
calculate: conFdence interval
conclude: in context
1-sample z-int for p: conFdence interval for one proportion
2-sample z-int for p: conFdence interval for difference between two proportions
calculator functions: 1-prop z-test and 2-prop z-test (double check this)

chapter 9: significance tests for


proportions
content overview
9.2 tests about a population proportion
power of a test: probability that a test will Fnd convincing evidence for Ha when a speciFc
alternative value of the parameter is true (probability that you avoid a type II error)
increasing power
● n larger
● ɑ larger
● effect size is larger (null and alternative values farther apart)
● making wise decisions in collecting data
sample size for a study depends on
● signiFcance level (risk of type I error we’re willing to accept)
● effect size (how large a difference is important to detect)
● power that we want
note: beta is prob of type II error

9.3 tests about a difference in population proportions


*use pooled proportion in checking the LLC and calculating the standard error!
p-hat = (sum of successes) / (sum of trials) OR ((n1p-hat1) + (n2p-hat2)) / (n1 + n2)

these signiFcance tests don’t always match up with the conFdence interval bc you’re using the
pooled proportion

frq answer formulas


conclusions (interpreting signi/cance tests)
two parts: Interpret p-value & then state conclusion
Assuming that H0 is true (p = __), the probability of obtaining results as extreme as or more
extreme than what we observed in our sample is [ P-value ]. Thus since these results [ are / are
not ] signiFcant at the α = [ ] level, we [ reject / fail to reject ] H0 as we [ do / do not ] have
convincing evidence that the true proportion of [ parameter in context ] is [ state alternative
hypothesis – greater than, less than, or differs from ___ (value of H0) ].

state type I and II errors in context & provide a potential consequence of each

types of mcqs (or frqs)


● link between conFdence intervals and signiFcance tests
○ conFdence intervals would generally lead you to the same conclusion as a two-
sided test at the (100 – ɑ)% signiFcance level
○ but they provide more information (a range of plausible values) as well
:
tips
● choose
○ always interpret parameters correctly! it is often conditional tense, NOT past
tense.
○ for two-sample: deFne which proportion is which & order of subtraction
○ hypotheses ALWAYS for population parameters
● conditions
○ APPROXIMATELY NORMAL (when describing the function of the LLC)
○ use P0 to check LLC (bc you’re assuming Ho is true)
● calculate
○ ALWAYS need to list the standardized test statistic
■ (p-hat – p0) / (standard error using p0)
○ and make a drawing of normal distribution
○ remember to double the p-value for two-sided tests! COMMON MISTAKE
● conclude
○ NEVER ACCEPT THE NULL (can accept alternative though)
● trade-off between type I and type II errors (as probability of one increases, probability of
the other decreases) (so consider consequences for each before selecting signiFcance
level)
1-sample z-test for p: signiFcance test for one proportion
2-sample z-test for p: signiFcance test for a difference between two proportions
calculator functions: 1-prop z-int and 2-prop z-int (double check this)
choose: inference procedure, parameter, hypotheses, signiFcance level
conditions: LLC, 10%, random sample
calculate: sample statistic, standardized test statistic, p-value
conclude: in context

chapter 10: confidence intervals for means


content overview
10.1 estimating a population mean
x-bar ± t*(sx / n)

t-distributions
● family of curves (diff one for each df)
○ df: n-1
○ more and more normal as n & therefore df increase
○ use table B (round down to closest if exact df not there) OR invT
● because using sx to estimate for σ adds too much variability
○ t* values are larger to account for it (bc t-dist have more area in tails)
conditions: random (generalize to population & independence), 10% (independence), normal
condition (~normality)
when violated: random (no point), normal (too narrow), 10% (too large)
the normal condition can be satisFed in one of three ways:
● central limit theorem (n ≥ 30)
● population distribution approximately normal
● distribution of sample data approximately normal
○ plot to check for strong skewness or outliers
○ always include this plot on the FRQ as proof
all the same concepts about moe still apply
conFdence intervals & conFdence levels: see above answer formulas & tips
:
10.2 estimating a difference in means
same as usual
df: either use tech (2SampTInt) or choose the smaller df (more conservative)
µdiff (mean difference): for paired data / matched pairs designs
summary statistics: x-bardiff and sdiff
1-sample t-int for mean diff OR paired t-int for mean diff
conditions: same
but 10% and normal focused on number of differences, not individuals (usually
not independent bc pairs, but differences should be independent)

overview
1-sample t-int for μ: conFdence interval for one mean
2-sample t-int for μ: conFdence interval for difference between two means
calculator functions: t-int and 2-samp t-int
and invT (gives you the t-score for a speciFc distribution. plug in tail value)
choose: inference procedure, parameter, conFdence level
conditions: normal condition, 10%, random sample
calculate: conFdence interval
● use moe formula and / or calculator functions
conclude: in context

tips
● for two-sample: deFne which mean is which & order of subtraction
○ although order of subtraction does NOT matter in terms of the results you get
● Always interpret parameters correctly! It is often conditional tense, NOT past tense.
● APPROXIMATELY NORMAL (when describing the function of the normal condition)
● for normal condition: when plotting sample data, always include the plot on the FRQ as
proof that the condition is met
● technically, if you knew σ, you would use z* values (but this never happens)
types of problems:
● conFdence intervals
● assess claims about a parameter using a conFdence interval
● given moe Fnd the minimum required sample size (use a known reasonable estimate
for σ instead of sx & use z* instead of t* – these account for issues you might encounter)

chapter 11: significance tests for means


content overview
signiFcance tests: see above answer formulas & tips
1-sample t-test for μ: signiFcance test for one mean
2-sample t-test for μ: signiFcance test for a difference between two means
paired t-test for mean difference: signiFcance test for mean difference
(matched pairs experiment design)
calculator functions: t-test and 2-samp t-test
tcdf also useful for signiFcance tests (plug in standardized test statistic as the bound)
choose: inference procedure, parameter, hypotheses, signiFcance level
conditions: normal condition, 10%, random sample(s)
calculate: sample statistic, standardized test statistic, p-value
conclude: in context
:
11.1 testing claims about means
additional info
● beware of the difference between statistical & practical signiFcance
● beware of p-hacking

types of mcqs (or frqs)


● link between conFdence intervals and signiFcance tests
○ conFdence intervals would generally lead you to the same conclusion as a two-
sided test at the (100 – ɑ)% signiFcance level
○ but they provide more information (a range of plausible values) as well

tips
● for two-sample: deFne which mean is which & order of subtraction
○ ESPECIALLY deFne order of subtraction for µdiff
● choose
○ always interpret parameters correctly! it is often conditional tense, NOT past
tense.
○ for two-sample: deFne which mean is which & order of subtraction
○ hypotheses ALWAYS for population parameters
● conditions
○ APPROXIMATELY NORMAL (when describing the function of the normal
condition)
● calculate
○ ALWAYS need to list the standardized test statistic
■ (p-hat – p0) / (standard error using p0)
○ and make a drawing of normal distribution
○ remember to double the p-value for two-sided tests! COMMON MISTAKE
● conclude
○ NEVER ACCEPT THE NULL (can accept alternative though)
● trade-off between type I and type II errors (as probability of one increases, probability of
the other decreases) (so consider consequences for each before selecting signiFcance
level)
chapter 12: inference for distributions &
relationships
content overview
12.1 chi square tests for goodness of Ft
use for: categorical variables with two or more categories (not just success / failure)
to check whether a hypothesized distribution seems valid
● how likely it is to get a sample distribution for a sample size n that differs as much as
your sample does
● compares observed counts from sample with expected counts from hypothesized
distribution (if Ho is true)
choose:
● procedure: chi-square test for GOF
● hypotheses (Ho: hypothesised distribution true, Ha: at least one of the categories does
not match the value stated by Ho)
● signiFcance level
● state categorical var & population of interest (no parameter)
conditions: random, 10%, LLC (all expected counts at least 5)
calculate:
● expected values for each category (npi)
● calculate chi-square test statistic
:
○ measures how far observed and expected counts are, relative to expected
counts
■ always 0 or positive
■ large chi-square test statistic means more evidence for Ha

● once you have chi-square, compare to sampling distribution of chi-square


○ sampling distribution of chi-square: always positive & right-skewed
○ speciFed by df (categories - 1)
○ shape: as df increase, less skewness, larger values more probable
○ mean: always at df
■ mode: for df > 2, mode (peak) of chi-square density curve at df - 2
● get p-values using tech or table c
○ probability of getting a value of chi-square as large or larger when Ho true
○ 2nd -> vars -> x^2cdf(test statistic, upper bound, df)
conclude: same as usual!
follow-up analysis: if you reject the null, analyse how distribution is different from Ho
take a look at the categories that contribute the most to the test statistic (“components”)
describe how observed & expected differ in the category (& note direction of difference)

12.2 inference for two-way tables


chi-square test for homogeneity: compares distributions of a single cat var over multiple
populations / treatments (multiple independent samples)
choose:
● procedure: chi-square test for homogeneity
○ to Fnd if the diff are convincing evidence OR if they just occurred by pure chance
(and there is no diff)
● signiFcance level
● hypotheses (Ho: no diff in true distribution of [cat var] between populations /
treatments, Ha: there is a difference, at least one dist different)
conditions: random & independent samples, 10% for both samples / populations (when sampling
without replacement), LLC (all expected values greater than or equal to 5)
calculate:
● expected values
○ (row total * column total) / total total
● chi-square test statistic
○ add up test statistics for all cells in the table
○ use calculator to simplify this process
● Fnd p-value
○ df: (number of rows - 1)(number of columns - 1)
○ table c or technology (see notes for the steps if using technology)
follow-up analysis
chi-square test for independence & association: compares distributions of two cat var
(association) in a single population (one sample)
choose: same
● does association hold up in the larger population or did we observe just by chance?
● hypotheses: no association in the population / an association in the population
○ OR independent / not independent
conditions: random: single random sample, other two conditions: same
calculate:
● expected values (same)
○ can use principles of probability (if A and B independent P(A | B) = P(A) )
● chi-square test statistic & p-value (same)
:
conclude: same (include follow-up analysis if applicable)
● remember that association does not equal causation!
● convincing evidence of an association / not convincing evidence of an association
differentiating between the two:
● homogeneity: one set of totals initially known
● independence: only total total known
● best method: how the sample was selected / how data was produced
● homogeneity usually experiments, independence usually observational studies
● don’t get thrown off if a question asks about independence even for homogeneity and
vice versa
more tips / miscellaneous
comparing two proportions
● test for homogeneity based on 2x2 table: same results as two-sample z-test for p1 - p2
TWO SIDED
● larger than 2x2 table: only option is homogeneity
● 2x2 & one-sided hypothesis: use two-sample z-test for p1 - p2 (NOT chi-square test)
● 2x2 & conFdence interval for difference in proportions: only option is two-sample z-int
comparing quantitative data: can group to make it categorical
● gives you more info if you compare whole distribution with a signiFcance test (not just
mean)
what if my expected count is less than 5?
● collapse 2+ rows or columns together so that it is met

types of mcqs (or frqs)


● statements about the chi-square distribution

types of frqs
● conduct a follow-up analysis

tips
● always write out the general formula AND the test statistic!!

the calculator functions


Calculator Functions for the AP Stats Exam

notation guide
samples
x̄ - sample mean
sx - sample standard deviation
p̂ - sample proportion
n - sample size
population
µ - population mean
σ - population standard deviation
p - population proportion
N - number of individuals in population
sampling distributions
proportions
µ of p-hat
σ of p-hat
means
:
µ of x-bar
σ of x-bar
standard errors
s or SE of p-hat: standard error of the sample proportion
s or SE of x-bar: standard error of the sample mean
linear regression
samples
ŷ - predicted y
b - slope from sample data
a - y-int from sample data
population
y - y-value of true linear regression line
β - slope of true linear regression line
ɑ - y-int of true linear regression line
tests
ɑ - probability of making a type I error (alpha-level)
β - probability of making a type II error (beta-level)
power: probability of avoiding a type II error (1 - β)
:

You might also like