0% found this document useful (0 votes)
5 views23 pages

Stat Notes

The document provides an overview of data classification, focusing on categorical and quantitative data types, including their definitions and examples. It discusses data distributions, techniques for displaying data, and numerical summaries such as measures of location (mean, median, mode) and variability (range, interquartile range). Additionally, it covers graphical representations like boxplots and scatterplots to analyze relationships between variables.

Uploaded by

cswf5xq592
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views23 pages

Stat Notes

The document provides an overview of data classification, focusing on categorical and quantitative data types, including their definitions and examples. It discusses data distributions, techniques for displaying data, and numerical summaries such as measures of location (mean, median, mode) and variability (range, interquartile range). Additionally, it covers graphical representations like boxplots and scatterplots to analyze relationships between variables.

Uploaded by

cswf5xq592
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

j

Unit1 :
Examining Distributions Classifications of data:

categories no numbers
,

Statistics :
set of methods for obtaining
,
.

Categorical data represents values :


of

organizing summarizing data


,
.
categorical variables aka. qualitative variables
Data: Comes from characteristics measured on that place individuals into of those
one
groups.
individuals or units can be people , animals places
things
, ,
ex
-
gender of baby ,
reason for taking course , eye color, far to show

I
,

Population the totality


:
of individuals about which we want Categorical and Ordinal : if there is logical


to ordering makes sense for the values of
information. ordering
of categorical variable.
sample : subset of individuals/units in a population a

that is actually examined in


it
order to gather information . ex :

Placing in tournament 1st


,
2nd 3 , ,
etc .)

want " next Service fair, poor


ex 1000 voters are asked which candidate
.

they'll support in election. rating good ,

letter grade in A+, A Bt, F


·

a course ,
. . .
.

Variable : a characteristic or property of an individual


Categorical and nominal : if there's no
logical
eX· time until
light bulb runs out anything ordering ordering doesn't make sense .
distance travelled
·

by taxi driver in one


day ex
-
gender of baby ,
reason for taking course , eye color, far to show
has to be
whole #
number of heads in five coin tosses. have to be able to do math with it.
·

hair colour , . Quantitative data


2 :
represents values of
your grade in this course.

quantitave variable for which arithmetic operations make sense.

ex ·
final exam score
, height ,
volume of airinballoon , sum of

the numbers shown on two rolled dice.

Classifying Data/Variable Types


Types of Quantitative Variables :

has decimals
usually ,

Continuous can't be whole .

take
·
:
value within given range
can
any a .

eX
weight age distance,etc
, ,
.
onlycertain values
not finite.
,

Discrete : can only take a


·

countable number of values


max 31
eX
. # of children in a month ,
family the of days of rain in ,
a and the

highest denomination of bill in someone's wallet.


j

Data Distribution
Techniques to data set :
display a

Distribution of a data set tells us what values :

a variable takes and how often it takes thesebar charts displays values on : one axis , frequencies on other.
values· Note here "value" doesn't have to be quantitative
:
,
. note: spaces between bars to
ex .
"blue" is a value of the variable "eye colour . "
not imply continuity.

Frequency Distribution : a count of now


many ofPieCharts : visual representation of the relative
frequency of

our data values fall into various pre-determined the observed values for a
categorical
choose ourselves so we
get "nice" of data
summary
·
a

classes or intervals . should


usually choose 5-10 intervals
equal length ·
We
be .
.

variable.
first interval includes lowest data value minimum
·

last interval includes highest data value maximum)


·

histogram: a form of bar


graph with no spaces between

note : each interval includes left endpoint,but not the bars to reflect continuity of data
light ie . . To 80 .

to ensure that the distribution reflects of data.


·
for quantitative data
frequency continuity
height represents the
·

frequency
Relative frequency or proportion : values between O and I ,
or relative frequency.
inclusive and
·

decimal representations of fractions. base each


,
are
rectangle represents
length of interval (all intervals
·

Can
getRelative frequency
or proportion of individuals in each interval by dividing the
number of data values in each interval by the total number of should be of equal length
data values.
h
Relative frequency distribution frequency can be used to characterize the shape
·
: FD RED : : n =
relative frequency. of

a data distribution.

symmetric if its center divides into


:
·
must leave in order !

ime lots: displays values over period of time two approximate mirror
images.
ex Seasonal variation
·
·
skewed to the right When the right side of histogram
:

extends much further than the


left side.

decreasing :
$ even though the graph increases

in some interval , it decreased overall


time
-
Numerical Summaries : features of a data set location and
:

variability

Location : determined by where the center of ourdata falls.

I
There are 3 measures of location : examples
Mode : the most
.

frequently observed data value


not best option
·

data set
·

can have more than 1 mode.

.mostine en
2 : the middle value in an ordered data set

outliers do NOT affect value of the median .


·
extreme values aka , ,

resistant to the effect of outliers.


for
say the median
this reason we is
,

To find the median :

order the data from smallest to


largest n+1
2
Countin" ,
the number of data values individuals, and compute 2
n+1
3
Count 2 data values up from the lowest value .
n+1
if h is odd , median is in position 2

if n is even , median is the


average of the two values on both sides of

* median : it is the data value in that position .

info since we all data values from sample


gives us more use

. Mean the "center of mass" or "balance point" of the data


3 : .

strongly affected by an extreme value . It is NOT resistant to outliers.

units of are same as units of X

mean of total
individuals

formulas :n
·

NX tha, ...

Weighted Mean :
calculated when some data values are given more
weight than others.

the fact that some


may be due to
·

values are observed more


frequently or

WiX :
because some values are (in some sense) more important than others.

it data value
·

XoW , where Wi is the


weight given to the

Mean Vs Median : .

in a
symmetric distribution mean and median are
equal ght skewed median
·

,
VI :
mean

in skewed distribution the mean is closer to thetail left skewed mean median
·

a .
:
,
Numerical Summaries : features of a data set location and
:

variability

T
-

examples
vary

Variability :

·
one sample
also called spread
will differ from another sample

1 Range .
:
R maximum-minimum
=

heavily affected by (outliers)


·

most extreme values

. Interquartile
2
Range : measures the middle 50 % of the ordered observations.
does NOT consider
·

outliers

first quartile (Q1)


·

of data set is the ordered observation where at least 25 % of data values are

25thpercentile as smaller or smaller and at least 75 % of the values are as large or larger.
third quartile Q3) of data set is the ordered observation where 75 % of data values are
·
: at least

75th percentile as smaller or smaller and at least 25 % of the values are as large or larger.
second quartile is the median.
·
:

Percentiles : the pt percentile of data set is the ordered observation where at least p % of data values are as small or

smaller and at least 100-p) % are as


large or
larger.
To find the interquartile
·

range :

Q1 take median
·
:
of all data values lower in position than the median.
·

Q3 : take median of all data values


higher in position than the median.
·

IQR Q3-Q=
-

. Five Number Summary


3 :
includes the minimum , QI ,
the median Q3 , and the maximum
,
.

good numerical description of the distribution of data valuess


·

describes the center shape and spread of , ,


our data.
·

used to make a box plot


.

mu mum
Boxplots :
there are 2 kinds

1 Quantile Boxplot
.
:

consists of :
·

a line at the median ,

box that covers IQR


·

lines whiskers that extend from the box out to the min and
·

max
.

can be
displayed horizontally vertically
·

either or

allows us to characterize the shape of the distribution :


·

right skewed :

left skewed :

symmetric :

two
Side by Side Boxplots :
used to compare multiple data distributions .

used for the variable for samples


representing different populations.
·

plotting same

compare the distributionsWith respect to center shape and spread.


·

enable us to , ,

must be constructed with uniform scale to the comparison.


a
legitimize

. Outlier Box plots


2 :
also based on
quantiles but takes outliers into account .
median
heights and IQR are equal

consists of the middle three lines (box) from


·

:
same quantile box plot range for fielders is
higher
for Q1 median , Q3
,

values (if outliers).


·

Whiskers do NOT extend to max and min

instead , construct a "fence" with lower fence ((F) and upper fence (UF) values

(F Q1 1 5 (IQR) Q1 5(Q3 Q1)


·

=
-

.
= -
1 .
-

UF Q3 + 1 5 (IQR) Q3 + 1 5 (Q3 Q1)


·

= .
=
.
-

whiskers from left side of box extends to lowest data value that's still (f
greater than
whiskers from right side of box extends to highest data value that's still less than UF
·

whiskers extend to "new" max and min values wo outliers

do NOT include vertical lines at ends of whiskers .

outliers included points along


·

are as axis
·

outliers are
any data points that full outside the fences

-
loose definition

Sample Variance aka Variance "average squared : deviation" from mean.

units are squared units of the observations

"Xi ** -

formula S : = i= S2 (value
= ,
-
means + (value ...
-
"
)
mean

n 1
n 1
-

steps to solve: I calculate sample mean I

2 calculaten deviations Xi -*

3
square the deviations xi-x)
2
-

4 add the squared deviations X: -

5 divide n-
by

Standard Deviation : the positive square root of the Variance


·

units are the same as the individual observations


2

* Xi *-

formula : S= s = i=

n 1
-

to solve square root the variance


:

NOT the
average deviation from the mean even
though the variance is
loosely
defined as the
average squared deviation from the mean.

Numerical Summaries
Summary
shape of distribution which numerical summary to report Why

mean and stader.

symmetric with mean and standard carry more into about a

no outliers deviation sample than five num.

outliers and / min , Q1 , median, Q3 ,


may

or skewed five number


summary are not affected
by outliers
Unit 2 :
Examining Relationships
Scatterplots displays the values of two different
:

Examining relationships between several quantitative variables measured on the same individuals .
·

variables for the same population explanatory variable X-axis if neither, choice of

two situations : response variable y-axis axes is arbitrary


We're simply interested in the nature of the
What to look for When
relationship.
examining a scatterplot :

2
One of the variables
may explain or predict
the other. 1) direction :
negative slope negative association

Explanatory Variable : denoted by X helps explain


:
positive slope positive association
the outcome of a study
.
Response Variable : denoted by :
takes values 2) form type of relationship.
:
ex. linear

representing the outcome of a study


.
3
if there is no
explanatory response variable we are Strength determined by how close the points lie
·
:
,

simply interested in the nature of the relationship .


to it's simple form ex a .

straight line .
close to the
strong relationship all points fall line
doesn't require response/explanatory variables
weak relationship points fall far from the line

Coefficient of Correlation numerical measure : of points appear randomly scattered .


strength and direction oflinear relationship the scale
·

can affect how the strength appears.

between twoquantitative variables.


4
Outliers several types of outliers for bivariate data.
·

no info about the Casual nature of the


:
gives relationship .

=
an observation may be outlying in either
X -X
yy
or y-directions or both
X- ,
.

when an observation simply falls outside


ii
·

Steps to solve: 1) Calculate , Sx Sy the


, ,

general pattern of points


2) calculate deviations Xi-X and
yi-y .

3) -z) (yi
*
multiply deviations X:
-

j)
* 4) add n products
...
(xi -
= )(yiy) Association doesn't imply causation.

5) divide
by n-1 SxSy Association Vs Causation .
:
cannot conclude that

properties i)
y-value to/ (could be smth didn't measure
·
:
positive indicates positive association ,
while X-variable causes . we

Lurking Variable : one that helps


·

negative indicates negative association .


explain the relationship
ii) falls between -Land 1 ,
inclusive . between variables in a
study ,
but isn't included in the study
.
; weakIvs close to to avoid perform experiment instead of observational
strongers :r closeto-1 : zero .
:
study when

ii has no units so ,
changing units of XkY has no effect on it. the value ar randomly "assigned to sample units
is .

iii) makes no distinction between X and Y Exp/Resp. Variable aren't if there's still X caused Y
.
necessary. a
strong + association we ,
can then say

iv
only measures strength of linear relationship because we diversified
away the Similarities Within the
groups.
v NOT resistant to outliers
Linear Regression Residuals comparisonbtwn the actual Y value
:
Y
the predicted value y

Regression Line : reflects the error of our prediction for any X-value .
residual for the ith observation
always has explanatory X and response variables Yi-Yi =

residual indicates that the point falls above line .


Cunlike correlation +
regression
given X ,
it's used to predict corresponding.
*A
-

residual indicates that the point falls below


regression line .

slope of regression line : the predicted increase in


y when X * it is the sum of squared residuals that is minimized in

increases by 1 unit calculating the least squares regression line .


b =
usb=
intercept of regression line : the predicted value of y
·

when .
X=0

Extrapolation the process of predicting


: a value for a value
bo y
=
bX,

of X outside our of data


range
the line that minimizes the sum of squared deviations in the
·

vertical directio n. ... Y: Y ;


Types of Outliers
·

estimate of the "true line" is


y bo + bex 1) outlier in the
y-direction
slope little effect on
predicted value of intercept regression line .

"least
11
2)
squares regression line bivariate falls outside the pattern of
:
points
little effect on
regression line .

3) an outlier in the X-direction

Coefficient of Determination v2 : the fraction of


strong effect on regression
line .

variation in7 that is accounted for by its


regression on X an influential observation

in least squares
regression . I re O Influential :
an observation if removing it would dramatically
we don't know the other factors that accounts for alter the position of the regression line and the value of r

remaining Variation
.

if V = -I or 1 ,
then r = I
meaning regression on X accounts Categorical Variables on Scatterplot
predict y exactly for any
·

all variation in sometimes two or distinct


for and we can a scatterplot may actually be displaying more

given X. relationships.

meaning regression on X
if r = 0 , then r = O tells us
nothing ex .

about

percentvariationuse
we should be careful when relationship to ensure that the data belongs
examining a

to
only one population. In this case , make separate regression lines.
1) five number
summary-min Q1 , median Q3 max provide descriptions

3
, , ,
both

of the center and


variability
2) Sample mean and standard deviation of a data distribution

How to decide which one to report?


·
it depends on the SHAPE of the data distribution
·

symmetric With no outliers : mean and SD

·
distribution report outliers and/or skewed : five number
·
if the of data values is
reasonably symmetric with no outliers ,
summary
the mean and SD since it carries more into about the sample than the median IQR , and range
,
.
variance and standard deviation
·

are NOT resistant to

·
for skewed distributions or if there are outliers report the five number summary. outliers. They are more effected
, by mean .
Unit 3 :
Sampling and Experimental Design Simple Random Sampling : SRS of sizen consists of n

individuals from the chosen in


-
population a
way that every group

Sampling the process of collecting data from


: a sample to of n individuals has an
equal chance to be the sample selected .

make statements about the population . :


each individual has an equal chance to be selected into sample .
needs to be representative of the population we wish to like putting names in a hat and pull out n of them but with computer software.

examine . If it isn't ,
we cannot infer
any conclusions.
GOOD

best
Stratified Random Sampling :
used when our population is
Voluntary Response Sample :
when people who choose to naturally divided into strata . A stratum is group a of similar individuals

include themselves into the sample by responding to a within each of the strata , We take an SRS of sized i.

question survey
. note : total sample is not SRS as it doesn't match either definition.

the data in this sample represents the opinions of those who feel Can select a sample size from each stratum proportional

about the subject. to it's population size. This would allow for each individual in
strongly
biased and BAD population has the same chance to be selected still not SRs
When it
systematically
favors certain outcomes
don't have to choose same amount from each stratum.
over others
Sampler
Convenience Sample : when survey or chooses individuals GOOD ,
bias more
strategically eliminated than SRS

who are easiest to reach .

biased and BAD Multistage Sampling often


:
used when we need
large numbers in close
proximity to one another.
first SRS Canadian cities then SRS
ex ·

, neighbourhoods in those cities then SRS


,

Leading Question : when the question is asked after


stating something city blocks in each
neighbourhood then survey occupants of allhouseholds
,
there .

that will influence the respondent. Results are obtained "fabricated" note : total sample is not SRS as it doesn't match either definition

GOOD
of interviewer

bias not as SRS


wording tone/brevity, race
,
can a respondent's answer .
,
good as ,
sometimes our only option due to time/cost.

some respondents will or lie .


guess

Nonresponse When respondents


:

refuse to answer the question(s). Systematic Random Sampling Start with numbered list of all N individuals

Undercoverage When some :


units in the population have no chance of in the population . To select sample of n individuals , randomly select a number from

being included in the sample. 1 to k where


,
K =. The sample will consist of the 1th individual on the list and

1individual after that. The value is the .


every sampling interval
note : total sample is not SRS
·

remember when
selecting a
'good' sample. because not
every group of n is
likely
to

selecting sample in unbiased and representative manner . be chosen even


though each individual has same chance
proper interview training advantages easy/fast to select and we don't need
: a list of entive population .

good (non-leading
, easy to understand wording of questions be careful when there is a pattern in population list

GOOD
·

try to make non-response a non-issue

include all population units in


your possible sample
voluntary response and convenience samples aren't appropriate Census "sample" of entire population
:

. Ideal type but usually impossible due to time cost

if you can do better than SRS , do it.


Experimental Design Principles of experimental design :

Observational Study :
simple measures values of variables On individual ·

Randomization :

randomly assign experimental units to various treatments.

nothing was done randomly .


only distinction between treatment groups should be the treatments.

Confounded two variables explanatory or lurking


: when Groups receiving each treatment must be similar with respect
their effects on the response cannot be separated . to all other variables.

"mixed up" and most often the result of an observational


study .
·

Control of the effects of


lucking variables on the response by comparing
several treatments , one of which may be a control treatment

Completely Randomized design :


a special type of experiment where Control Group :
the one in which we compare our treatment of interest .

the various treatments received


all units are
randomly assigned to receive can be the
group for which no/"fake" treatment is or the

simplest form of design . (known) treatment is received


group to which a standard

only absolutely necessary when we


only have one treatment to examine

to eliminate the effect of


any potential lurking variables
Experiment : imposes treatment on an individual to observe
·

Replication : the administration of each treatment to more than one unit to

their response to the treatment. reduce variation in the results .

distinction between
explanatory/response variables is necessary give treatment to as many individuals as possible for reliable results .

can examine the effect of several variables on the response


variable. this allows us to examine
any interaction
among factors.
units are not selected randomly they must volunteer
,
but we need to Double-Blind Experiment :
neither subject nor the person

view them as representative of the population they're from


administering the treatment knows which treatment is the one being applied.
comparison of treatments is eliminates bias administrator of treatment
leading principle experimental
in o
a
by
design .
Placebo Effect : a dummy treatment that is known to

Blocking is to experimental design as stratification is to sampling. have no


physical effect.
Randomized Block Design the random
assignment of treatments to units it have beneficial effects.
: is
may psychological
carried out ethics : must tell subjects they could
separately within each block.
get the placebo
of further ensure the exclusion of the effect of
uses principle blocking to
lurking variables . experiment vocab :

experimental unit the individuals which the experiment is performed


·

on
·

block : a
group of experimental units or subjects that are similar in ways that are
:

subjects if the individuals are people


·

expected to effect the response to the treatments . :

NOT formed randomly


·

treatment : specific set of experimental conditions applied to units ; the

.
combo of factor levels applied to unit

# of treatments-fle xfle ,
...
factory factor (v)
·

factors :
explanatory variables in experiment
·

factor levels different values


: of the factors
Summary of Good experimental
Design
Unit4 Density Curves and
: Normal Distribution
·

Parameters : a number that describes an entire population .

usually unknown bo we are


dealing with very large populations

Density Curve the curve describing the overall shape of


:
a distribution values such as M can be +or
-

and G must be positive

of continuous variables for "infinite" large populations .


time can be
any number
.
like
height ,

Statistics estimators
·

discrete would be
rolling 2 dice or :
a number computed from sample data .

used to estimate the values of parameters

Three properties of all density curves: values such as X and S

1) lies
the curve strictly on or above the x-axis proportions Can't be negative
*
a) the total area underneath the curve and above the X-axis is equal to one
1

Uniform Distribution for some variable Xo


·

:
this are a represents the proportion of all values within that interval ?

3)
the curve represents a
proper function for each x-value there is a unique y value area :
(base) (height) = 1 > X
4

ex . P1 6(X(3
.
.
3 = (3 3 -1
.
. 6 (0 25.
= 0 425
.

ex .
27 % of the time the person spends more than how many minutes in the shower ?
area of interval
proportion of values in interval = total area under curve =
1
0 2 .
p(xx() p(x10 = =
0 27
. 10 c 0 2
-
. =
02
.

73 i

Por tion
of values in interval area of interval
0 .

8 .

10 c
-
=
0
22
.

.
=
1 35
.
C =
10 1 35 -
.
=
8 65 .

& C10

Triangular Distributions for


·

some variable X .

median point
·

:
on the X-axis with are O S under the .
curve on each side . area = 0 5 .
(base) (height) = 1
X
O 5
to be valid
density curve :

mean of continuous distribution the balance point along the X-axis A = 0 S(s)(h) 1 1 = n 4
:
.
. = + 2 Sn
.
= = 0 .

if skewed ,
mean is closer to tail .
·
bell-shaped

Normal Distribution
1
.
·

sample mean :
Y :
most common/important

population mean true mean :


M . characterized by parameters mando > X

does NOT affect distribution family of normal distributions : set of all


height of normal distributions the infinite amount of

combos of M and o

Variance of continuous distribution :


measure of spread and variance if normal variable X has a normal distribution with Manda : XN u ,
2

"-"
the observation from the X follows normal distribution (N)
average squared deviation of of the distribution .
:
an mean a

sample standard deviation : s area determined by # of standard deviations an observation is


away from its mean .

population standard deviation : O


"Sigma"
population variance 0 :
·

Standard Normal Deviation. when normal distribution has M O and G I


= =

denoted
affects
height : low st is tall
,
highsd is short by 2 :
# of d an observation X is from M

X -
M

to find z-score aka standardized value of X :


Z= o

decimal places : when we


area/proportions 4 know
everything
·

z =
2
Normal Distribution continued The 68-95-99 7 Rule .

applies to all normal distribution.

for continuous distributions approximately


·

:
:

=> O b) it is the area of a line ·

PX(x =
p(X(x) + p(X =
x) 68 % of all values full within I standard deviation of the mean M 1
+

95 %
·

of all values full within 2 standard deviations of the mean +


M 2

99 7 % of all
. values full within 3 standard deviations of the mean M 3
+

Table 1
·

provides proportions understandard normal curve to the left of z.

:
rephrase the problem so it
only involves areas to the left.
Comparing Normal Variables
proportions less than -3 49 and
greater than
·

. 3 49 are like
. zero.

if exact proportion isn't on table , take closest value of calculate z-score for both normal variables so that
they've on same scale

rare) if proportion is
exactly halfway ,
take
average of two values on either side. now we can compare which is relatively "better or worse"

finding
·

proportions under the standard normal curve :

bEC)
finding
·

for values b and C


any a
:

1) P 2 b table
entry corresponding to z= b Step 1 find such that P(2) given
=
:
z = value

2)
P2b 1- table
entry for b
x M
-
-

Step 2 Solve for


=
: : z =
2
=
e

3)
Pb I c table entry forc-table entry for b Step 3 Check P(Z) with found value if it matches value.
o
given
:

4)
Pl b =
=
0

finding M
·

finding such that P 22


·

I is equal to some specified value .


Step 1 find :
z such that P(2) given = value

x -
M
work backwards Step 2 Solve for
: : z =
e M = x -

20

1)
P[<2) value Step 3 Check P(Z) with found M value if it matches value.
given
=
:

body of table for proportion value to


Search determine
corresponding 2 value

2) P(z >z) =
value

search value of 2 with 1-(proportion value) ; P(2<) =


1-(proportion value examples : remember M = 0

3)
interquartile range
first Q1 :
is z such that P(z(z) =
0 2500
. : z = - 0 67
.
1) P2>2) = 0 05/2 .
= 0 025
.

to left is Q3= 0 67 bc P(-2 < z(2)


next : by symmetry value of area 0 75 0 95
=
,
z with . . .

last : IQR Q3 -Q1 = =


0 67.
-
(-0 67) .
= 1 34. and . P(2-2) P(z)2) = + 1 -
0 95
.
= 0 . 05

2) P -

z([(1 = P( z(20) + p(0 < z()) =


-

0 95/2
.
+0 .
68/2 = 0 815
.

finding PCX)
·

a such that is some


given value
Step 1: find 2 such that P(2-2) given value =
3) P(178(X(196) = P1818 XM 1968 =

DOZL3 = 0 19
.

Step 2 transform tox


: : z = x =
M + 20

Step 3 Check
:

P(X < found x) given value =


Unit 5 : Probability and
Sampling Distribution of Event :

any
subset of outcomes in the sample space

the Sample Mean it's


probability can be found
by adding the probabilities of all outcomes

.
contained in the event

Probability of any outcome of phenomenon the proportion of


·

a random ↑
:

times the outcome would occur in an infinitely long series of trials. Complement A of event A is the event
consisting of all outcomes in the

ahead of time but can be described sample space which are not contained in A
the outcome cannot be predicted by .

a
regular pattern only happens after many repeated trials
basically the opposite/everything else

a phenomenon is random if individual outcomes are uncertain Probability distribution gives the values of :
some variable and the

of each value .
BUT there is a
regular distribution of outcomes in
large It of repetitions probability
if X-N(M 2) we
say X follows a normal probability distribution
:
, ,

Experiment any
·

Random and
:
process/activity in which there is
uncertainty
has two or more possible outcomes.

Random Variable : a numerical description of the outcome of a

statistical experiment. ex the sum of the rolls of two dice


.

Proportion : a known or observed value

speak of it in present tense


Probabilities for Continuous Variables
Probability : a theoretical value of a proportion after infinitely
many
trials

relates to future events S = Sall values of s such that a 203 for ex . time it takes for ...

outcomes
to (even there are possible).
·

outcomes in S
assign probabilities to the ,
a

like for

Probability Model : a mathematical model to describe random behaviour


assign probabilities to intervals of values instead of individual values discrete

composed of the areas under


:

density curves represent probability


i) a list of possible outcomes

ii)
for
a
probability each outcome

Sample space S : the set of all possible outcomes of a random

phenomenon
·
11
eX .
if we toss a coin three times :

suppose S =
E , .. . . .,
n) where the prob of outcome i is Pi

the probabilities of the outcomes must two conditions :


satisfy
i
0 P: 1 for all i = 1. 2
. . ... n

ii + Pat 1
p ,
...
Pn =

does not need to be finite can use discrete and continuous


DISTUTUOfThSAMMana Central Limit Theorem :

sample mean falls within some


range of values take an SRS of size n from any population with mean u and Sd0 .

sampling distribution
·

the idea of
repeatedly taking samples of the same size n from the When his large ,
the of sample mean is approx normal :
.

11

and "approx follows


population calculating X ex
.
.


.

X =
Nu ,
n

three characteristics
·

sample size required depends on


og distribution .

the distribution of I is still normal


·

lower values ofa required for symmetric than skewed distributions


of the distribution of X is the same the of for this
·

the mean as mean course : apply CLT When ne 30

the population of X
z
the standard deviation of
·

is :
n

sd is lower be the
averages are less variable than is the population of X normal ?

individual observations .
·

yes no/idk

practice problems include "average"


·

is n230 ?
I exactly
·

if distribution of distribution of is
X is skewed , the
sampling
will approach a normal distribution as the sample size increases. normal for

the form of pop distribution of X doesn't matter be the sample


any
: .
n

yes no
mean will always be approximately normally distributed when sample size

is sufficiently large .

I is approx . I is not normal


ALWAYS TRUE for any distribution :
normal (CLT) :
nothing we can do

of X
·

mean =
M

So of = F

if asked about total convert to calculating


·

, average before
Unit 6 : Confidence Intervals 95 % Confidence Interval :

using 68-95 99 7 rule if -N


. .
:
M , , -95% of
Statistical Inference provides all values of I fall withinon
:
methods for
drawing of the mean .
M

conclusions about a population from sample data. if X 2sd of M then M is within 2sd of
·

is within ,
.

to make up for 95% confidence interval for


use the
lang. of probability uncertainty M :

z
95 % of all samples M lies within I2
that the population is fairly represented in .
n

foundation of inferences lies on


long-run predictable behaviour .
Only two possibilities
·

reporting sample mean alone givesNo info as to how

accurate believe estimate to be. 1) the true value of M falls within this interval .
we our
actually
2)
this is one of the rare samples 5% that produces an

interval which excludes the true value of M


.
estimate
Confidence Intervals :
our
of M

form of C for population
·

mean u : estimate margin of error =


X + 2*

n
·

construct an interval of values to estimate population mean . estimate : our best guess at the true value of M

in
way that M is in the interval for Most samples. of error reflects how accurate believe our estimate to be
a
margin we
-
:

PC-2*< 2 ) C
·

we'd like to be confident that the interval we construct where z is the value of I such that <
z
*
=

contains the value of the parameter we've


trying to estimate

confidence level C which interpretation


·

each confidence interval has


·

gives
a :
,

the probability that the interval will capture the true value of M .
1)
If we were to take repeated samples of n and compute the
O

interval in similar manner , then C of such intervals would contain

A
a

the true mean .


ex

2)
in the
long run , c of similarly constructed intervals would contain

the true mean ex


.

Critical values
*
values that mark off specific area und
·
: 2

standard normal curve . Sample Size


and obtain narrow confidence level
increasen to use
high
·

to find C confidence interval for M


margin of error by it
·

by factor of K reduces
·

increasing n ,

1) ·

add area of confidence (vI with area of what's left over on the left. if we reduce moe
by factor of I ,
need sample K2 as
large.
2
doesn't matter when
population
estimating
·

find corresponding z value size u; sample size does

3) X+ 2 * =

&
assuming equaled for two pops ,
a C confidence interval for pop #

4) state interpretation will have the of error as #2


same
margin pop
22
if m = 2 then n =
m

round up
always n
Cautions :

17 formula for confidence interval if data


only holds was collected
using SRS

2) confidence interval is
*
strongly influenced by outliers.

3) we use the true population standard deviation o in our calculations but this is NOT a realistic assumption .
4)
the margin of error
only reflects the natural variation in the sampling distribution of X
.

it doesn't reflect of other forms of bias.


any degree undercoverage nonresponse
, ,
or
Unit 7 :
Hypothesis Testing note we've not
:
concluding that the population mean is (the value of Ho),

sufficient evidence to believe claim


we are
concluding whether there is
Hypothesis Testing type of inference that helps us assess
:
:
note : there's certain chance that we're our conclusion
always a wrong in
the evidence provided by some claim concerning population .
·

a note : keep in mind the shape of distribution and sample size .

AND an outcome that'd occur if an assumption were true is sufficient


rarely
evidence that the assumption is not true why low probability provides evidence
* it's foundation : "If assumption true , how
our initial was
likely would the lower the probability ,
the
higher the confidence interval is
"
it be to observe an estimate this extreme ?

to determine whether there is evidence that supports claim made


·

about the value of some parameter M vocab


·

Can NEVER PROVE that parameter has any specific value null Ho
·

hypothesis the statement being tested of "no difference/effect "


a
:

verifying substantially
·

When take sample and sample u needs to be the test to assess the of evidence Ho
designed strength
:
is
against
different than the claim to believe it
always expressed as an equality in terms of the population parameter u
.
No is claim

hypothesis Ha
·

alternative the statement


trying to support
:
claim we're
making the
always expressed as an
inequality in terms of the population parameter u
.

steps of hypothesis testing/tests of significance


P-value
·

of the test
the
probability
:

1) State the level of


significance :
let z =
the lower thep-value the less likely it would have been to observe
,

2)
Statement of hypothesis uppertail
a sample mean as extreme as the one observed if Ho were true .
Mo
Ho the lower the P-value the
M claim VS is P(z(z)
stronger evidence
(against Ho)
:
=
.
Ha :
MMo ,
our or
lower tail
ii)
Ha : M Mo P(zcz) in favour of the alternative
-double
HaiM # Mo 2P(2x21) interpretation : "if Ho
iii) tail area
probability value of the
was true , the of
observing a

3 high ou 3)

Statement of the decision rule rejection rule sample mean at least as extreme as we did would be p-value

"reject
·
"

Ho if P d =
given level of
significance
4
Calculation of the test statistic level of
·

significance
C the value compare the p-value to
: 0 1 0 05 , 0 01
we .
.

,
. .

the hull
provides a measure of the
compatibility between the maximum P-value for which the null hypothesis will be rejected

hypothesis and our data . if P &


reject Ho in favour of Ha the lower the value of alpha , the
evidence we need to reject
G
2 Mo
the null (claim) if P
stronger
&
z =
on ,
assuming hypothesis is true & fail to rejectHo the null
hypothesis
test statistic

statistically significant : results that lead to the rejection of Ho


5
Calculate the P-value statistical significance it would
: an effect so
large , rarely occur by chance alone

if the conclude claim is true


probability is low ,
we can

·
Conclusion

1) What conclude 2)
we Why
reject
·

"Since the p-value C


,
We fail
to reject Ho At .
the d

"
sufficient
level of significance ,
we have insufficient evidence that Ha is true .
Two Sided Tests
the
observing
·

the P-value of
is
probability a value of the test

statistic at least as extreme in either direction


given that Ho true
=

find P-value
by doubling the probability to the left/right of a whichevera
bc were interested in value of sample mean far fromMo in either direction
we
being
Step 2) Ho :
M =
Moversus is Ha : M >Mo

Methods to solve two-sided test


1)
hypothesis test
Confidence Interval Method :

conditions 1) test must be two sidded


·

2) confidence level level of 1


+
significance
=

sided test
·

two rejectsHo if Mofalls outside the

100 (1-2)% confidence interval


Unit 8 Inference for M when
:
o is Unknown do I use Z or ? did the question provide o or s ?

until now , we used the unrealistic assumption that we know o do I know a ? 2

T -

M -
use use
·

if s is unknown then estimate it,


by SAMPLE standard deviation "s" : Ze S/
no
yes

z E
t Distribution
M
is
standard error of sample mean
if X-N(M 2) then the variable T ,
=
estimated standard deviation

P-value
8

shape : similar but slightly greater than the standard normal curve (2 distribution). :
suppose Ho :
M= Mo VS Ha : .

M Mo
as n increases
P(T (15) 53)
·

closer to2 curve


as on increases, shape gets n= 16 ; + =
0 53 ; P-value
.
: =0 .

less area near center , more in tails than standard normal distribution table 2
says PCT(1S)
= 0 .
691) = 0 25.
and 0 691 is the lowest value in the row.
.

+ statistic + =sin has t distribution with P-value must be


n 1
degrees of freedom :
our value is even lower,
greater than 0 25
-

P(T (24) = 4 58).


·

strongly influenced by outliers n = 2S it = 4 98 ;


. P-value is .

table 2 says P(T(24) < 3 745 .


= 0 00s . and that 3 74s is
.

highest value in this row .

the P-value must be less than 0


higher
:
our value is even ,
. 0005 .

Confidence intervals when m and o are unknown


·

alt ,
where is the upper critical for +(n-1) distribution Paired Data
question ·
each pair ·
interpretation stays same when data is collected in pairs
·

repeatedly took samples a


: if we of context
,

Paired data data that was observed


·

and constructed an interval in a similar manner, the Clevel of such intervals :


in natural pairs. occurred
by :

question
1
would contain the true mean context
.
M two diffvariables are measured for each individuals we want the . differences between

values of the two variables.


2)
each individual is measured twice the two measurements of the same .

Tests of
Significance/Hypothesis tests characteristic are made under different conditions (orat different times).
3)
similar individuals are placed in pairs and each member of the pair then receives

uppertail different treatment. The same response variable is measured and compared for
·
:
a

the two individuals in each pair .

lower tail Matched Pairs t Procedures used to detectestimate differences between responses to
·
:
:

the two treatments by making one comparison for each of the n pairs
two-sided
·

parameter Ma
·

: :
,
the true mean of all differences of all pairs in the population
·

assumptions differences follow normal distribution with mean Maand


: &d &d.

difference and
we estimate the population mean difference Ma
by sample mean
and the population standard deviation of differences
& by samplesd of differences Sd

Dependent variables are related


·

:
ex. for any given car ,
mileage for methods to construct confidence intervals and hypothesis tests the same.
differences -

premium gas vs . for


regular gas but now we're examining differences instead of individual observations N distribution

Independent :
variables are unrelate d ex for any two cars
-
-
, premium
·

test statistic :

mileages are independent t =


Ud-Mao ,waysofference
/
Unit 9 :

Sampling distribution and inference for proportions Hypothesis Tests HopPoste p

population proportion :

p # of individuals that (successes)


1) let =
My
characteristics
-possess
some

sample proportion :P :
1

p = -
sample sizen 2) Ho :
p
=
vs .

Ha: p
1 = estimate/predict 3)
reject Ho if p-value -d =
-

mean of F :
M
,
=
P 1) calculate = for test statistic

standard deviation of p :
c
= Pp'p z =
- Po
PoCIP

5) calculate p-value
Sample Size 6) conclusion

Z = variablemean -
NCO , think of p as a kind of sample men is

sowhennis high p ,
:
Np , ** 2 . P N10 1 ,

Summary :

if Up 10 AND nC1-p 10 then population is


large
:

F -

P
to sample to use normal distribution 1)
enough compared knowp and want probabilities for p z
:
PC-P)
=

2) confidence interval for p :


piza
Confidence Intervals (C)
margina of
= -p *

- 3) sample size : n
·

a100(1-2)% C for p is piz


2
4)
hypothesis test for p-test statistic : z =
"poc
= -P
for
large enough of specified pandm
·

n :
n
*
however , we don't know port ,
so we use p

usually assume p* =
0 5 unless
. specified otherwise like
"

"Suppose we believe the true proportion is close to 0 02 . . (p 02)


= 0 .

cannot be used to conduct hypothesis tests b) the formulas use


·
c

different versions

You might also like