Stat Notes
Stat Notes
Unit1 :
Examining Distributions Classifications of data:
categories no numbers
,
Statistics :
set of methods for obtaining
,
.
I
,
↑
to ordering makes sense for the values of
information. ordering
of categorical variable.
sample : subset of individuals/units in a population a
a course ,
. . .
.
ex ·
final exam score
, height ,
volume of airinballoon , sum of
has decimals
usually ,
take
·
:
value within given range
can
any a .
eX
weight age distance,etc
, ,
.
onlycertain values
not finite.
,
Data Distribution
Techniques to data set :
display a
a variable takes and how often it takes thesebar charts displays values on : one axis , frequencies on other.
values· Note here "value" doesn't have to be quantitative
:
,
. note: spaces between bars to
ex .
"blue" is a value of the variable "eye colour . "
not imply continuity.
our data values fall into various pre-determined the observed values for a
categorical
choose ourselves so we
get "nice" of data
summary
·
a
variable.
first interval includes lowest data value minimum
·
note : each interval includes left endpoint,but not the bars to reflect continuity of data
light ie . . To 80 .
frequency
Relative frequency or proportion : values between O and I ,
or relative frequency.
inclusive and
·
Can
getRelative frequency
or proportion of individuals in each interval by dividing the
number of data values in each interval by the total number of should be of equal length
data values.
h
Relative frequency distribution frequency can be used to characterize the shape
·
: FD RED : : n =
relative frequency. of
a data distribution.
ime lots: displays values over period of time two approximate mirror
images.
ex Seasonal variation
·
·
skewed to the right When the right side of histogram
:
decreasing :
$ even though the graph increases
variability
I
There are 3 measures of location : examples
Mode : the most
.
data set
·
.mostine en
2 : the middle value in an ordered data set
mean of total
individuals
formulas :n
·
NX tha, ...
Weighted Mean :
calculated when some data values are given more
weight than others.
WiX :
because some values are (in some sense) more important than others.
it data value
·
Mean Vs Median : .
in a
symmetric distribution mean and median are
equal ght skewed median
·
,
VI :
mean
in skewed distribution the mean is closer to thetail left skewed mean median
·
a .
:
,
Numerical Summaries : features of a data set location and
:
variability
T
-
examples
vary
Variability :
·
one sample
also called spread
will differ from another sample
1 Range .
:
R maximum-minimum
=
. Interquartile
2
Range : measures the middle 50 % of the ordered observations.
does NOT consider
·
outliers
of data set is the ordered observation where at least 25 % of data values are
25thpercentile as smaller or smaller and at least 75 % of the values are as large or larger.
third quartile Q3) of data set is the ordered observation where 75 % of data values are
·
: at least
75th percentile as smaller or smaller and at least 25 % of the values are as large or larger.
second quartile is the median.
·
:
Percentiles : the pt percentile of data set is the ordered observation where at least p % of data values are as small or
range :
Q1 take median
·
:
of all data values lower in position than the median.
·
IQR Q3-Q=
-
mu mum
Boxplots :
there are 2 kinds
1 Quantile Boxplot
.
:
consists of :
·
lines whiskers that extend from the box out to the min and
·
max
.
can be
displayed horizontally vertically
·
either or
right skewed :
left skewed :
symmetric :
two
Side by Side Boxplots :
used to compare multiple data distributions .
plotting same
enable us to , ,
:
same quantile box plot range for fielders is
higher
for Q1 median , Q3
,
instead , construct a "fence" with lower fence ((F) and upper fence (UF) values
=
-
.
= -
1 .
-
= .
=
.
-
whiskers from left side of box extends to lowest data value that's still (f
greater than
whiskers from right side of box extends to highest data value that's still less than UF
·
are as axis
·
outliers are
any data points that full outside the fences
-
loose definition
"Xi ** -
formula S : = i= S2 (value
= ,
-
means + (value ...
-
"
)
mean
n 1
n 1
-
2 calculaten deviations Xi -*
3
square the deviations xi-x)
2
-
5 divide n-
by
* Xi *-
formula : S= s = i=
n 1
-
NOT the
average deviation from the mean even
though the variance is
loosely
defined as the
average squared deviation from the mean.
Numerical Summaries
Summary
shape of distribution which numerical summary to report Why
Examining relationships between several quantitative variables measured on the same individuals .
·
variables for the same population explanatory variable X-axis if neither, choice of
2
One of the variables
may explain or predict
the other. 1) direction :
negative slope negative association
straight line .
close to the
strong relationship all points fall line
doesn't require response/explanatory variables
weak relationship points fall far from the line
=
an observation may be outlying in either
X -X
yy
or y-directions or both
X- ,
.
3) -z) (yi
*
multiply deviations X:
-
j)
* 4) add n products
...
(xi -
= )(yiy) Association doesn't imply causation.
5) divide
by n-1 SxSy Association Vs Causation .
:
cannot conclude that
properties i)
y-value to/ (could be smth didn't measure
·
:
positive indicates positive association ,
while X-variable causes . we
ii has no units so ,
changing units of XkY has no effect on it. the value ar randomly "assigned to sample units
is .
iii) makes no distinction between X and Y Exp/Resp. Variable aren't if there's still X caused Y
.
necessary. a
strong + association we ,
can then say
iv
only measures strength of linear relationship because we diversified
away the Similarities Within the
groups.
v NOT resistant to outliers
Linear Regression Residuals comparisonbtwn the actual Y value
:
Y
the predicted value y
Regression Line : reflects the error of our prediction for any X-value .
residual for the ith observation
always has explanatory X and response variables Yi-Yi =
when .
X=0
"least
11
2)
squares regression line bivariate falls outside the pattern of
:
points
little effect on
regression line .
in least squares
regression . I re O Influential :
an observation if removing it would dramatically
we don't know the other factors that accounts for alter the position of the regression line and the value of r
remaining Variation
.
if V = -I or 1 ,
then r = I
meaning regression on X accounts Categorical Variables on Scatterplot
predict y exactly for any
·
given X. relationships.
meaning regression on X
if r = 0 , then r = O tells us
nothing ex .
about
percentvariationuse
we should be careful when relationship to ensure that the data belongs
examining a
to
only one population. In this case , make separate regression lines.
1) five number
summary-min Q1 , median Q3 max provide descriptions
3
, , ,
both
·
distribution report outliers and/or skewed : five number
·
if the of data values is
reasonably symmetric with no outliers ,
summary
the mean and SD since it carries more into about the sample than the median IQR , and range
,
.
variance and standard deviation
·
·
for skewed distributions or if there are outliers report the five number summary. outliers. They are more effected
, by mean .
Unit 3 :
Sampling and Experimental Design Simple Random Sampling : SRS of sizen consists of n
examine . If it isn't ,
we cannot infer
any conclusions.
GOOD
best
Stratified Random Sampling :
used when our population is
Voluntary Response Sample :
when people who choose to naturally divided into strata . A stratum is group a of similar individuals
include themselves into the sample by responding to a within each of the strata , We take an SRS of sized i.
question survey
. note : total sample is not SRS as it doesn't match either definition.
the data in this sample represents the opinions of those who feel Can select a sample size from each stratum proportional
about the subject. to it's population size. This would allow for each individual in
strongly
biased and BAD population has the same chance to be selected still not SRs
When it
systematically
favors certain outcomes
don't have to choose same amount from each stratum.
over others
Sampler
Convenience Sample : when survey or chooses individuals GOOD ,
bias more
strategically eliminated than SRS
that will influence the respondent. Results are obtained "fabricated" note : total sample is not SRS as it doesn't match either definition
GOOD
of interviewer
refuse to answer the question(s). Systematic Random Sampling Start with numbered list of all N individuals
remember when
selecting a
'good' sample. because not
every group of n is
likely
to
good (non-leading
, easy to understand wording of questions be careful when there is a pattern in population list
GOOD
·
Observational Study :
simple measures values of variables On individual ·
Randomization :
distinction between
explanatory/response variables is necessary give treatment to as many individuals as possible for reliable results .
on
·
block : a
group of experimental units or subjects that are similar in ways that are
:
.
combo of factor levels applied to unit
# of treatments-fle xfle ,
...
factory factor (v)
·
factors :
explanatory variables in experiment
·
Statistics estimators
·
discrete would be
rolling 2 dice or :
a number computed from sample data .
1) lies
the curve strictly on or above the x-axis proportions Can't be negative
*
a) the total area underneath the curve and above the X-axis is equal to one
1
:
this are a represents the proportion of all values within that interval ?
3)
the curve represents a
proper function for each x-value there is a unique y value area :
(base) (height) = 1 > X
4
ex . P1 6(X(3
.
.
3 = (3 3 -1
.
. 6 (0 25.
= 0 425
.
ex .
27 % of the time the person spends more than how many minutes in the shower ?
area of interval
proportion of values in interval = total area under curve =
1
0 2 .
p(xx() p(x10 = =
0 27
. 10 c 0 2
-
. =
02
.
73 i
Por tion
of values in interval area of interval
0 .
8 .
10 c
-
=
0
22
.
.
=
1 35
.
C =
10 1 35 -
.
=
8 65 .
& C10
some variable X .
median point
·
:
on the X-axis with are O S under the .
curve on each side . area = 0 5 .
(base) (height) = 1
X
O 5
to be valid
density curve :
mean of continuous distribution the balance point along the X-axis A = 0 S(s)(h) 1 1 = n 4
:
.
. = + 2 Sn
.
= = 0 .
if skewed ,
mean is closer to tail .
·
bell-shaped
Normal Distribution
1
.
·
sample mean :
Y :
most common/important
combos of M and o
"-"
the observation from the X follows normal distribution (N)
average squared deviation of of the distribution .
:
an mean a
denoted
affects
height : low st is tall
,
highsd is short by 2 :
# of d an observation X is from M
X -
M
z =
2
Normal Distribution continued The 68-95-99 7 Rule .
:
:
PX(x =
p(X(x) + p(X =
x) 68 % of all values full within I standard deviation of the mean M 1
+
95 %
·
99 7 % of all
. values full within 3 standard deviations of the mean M 3
+
Table 1
·
:
rephrase the problem so it
only involves areas to the left.
Comparing Normal Variables
proportions less than -3 49 and
greater than
·
. 3 49 are like
. zero.
if exact proportion isn't on table , take closest value of calculate z-score for both normal variables so that
they've on same scale
rare) if proportion is
exactly halfway ,
take
average of two values on either side. now we can compare which is relatively "better or worse"
finding
·
bEC)
finding
·
1) P 2 b table
entry corresponding to z= b Step 1 find such that P(2) given
=
:
z = value
2)
P2b 1- table
entry for b
x M
-
-
3)
Pb I c table entry forc-table entry for b Step 3 Check P(Z) with found value if it matches value.
o
given
:
4)
Pl b =
=
0
finding M
·
x -
M
work backwards Step 2 Solve for
: : z =
e M = x -
20
1)
P[<2) value Step 3 Check P(Z) with found M value if it matches value.
given
=
:
2) P(z >z) =
value
3)
interquartile range
first Q1 :
is z such that P(z(z) =
0 2500
. : z = - 0 67
.
1) P2>2) = 0 05/2 .
= 0 025
.
2) P -
0 95/2
.
+0 .
68/2 = 0 815
.
finding PCX)
·
DOZL3 = 0 19
.
Step 3 Check
:
any
subset of outcomes in the sample space
.
contained in the event
a random ↑
:
times the outcome would occur in an infinitely long series of trials. Complement A of event A is the event
consisting of all outcomes in the
ahead of time but can be described sample space which are not contained in A
the outcome cannot be predicted by .
a
regular pattern only happens after many repeated trials
basically the opposite/everything else
a phenomenon is random if individual outcomes are uncertain Probability distribution gives the values of :
some variable and the
of each value .
BUT there is a
regular distribution of outcomes in
large It of repetitions probability
if X-N(M 2) we
say X follows a normal probability distribution
:
, ,
Experiment any
·
Random and
:
process/activity in which there is
uncertainty
has two or more possible outcomes.
relates to future events S = Sall values of s such that a 203 for ex . time it takes for ...
outcomes
to (even there are possible).
·
outcomes in S
assign probabilities to the ,
a
like for
ii)
for
a
probability each outcome
phenomenon
·
11
eX .
if we toss a coin three times :
suppose S =
E , .. . . .,
n) where the prob of outcome i is Pi
ii + Pat 1
p ,
...
Pn =
sampling distribution
·
the idea of
repeatedly taking samples of the same size n from the When his large ,
the of sample mean is approx normal :
.
11
②
.
X =
Nu ,
n
three characteristics
·
↑
of the distribution of X is the same the of for this
·
the population of X
z
the standard deviation of
·
is :
n
sd is lower be the
averages are less variable than is the population of X normal ?
individual observations .
·
yes no/idk
is n230 ?
I exactly
·
if distribution of distribution of is
X is skewed , the
sampling
will approach a normal distribution as the sample size increases. normal for
yes no
mean will always be approximately normally distributed when sample size
is sufficiently large .
of X
·
mean =
M
So of = F
, average before
Unit 6 : Confidence Intervals 95 % Confidence Interval :
conclusions about a population from sample data. if X 2sd of M then M is within 2sd of
·
is within ,
.
z
95 % of all samples M lies within I2
that the population is fairly represented in .
n
accurate believe estimate to be. 1) the true value of M falls within this interval .
we our
actually
2)
this is one of the rare samples 5% that produces an
n
·
construct an interval of values to estimate population mean . estimate : our best guess at the true value of M
in
way that M is in the interval for Most samples. of error reflects how accurate believe our estimate to be
a
margin we
-
:
PC-2*< 2 ) C
·
we'd like to be confident that the interval we construct where z is the value of I such that <
z
*
=
gives
a :
,
the probability that the interval will capture the true value of M .
1)
If we were to take repeated samples of n and compute the
O
A
a
2)
in the
long run , c of similarly constructed intervals would contain
Critical values
*
values that mark off specific area und
·
: 2
by factor of K reduces
·
increasing n ,
1) ·
add area of confidence (vI with area of what's left over on the left. if we reduce moe
by factor of I ,
need sample K2 as
large.
2
doesn't matter when
population
estimating
·
3) X+ 2 * =
&
assuming equaled for two pops ,
a C confidence interval for pop #
round up
always n
Cautions :
2) confidence interval is
*
strongly influenced by outliers.
3) we use the true population standard deviation o in our calculations but this is NOT a realistic assumption .
4)
the margin of error
only reflects the natural variation in the sampling distribution of X
.
Can NEVER PROVE that parameter has any specific value null Ho
·
verifying substantially
·
When take sample and sample u needs to be the test to assess the of evidence Ho
designed strength
:
is
against
different than the claim to believe it
always expressed as an equality in terms of the population parameter u
.
No is claim
hypothesis Ha
·
of the test
the
probability
:
2)
Statement of hypothesis uppertail
a sample mean as extreme as the one observed if Ho were true .
Mo
Ho the lower the P-value the
M claim VS is P(z(z)
stronger evidence
(against Ho)
:
=
.
Ha :
MMo ,
our or
lower tail
ii)
Ha : M Mo P(zcz) in favour of the alternative
-double
HaiM # Mo 2P(2x21) interpretation : "if Ho
iii) tail area
probability value of the
was true , the of
observing a
3 high ou 3)
Statement of the decision rule rejection rule sample mean at least as extreme as we did would be p-value
"reject
·
"
Ho if P d =
given level of
significance
4
Calculation of the test statistic level of
·
significance
C the value compare the p-value to
: 0 1 0 05 , 0 01
we .
.
,
. .
the hull
provides a measure of the
compatibility between the maximum P-value for which the null hypothesis will be rejected
·
Conclusion
1) What conclude 2)
we Why
reject
·
"
sufficient
level of significance ,
we have insufficient evidence that Ha is true .
Two Sided Tests
the
observing
·
the P-value of
is
probability a value of the test
find P-value
by doubling the probability to the left/right of a whichevera
bc were interested in value of sample mean far fromMo in either direction
we
being
Step 2) Ho :
M =
Moversus is Ha : M >Mo
sided test
·
T -
M -
use use
·
z E
t Distribution
M
is
standard error of sample mean
if X-N(M 2) then the variable T ,
=
estimated standard deviation
P-value
8
shape : similar but slightly greater than the standard normal curve (2 distribution). :
suppose Ho :
M= Mo VS Ha : .
M Mo
as n increases
P(T (15) 53)
·
less area near center , more in tails than standard normal distribution table 2
says PCT(1S)
= 0 .
691) = 0 25.
and 0 691 is the lowest value in the row.
.
alt ,
where is the upper critical for +(n-1) distribution Paired Data
question ·
each pair ·
interpretation stays same when data is collected in pairs
·
question
1
would contain the true mean context
.
M two diffvariables are measured for each individuals we want the . differences between
Tests of
Significance/Hypothesis tests characteristic are made under different conditions (orat different times).
3)
similar individuals are placed in pairs and each member of the pair then receives
uppertail different treatment. The same response variable is measured and compared for
·
:
a
lower tail Matched Pairs t Procedures used to detectestimate differences between responses to
·
:
:
the two treatments by making one comparison for each of the n pairs
two-sided
·
parameter Ma
·
: :
,
the true mean of all differences of all pairs in the population
·
difference and
we estimate the population mean difference Ma
by sample mean
and the population standard deviation of differences
& by samplesd of differences Sd
:
ex. for any given car ,
mileage for methods to construct confidence intervals and hypothesis tests the same.
differences -
Independent :
variables are unrelate d ex for any two cars
-
-
, premium
·
test statistic :
population proportion :
sample proportion :P :
1
p = -
sample sizen 2) Ho :
p
=
vs .
Ha: p
1 = estimate/predict 3)
reject Ho if p-value -d =
-
mean of F :
M
,
=
P 1) calculate = for test statistic
standard deviation of p :
c
= Pp'p z =
- Po
PoCIP
5) calculate p-value
Sample Size 6) conclusion
Z = variablemean -
NCO , think of p as a kind of sample men is
sowhennis high p ,
:
Np , ** 2 . P N10 1 ,
Summary :
F -
P
to sample to use normal distribution 1)
enough compared knowp and want probabilities for p z
:
PC-P)
=
- 3) sample size : n
·
n :
n
*
however , we don't know port ,
so we use p
usually assume p* =
0 5 unless
. specified otherwise like
"
different versions