Statistics Notes for University of Pretoria
Statistics Notes for University of Pretoria
Notes
Chapter : 12
BSERVED FREQUENCIES :
26 Esi25985 Sij
635 102 C 1s
:
INDEPENDANT
Pa(7) P(a)P(7)
=
=
(Row total) (Column total)
total
grand
W
total -
grand eij must be , 5
1 234567
F 26 657129 20 2831270
16 88 31 7327 68
2
. . .
M
% Go t 138 12515 2813 33
.
51910138
. .
72 0 as a 100
fij -lij
1 2 3456 7
22 == zij
F 26 65 71 29 20 2831 270 j
48 6
. 67 550 63
. . 2716 88 31 7327 68
. . .
11 5 19 10130 26 18 621 -
13 332
M 46
-
35
.
%
.
+... +
23 4
. 32 5
. 24 3
.
138 12515 2813 33
.
. . .6
8 13 33
.
Chi-Square table :
Ca
-
ST F1 DEPE DAC - :
Dependant
State Ho and Ha
Calculate test statistic : Ch
fij -lij 26 -
48 62 1 -
13 .
332
2
.
=
+... + = 62 668
.
eij .6
8 13 33.
=j
Rejection Critered
a .
CCv2 :
density curve
Decision
2 eject Ho at a 5%
0 05 .
Conclusion
12 592
.
dependent x
HYPOTHESIS TESTUSIGP-ALUE ·
:
Reject Ho .
-
X(e :
Big Samples ·
Small samples
· Quantitative data ·
No assumptions about
form of distribution
Categorical
·
or
quantitative
Ho : u g Ho : Median = 450
Ha :
M Ha : Median = 450
Ho :
M1 Ma Mann-Whitney-
Ha :
M1 M2 Wilcoxon test
M1 125
,
12 135
1 .
= -
(u1 ,
0 ? )
-
H :
M1 =
Mz
Ha :
M1 r
or
/2
12
-
12 ,
0?
---ST :
Criteria : Ho :
th Mz ...
UK
· k independant populations. K 2 Ha : -
cleast I mean differs
·
Random samples from each population
·
Response variable :
ormally dist 200 200
15002 15002 (2
(2
·
ariance of available responses same . (41 (41
300 300
145 145
It It
SS SSR 33
factors we are :
·
first estimate is indp of Ho ·
Second estimate based on assumption -
· -
S-
17
SS-
<
SS- ·
Mean
square due to Treatment:
k
2
S2SSIR SS12
1,, x
k1 j = 1
erall mean :
M
1 T 12
x Cij ...
Tk
77T
M11 Mrk
Mai
m nly if sample sizes =
j
~
variance
--
IABLE : Pop
E
ariati r Squares -
reedom S
SS df
S
-
r p SS T SS
S
T
I ta SS 2
SSTR
-
S2 k1
S SS-
11 k --
table
df d)
.
[USK- -
S -
S :
Sort small to
big
ranks in order mean ranks
2
Assign increasing .
hen ties ,
assigned.
3 Show rank
Sum ranks
Use chi table determine the
square distribution
to K- 1 of,
5 p-value , using
under the assumption of identical populations the distribution of H
,
sampling
can be approximated chi square distribution. Sample Sizes
k
2 R 5
317T
1+ 75 Ni 2 2 2
i=1
25 3 3
T number of a s 3
number f bservati
35 & 5
11 ; ns in sample i
T
6 6
1T ni = + tal number of observations in all samples
i = 1
5
6 G 9
2 sum fthe ranks for sample i 6 I 9
~
erf rmance -
valuatin ratings 6 I
7 2
>
-
3 E
7 2 2
25 3 6 S 5 3 2
7 2 2 2 2 5
6 & 3 & 6 I 5 .5
5
85 15 1 55 G .5
5
95 2 6 I .
5 5
& 5 35 S 2 I g
55 S I 9 I
95 2 2
EXAMPLE:
candies are often high in
calories. Assume that the following data show
the calorie content from samples of M&MÕs.
Kit Kat, and Milky Bar.
Test for significant
differences among the calorie content of
these three candies.
At a 5% level of
significance, what is your conclusion?
Sort small to
big
ranks in order mean ranks
2
Assign increasing .
hen ties ,
assigned.
3 Show rank
& Zank Zank 3
Sum ranks
23 1. 5 220 & 2 3
e C
2 25 5 2 l
e C 2 25 22
13 14 4
k 25 15 235 12 2
2 R
317T
1+ 75 Mi 23 10 5 22 g 1
i=1
.
-
56 = 48 = 16
2
2 52 6
3 -
je
565 5
2
Ce ve
3
2 ale .
S , reject /0 S
)
C - -
2 :
1
. Horizontal Pattern :
-
.
2 Trend:.......
Gradual shift lower values
to
higher or
.
3 Seasonal Pattern : Beach in Summer
4
. Tre n d and Seasonal Pattern : Combo
.
5
Cyclical :
Colum :
1 t Time Series
2 Tre n d t : 5
day moving average
Add 5 it : 5
day
3 Seasonal Ye : Te
rregulare :
Row 1 : 3
S
1
5 Deseasonal : = S = T Row 1 = 4
6 t : Count f
1 to amount entries
7 T same as 5
8 Fitted T :
b it
9 F = 1xS Row , x
T 2 3 4 5 6 789
1 =3 =
55 69 125
.
TYPES OF FORECASTING:
·
Moving average
· Tre n d projection
·
Multiplicative model
1. MOVING AVERAGE:
Only horizontal pattern
35 -
=
yi
-
yi =
92
2. TREND PROJECTION:
Trend
ny Irregular component
b it
3. MULTIPLICATIVE MODEL:
t + Se It
& Time Series
Colum :
1 t Time Series
2 Tre n d t : 5
day moving average
Add 5 it : 5
day
3 Seasonalt Ye : Tt Sl actual value Centered MA
rregular : :
=
S1 Sum for
Unadjusted = n
#n
S
1
5 Deseasonal : = S = T Row 1 = 4
6 t : Count f
1 to amount entries
7 T same as 5
8 Fitted T :
b it
9 F = 1xS Row , x
Calculator
SI =
actual value : Centered MA Fitted T : = A + Bt
SI form
Unadj =
sum n F =
TX S
Seasonal s)--djusted .
factor
-
djust =
Unadjusted Six Correction
CHAPTER: 18
SIGN TEST:
Compare observation to
hypothesized value
of population mean
hypothesized :
hypothesized :
· :
eiminate the observation
·
lest number
stat =
signs
·
Decision rule Binomial : table ,
n = entries and p =
... or use
get get
Te s t stat or more .
EXAMPLE: -
: median =, 5 -
:
P = 5
-
a
: median =, 5 -
a
: P = 5
n = 10 , p =
0 . 5
Binom Dist .
6
,
0 5 , True
.
=. 8281
P alue = 1 -
-
8281
-
119
2 Sided :
p - 3438 . S
2 = 5
c =
2 .
5
,
2 +
=
m =
mp =
mp1 -
p 2 .
5
Y
23
EXAMPLE:
median and
hypothesized 22
-
: median - 236 -
:
P = 5
-
a
: median 236 -
a
: P = 5
u =
np = 56 =
3
0 = 60 . 5 .
5 = 3 873
.
22 .
53
p-alue = 22 5 .
note correction factor of . S
22 . 5 -
3
=
PZ 3 873 .
P1- 94) -
excels
1
Using S Dist
=
.
orm .
.
=
.
262
, reject
EXAMPLE:
1 :
3 1 :
21 : 9
-
: median- m =
mp = 53 =
15
-
a
: median O = 30 . 5 .
5 = 2 7386
.
p- alue
= 21 - 5 note correction factor of .
S
. 5
2 -
15
=
PZ 2 .
7386
=
P2 .
1) = 222
, reject
fx fy
1v1
1y 1
a
M My
·
a mount of shift #
F fr
#
? #
hen parametric anal gue tests fre la it, f medians in the t distributins .
EXAMPLE:
MANN WHITNEY WILCOXON TEST:
Two fuel additives are being tested to determine their effect on petrol mileage.
Seven cars were tested with additive 1 and nine cars with additive 2. Data shows
miles per gallon obtained with the two additives. Test whether there is a significant
difference between petrol mileage for the two additives.
-
:
cuatins are identical
- : " Justins identical
a are not
Sort small to
big
ranks in order mean ranks
2
Assign increasing .
hen ties ,
assigned.
3 Show rank 2
17 3 2 18 7 .5
8
Sum ranks for Sample
. .
4 1 as test statistic
18 .
G 17 8.
19 1 1 21. 3 15
57797595
.
ninn2 16 7 .
1 21 . 1
18 2 .
5 22 1.
16
1 1 18 G .
7 18 7.
8
.5
o
12771722227 1279797 17 5 . 3 19 8.
11
2
.7 13
94472
2
.2 12
3 c9d
Y
34 correctin factor f S
3
.
9 .
72
260
p-value =
.
2 J
H.
reject
arametric :
- --
IABLE : Te s t of independance :
P value = Pcc observed
Sum - Degres a Can as value)
Surcef square
-
:
F :
pop
test stat
- Ho
--
table
I reatments
-
r p
Sludres-reedom
SS2
SS T
52553
S
T
SS
52
S
to
22
:
independant
=
for
zij
Ha :
Dependant
Df =
#rows-1 # Columns-1
· Treatments factors
:
: I ta SS 2
= [25a)
df =
(2 1)(7 1) 6a
- -
= = 0 05
:
.
To t
Independant :
·
PARB =
P(A) P(B) .
Te s t Stat : = + ·
P-alue =
#- prob
=
row total x Colum total
total
· (n-test stat)'s prob grand
·
=. 5 a3 + or -
on Parametric test :
Binomial sample
u taken from certain distribution .
=
Pxmo =
np( p) -
#
One sample or matched Pairs
· · Two ind sample
·
Mann Whitney (Wilcoxon Rank Sum)
·
Sign test
·
Sampleizes
·
n
>ann2 , 072722nn2
·
Kruskal allis
·
Te s t Stat : Sum Ranks sample 1 "W"
W
:
Sort small to
big
order.
Ho Pop identialnatical
2
Assign ranks in
increasing hen ties ,
mean ranks
assigned . table
TS add correction) -
No
get pvalue
· z to -
3 Show rank Ow
tailed Px2
Always
:
· 2
Sum ranks
chi table determine
square distribution
5 Use to the p-value , using K- 1 of,
2 kR ?
317T
+ Mi E ce :
Chisq dist Re Hidt)
15 i
.
.
= 1
T number of ua S eC
eC
11 ; number f bservati ns in sample i
:
Jeterminant
C
ranspose
Property 1: Property 3:
multiplication of any one row or column by a scalar
Property 6: then -
is a
singular matrix .
then-is a non-singularmatrix.
is zero .
Cramer's rule:
& 113972xz =
Dy -
d213dzz2 D2 -
3th
an a a and -
a D1 913
be a , A
+ 1
el :
legeometrics method
L G aaehu -
S
Assembling time 25 5 17
Finishing fftime 35 3 15
umber funits x Y
c
rofit functi n: 1 x + 15 y .
#
24 feasible
B
Zestricti ms :
28
16
25x + 5y11 C 25x + 50x = 1100
12
M
35x + 3y15 g
&
x, D
Y
% 5 10 15 20 25 3035
1108
25x + 5y
= 44 22 and
= 11a+ x =
25 ..., ,
05
35x +
, =
3
: 35 and 3
3y15a+ x =
35 ,
25x5y = 11
y =
% Y
35 5
12 25
To the profit function the
.
max more
,
·
Profit 1 : x + 15 y = 1 19 5 .
+ 1512 25 .
= 378 75 .
· -
·
Finishing off time : 35x + 3 y =
35(19 5) .
+ 3 (12 .
25 = 15
·
The corners of the feasible region are the extreme points .
·
The solution of the linear equations can always be found in at least one of the extreme points .
·
If two points Cand D give both an
optimal solution then all the points on CD will give the optimal solutions .
·
ethod
16 D (3 , 10(30) + 75(0) =
300
195 12 25
C ,
.
25x 1100
+ 50x =
12
M
&
D
% 5 10 15 20 25 30 35
X 7 6 X, 7
D
2x
27 :
↑X .Setx
6 an fin 16. Sety an fin x 6 .
Crinates 6 and 6
Crinates 5 and
setx y CS latc X
.
7
T e
6
C eama
g
3
S
&
2
2 3 5 678 &