."
Ifthe space is two dimensional then the half space is kaown as half plane, The half space that is in
cone dimensional space is known as ray.
15
‘Scanned with CamScanner2 equation whi
‘derived from
an open hal esapcs.
ity that is strict will specify
Fax ob
ar inequality that is not strict és called closed half-space.
‘Scanned with CamScanner
38, Fay +e bay 2b
Consider the bclow piven two dimensional space.
H
+X,
ve half of plane
ae
i ‘An equation in two dimensions can be a line that can must be hyper plane. So equation ioncy,
can be written as, .
i Xn+b=0VX cline
In these two dimensions the line can be
xn, +10, +b=0
This line be extended on both sides even. If this is done the two dimensional space is divided iy,
| two spaces :Data Science unira
‘One space is at ane side of the line i.e, at right side and another space is at the other side of the Tne
ic., at let side, These two spaces are called half spaces, For example if there are points an one half space
‘and points on the other space, Is thereany characteristic that ean separaic them’ A solution for this would he
toperform certain computations on one half space fur al the points and obtain some result, Repeat the same
procedure on the other side and use the results to make the decisions, These type of situations are mostly
observed in classification problems. Consider a binary classification problem, to know on which half space
the point lies in. And now consider three points X,..X, and X, from the above figure and! distinguish their
positions. In the equation Xn +
0, nis said to be normal in thi
‘The above figure, ifn is considéred as normal in equation X7n +b =0 and ifthis equation is multiplied
by -I then normal is said to be defined to side of n, Otherwise normal is said to be defined in the opposite
direction of n, To know where the points X,, X, and X, lie, the equation Xn +b = 0 must be evaluated.
xTn+b
XIntb
Thtbo*
Forthe equation X7n +b 0, itis clearshatihe point fies on theline so it evaluates to0. Now consider
tie equation Xn +): In the above figure; take two points X' and X,, Here X, is a vector from to X,. Fron
vector addition, itcan be written as, :
xe
‘This must be substituted in the equation, :
weyyed
XT+b+¥'Tn
‘Scanned with CamScannerIfthe point ties in oF IV" quadeant, then the angel would be a positive @ angle, Foy in
| thoy
i a dot matrix ath = ‘aifbleos 8 the 0 angles might be between the two vectors. For any ping, ig *y
, ,
{ 270° to 360%, the equations Yn evaluats to a positive value since a°b is also positive .
| Xn b+yn > 0
‘ the points are at the opposite side ie., between 90° to 180° or 180° to 270°,
The cos 6 for angles between 90 to 270 would be anegative value. Therefore for any poj,, te
(on this side of line or half space, the computation Xfm + b would be less than 0.
Xtneheo
|
Example
Consider a2D geometry with n=
“
ix [Joa
X?+b=0
"end b=4
jj] andb=4.
x
fx tixygtd=0
ler three points, (—1,-1), (1,~1) and a 2), Substitute these points in the above equation
| Cb
x 4aepe aso
a+d=0
‘The point (-1,—I).is said to be on line.
2 GD
nt3n ded
1=3+4=2>0
|, ~1) is said to be in positive half space, —
13
3. Tie point
all
‘Scanned with Cam$cannerData Science
UNIT-1
2 Gd)
X,43x,440
1-6442-1<9
The point (1, -2) is said to be in negative half space
1.6 EIGEN VALUES, EIGEN VECTORS
Q15. Write about eigen values,
Answer : Model Paper-Ill, Q14{b)
Rigen Values +
Eigen values are the numerical values that can determine the number of features fo be retained. The
concept of eigen values is applicable to square matrices. It is considered as a very important topic. For
‘example let the principal component analysis is based on il. Each eigen val
definitely has corresponding «
‘eigen vector. The principal component analysis of a sy$tein of variables is performed hy computing the
eigen value of dispersion matrix or correlation matrix of variables, This principal component is considered
to be the linear combination of items of correspon:
i eigen vector.
‘The cigen values define the proportion of variance provided for every eigen vector that is derived
from transformations of original set of variables to orthogonal variables, This leads to a decrease in number
of variables that are used in determining the majority of total variance among, otiginal variables. If each
original variable contributes in the direetion of eigen vectors then the important variables can be summarized
in less number of vectors.
Consider the below mathematical formula,
Ax=hx -
Here, constant 4 (positive) represents the amount of stretch or shrinkage that the attributes x go through
the x direction.
‘Scanned with Cam$canner8X are called eigenvectors and their corresponding? ae called wigan ym
atrix, the eigen values and eigen vectors can be computed as fallowys,
The eigen vatties can be computed as follows,
AK = Ax. Atn ays xt 1)
AX —AIk = 0
(A-2)x = 0
Therefore the eigen values of the equation can be determined by using the below canis
|A-Al| =0
2. By substituting the eigen values in original equation the solution for cigen vector x an be com
Example
Consider the below matrix
_[8 7’
“bl
Bis] _ al |_ day
23) ,e}> “(x |7]an
a7 ro
[spol
Bae 7
2 3-2
=0
|A-All =
(@-NG-H-14=0
A-1IA+10=0
40,1)
-—_— TSS
20
‘Scanned with Cam$cannerData Science UNIT-1
R code is ns follows,
> RE-MALEAC(C (87,2, 31-2424 yEOWAT)
‘Therefor2, there are two eigen values,
To, comptite the eigen vectors considers the below process.
ee Pee
ESI) E
|
Bxy+7xy
Therefore the corresponding eigen vector to = Vis,
dnt Bi
xX 4X,=
Pe
IfaA=10
8 7]}x |_| tox e
23][m| [lon
8x, + Tp ]_[10x,
2x, 43x,] [10r,
21
‘Scanned with CamScanner‘Scanned with CamScanner
& RECrWEE AM (5EE47,2/21,2,2. YEON)
[> avcesgen ta)
Retationsh
between Eigen Values and Eigen Vectors
‘Theeigen values ean be complex numbers even fo real matrices, the eigen values become compley
than eigen vectors also become complex.
TF the matric is symmetric and if this symmetric is in the following,
AnaAT
then there are following properties
@ Ifthe matrix is symmet
thon cigen values will be real always
Gi) Eigen vectors of the symmetric are also real
For a matrix and for
Vp VovenV, for symmetric matrices
Q16. What is an
Answer : Model Papersil, 118)
gen values Ry 2, wy h, then there linenrly independent eigen veetors such ss
jen vector? Explain,
Eigen Vector
Jn linear algebra, if T is a Tinear transformation from a veetor space V over a field F into itselFand
is nonzero vector in V then v is called an eigen vector of T if T(v) isa scalar multiple of v ie, To)=
Where v isa scalar in F known as eigen value associated with eigen vector v. Eigen vectors are associated
22
alData Science UNIT-1
‘Tiih linenr models which are wnasial ih enginesring as first approximation, These veetors are wnrotated
uy is mostly vusefl in
aluerithi,
by transformation mavix, It is applicd on machine learning algorithm. This aly
handling the argc dla sets. The concep! of eigen veetery ix considered as a back bone of thi
For example, multiply a 2-dimensignal vector with a 2°2 mattis,
12)1.3
03)2 6
‘This particular operation on vector is called linear transformation, ‘The cofunn matrix represents a
jgetarwhase input one! eutpul vector directionsare not sume, The weetars whose dimension does notch
after applying finear transformation widh matrix are called eigen Vestors, This cangept is applicable only:
square matvices.
Finding Eigen Vector of a Matrix
Consider a matrix M and cigen veelor ‘e’ corresponding to the matrix,
“The direction of **remaiis unchanged when multiplied with anatrix, only has @ change in magnitude,
Consider the below equation,
Me
(M—C)e=0
‘Interms of (MC). C indicates an identify matrix of order equal to “MF that is multiplied by a sealer
*c'. There are two unknown “e" and ‘x’ and one equation, This equation can be solved by making the veetor
*e°as zero vector. Then there will be only a single choice that, (M-C) is a singular matrix. [t has a property
‘that ifs determinant is equal to 0, This property can be used to find the value of ‘c".
©
Det(M-C)=0
‘This produces an equation in ‘c* that is in the order based on matrix M, ‘This needs a solution
for equation. If the solutions are ‘cl", “e2" and so on then place *c1” in the eqnation and find vector “61”
corresponding to ‘c1*. The vector ‘el isan eigen vectar of M, This procedure anist be repeated with *¢2°,
“c3" and soon, .
Example
Ecit_ Mis
> He~maerix{e(80,31,20,$1,50,51, 60, 61,70) ,nrewss, byrow=T)
> xeceagen (Mf)
> xSveiues
£2) 147,737876 §.317459 -2.055095 i
> xSvectors
La 21 Lay
{2,} -0.3968974 0.9897557 -0.7447e185
(2)] -0.8497487 -0.8198420 -0.06303763
a) =0.7961272 0.366296 0.6643239. :
>
‘Scanned with CamScannera Scien aR
Q17. IMustrate thi usage of eigen vectors in data scient
Answer :
Theconceptofcigen vectors is applied i e., machine learning algorithm principal Component an,
re is data with huge set features It has high dimensionality. There mightbe redundant feature ina ab,
‘These. features make the eff ieney to reduce and disk space to inerease. But the PCA craps joo. ie
‘The cigen vectors help in defining these features.
ha
Consider the PCA alg
fo perform this are as follows.
‘ith for ‘n’ dimensional data that are to be reduced to *k* dimessiong 5
"ts
Step 1
Initially the data is mean normalized and feature scaled.
Step2
‘The covariance matrix of the data set is computed.
To reduce the number of features (dimensions) de, the features must be deducted. But this ead
Joss of information. So, loss of information need to be mi
ized and maintain the maximum varanes, fy
this, the directions of maximum variance must be determined. This is done in the next Step.
Step 3
Un this step eigen vectors of convariance matrix is determined. Since there is data in ‘n' dimes,
then ‘n’ cigen vectors corresponding to ‘n’ eigen values are deterinined, —~
Step 4
Select *k’ eigen vectors corresponding to ‘k’ largest eigen values and then build matrix in whieh evey
eigen vector that represents columns. This matrix is called as
In order to reduce a data point ‘a’ in the data set to ‘k’ dimensions, the transpose of the mairixU must
be determined and then multiplied with vector ‘a’. Then the desired vector in *k” dimensions is obtained.
24
a
‘Scanned with Cam$cannerSTATISTICAL MODELING
PART-A
SHORT QUESTIONS WITH SOLUTIONS
1. Define statistical modeling.
Answer: . Model Papers, 3
Statistical Modeling
Siatistica! modeling can be defined as the formalization of relationship in between the variables is
the form equations, ts actually about finding out the variable, It explains about how variables ae related
‘with each other, The relationship can be in the form oF mathematical equations. And the variable can be an
attribute such as height, weight or age ofa person.
| . The variables
| analyzing and applying it on varios circumstances,
Statisticel rodeling gives the introduction and illuminates the statistical reasoning that is uscd is
modem research throughout the natural s well as medicine, ecommerce, social sciences. government et. It
also focuses on the usage of inodels to untangle and quantify the variation on observed data.
Q2. What is a random variable?
Answer = Moulel Paperstl, 03
Random Variables
+2" Random variable is variable that takes particular value .e., numerical valuc with definite probability
It is obtained from the resull of rangom experiment. The random variables are denoted by capital letters and
the corresponding letters are denoted by srual letters. ‘
Example =
If'g fair dice is rolted and if'X* denotes the number obtained then *X” is called as random variable.
“Thus *X' can take any one of the particular values such as |, 2, 3, 4,$ of Geach with n probability 1/6. These
| Values are tabulated as follows. * ‘
‘Scanned with CamScanner
1 not be related accurately but ean be stochastically related, It consists of data .nce Using R
All the possible outcomes of random experiment together is called “Sample Spacg ay
* The sum of all probabilities of sample space is # always. "ay
oa
Random variates are of two types, they arc,
() Discrete random variable.
Gi) Contimious random variable.
3. Write in short about hypothesis testing.
Answer : Mode bop
Hl,
Hypothesis Testing 7
‘The statistical hypothesis can be defined as an assumption with Fespect to a populati
Ey Mot be true. It is a set of formal procedures that is used by stalisticians for accepting g-
Atwtistical hypothesis. Infact itis process of validing the hypothesis that is made by rescarchey,
the hypothesis, the complex population is considered.
10" Chan ni
"Sean
8. Fay
4 this process it makes use ef random samples from the poptlation, The selectiy
hy MOF recta
pothesis depends om the result of testing over the sample data. et
4. State the types of errors occur in hypothesis testing.
Answer : Model Papers a,
Types of Errors
‘There are two types of errors that exit occur in hypothesis testing.
1 Type! Error
Taccurs when the null hypothesis is rejected while its value it true. The probability of this enocen
be defermined through the term sighificance'level when the hypothesis is tested. The significance levis
denoted by the symbol ot (alpha).
2 Type tt ferer
Type Il error ean be defined as the acceptance of false null hypothesis H.'The term called poweraftsx
defines the probability of type LI error when the hypothesis testing is performed. It is represented by synto\
B (beta),
ee
QS. Define p-value.
Answer + ‘Model Paper-i, 4
p-value
‘The p-value can be dofinedas the probability of obtaining result that is equal [Link] more than observaon
from data when null hypothesis is true,
Hypothesis testing makes use of p-alus to actully use pvalo fo weight the strength of vie
data of population, The p-value can be computed for the given data through a statistical tes. Tt ele
compared with predetermined value i.e. alpha. usually the value of alpha will be 0.05. If itis less @! oak
then null hypothesis is rejected and if it is more or equal than alpha then rejection of null hypothes!
26
esl
‘Scanned with Cam$cannerStatistical Modeling UNIT-2
ae
fe PART-B
gore ESSAY QUESTIONS WITH SOLUTIONS
2.1 STATISTICAL MODELING
Q6. Discuss about statistical modeling.
Answer: Model Papers, @12(a)
Statistical Modeling
Statistical modeling can be defined as the formalization of relationship in between the variables in
the fonm equations. It is actually about finding out the variable. It explains about how variables are related
with each other. The relationship can be in the form of mathematical equations. And the variable can be an
ute such as height, weight or age of a person.
The variables might not be related accurately but can be stochastically related, Statistical modelin,
‘consists of data analyzing and applying. it on various circumstances.
Example
The attributes such as height and age are probabilistically distribited amang humans. They are
stochastically related i.e., if a person is of age 35 then this influences the chance of this person being 4 feet
{all and if a person is- of age-15 then this influences the chiaice of this persan being 6 fect tall.
Model 1
Height, = 6,+8,ape,+¢, “
Where, a
8
intercept,
6, is parameter that age is multiplied to generate a predi
€ is the error term and davutd
‘is subject,
Model 2
by bage, +b, sex, +e,
itistical reasoning that is used
sciences, governmentetc. It
"Statistigal’ modeling gives the introduction and il sli
modem research throughout the natural as Well as medicine} ‘eBrnimérce; 5
also focuses on the usage of models [Link] and quantify [Link] of [Link]: «
°? “Tesiplaté for statistical model mould be a linear regression model with independent and homoscedastic
errors.
ysi=sum_{j =.0}"p beta jx_ti}+ei, c ‘
oo . i
Fed, i
‘Scanned with Cam$cannerData Si
Where,
ce jare NEO (0. sigmna”2)
Inv vate terms tis eam be ete 8
y =X belate
Win
0, Hone
rdesign mati widneolumes
ig response weetor, Xis model mal
cariables.
More frequently X_0 mould be a column of ones by defining a
wn intercept term.
c
J! from potentially large sc of
Fieal models are ifusirated as jy?
“Dypes of Statistical Modeling
sing the mvnimal adele mod
uns cho' :
‘Various types of statist
simplification.
Statisiieal inodelling
models by using stepwise mo
Model Tnterpretation.
‘Saturated model ‘ne parameter for each dati point
Fit: Perfect
Degrees of freedom : None
Explanatory power of model : NONE
Treonsists all (P) factors, i Factions and covariates of any interest, Moy
of the madels ean be insignificant
Maximal model
Degtee of freedom : t-P—
Explanatory power of model: Depends.
simplified model with 1 gamma (20, 10)
J) (2) 8.322375 is.661ses 10.s27896 18.807450 LD.sa2sE2 B.1E1262 126780455.
{8} 10.709388 11.s49666 11.256586 16,979900 10,419608 15,895826 10.052508
Hal 8.436457 10,269957 6.191293 9.510985 8.270894 14.367074
>
5 recom
Ttretums n random numbers from geomettie distribution.
‘resom(n, prob)
Heren indicates n indicates number of observations and prob indicates probability a success in each
ackages 5 Windover»
> set .seed(2)
>-egeames, 1/6)
6 rlnorm
Il generates random amounts with a multivariate lognormal distribution or density of this particular
distribution at some specific point.
~rinorm(n, meanlog, varlog) :
Here n indicates number of data sets that are to be simulated, meatilog indicates the mean-vector of |
logs and varlog indicates the variance/covariance matrix of the logs.
S$.
é 31
‘Scanned with Cam$cannerData Science Using R
Example
Ble bait Misc Packages: Windows Helo
> doe (ztasen(s))
1) 0.210731885 o.oesa9s6q7
16) -1.246783429 9,99815995 0.580872"
\ 122) =1.4508639¢5 9.3s0909791 ~9.47452602
26) -1.087292503 2,03G203603 -0.926989232 sae
2-763246020 o_zeez02760 ~2.252558924 -1. 29956975
riS08a8sE13 o_s275¢0097 -o.sassae57s -0.9FE37EALS -0.7205¢5
30291196 o.eT7B¢a42 0.452793: 76 earl
2-85600373¢ a ogess922 g.a7G¢03855 9278215449 -2/87790294
~O.B26s26142 Lloia7rog6a 891277732 0-742002772 9147573408
e O-425365565 [Link] o.agi9ae754 9.225422912 -1. 010465085
-a,462689253 0.81083980 ~1.912248796
oo -0.216375791 -2.621957255
35402726:
7. logis
This function depicts information about logic distribution. Kt generates random devia
Hlogis(a, location0, seale=1)
ese minicates numberof observation, eeatonandscalehave0 and 13 Fm values nog
Example ;
}|> vax (=tagas (1000, 0, seals = 5))
8. rmvbin
ereates corretated multivariate binary random variables by thresholding a normal distributing
rvbin(n, bincor, margprob)
Here n indicates number of realization of variables that are to be simulated bincorr isa mains!
margprob indiestes the vector of some length,
Example
rmvbin( 10, margprob = C(03, 0.9))
«pois
Il generates values from poisson distribution and returns the results,
rpols(ob, rate = rate)
Here, ob indicates the number of observations and rate indicates estimated rate ef events for dss
32
‘Scanned with Cam$canner\> Statistical Modeling UMiie2
Example
It generates random compositions with uniform distribution.
if(9, win, max)
Here n indicates number of observations, min and max are by default 0 and | respectively.
Example as
[4] ~0.81133706 ~0.03129085 ~0.s7e Hine Bn, 8
‘There are even other types of hypothesis testing, they are as fullows,
Simple Hypothesis
Simple hypothesis ig a statistical hypothesis which completely specifies an exact Paraneter py
hypothesis is always a simple hypothesis stated as an equality specifying an exact value OF paramet, “
Example
bo Hw=y,
2 Hor y,—p
Complete Hypothesis
Composite hypothesis is stated in Trims of several possible values Lc, by an inequality. Aten
hypothesis is a composite hypothesis invalving statement expressed as inequalities such a8 <> ore,
Example
1 Aopen,
a H,
BSB
3. A pep,
Example of Hypothesis Testing
‘Consider an example, to check whether a coin was fair and balanced, According to mull hypothes:
the half flips would be of head and half would be of tails. And according to alternative hypothesis the iss
of head and tail may be different.
Hy: Ps05
H,:P#05
for 50 times, might result 40 heads and LO tails, Based on the result the null hypat®*
must be rejected and concluded according to the evidence that coin was not fair and balanced probebl-
—_—_—_——
Flipping of ©