0% found this document useful (0 votes)
5 views35 pages

Simple Linear Regression Overview

The document discusses logistic regression, a supervised machine learning algorithm used for binary classification tasks. It explains the logistic regression equation and the sigmoid function, which maps predictions to probabilities between 0 and 1. Additionally, it highlights the application of logistic regression in predicting outcomes such as the probability of heart attacks or university enrollment.

Uploaded by

Bhakti Sharma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views35 pages

Simple Linear Regression Overview

The document discusses logistic regression, a supervised machine learning algorithm used for binary classification tasks. It explains the logistic regression equation and the sigmoid function, which maps predictions to probabilities between 0 and 1. Additionally, it highlights the application of logistic regression in predicting outcomes such as the probability of heart attacks or university enrollment.

Uploaded by

Bhakti Sharma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

sd tat wethod stailial

uhist. dipeucbut
Regpuuion tinea Simple
vekibl dgundeut
cat
dahust
deuuivd
ML
Regauio< ditax Whats
Cubpnt (ne
Lineast Requsion ine
calld

tw

wheu

wd po-pz

ut wt xady valw ad
l u kke ed.

D D

P
X
neinu
hald be oa
ud ata valus

lyauclud

batsons, qau
vecos couai iy alhsu ih kdscooms ,

Eoaupl - Fid a
ata et.

8
hu

yi-7-) 2

-3.25 I|. 25
2. 3 -3

- 0.75 -0.+5

1. 25 J.25
L

Jo Jl. 25
13. 25

2+4+ 648 5

26 bo25

25-0.45 23.
15
B ll.25 + 1-25 +|·

b.25 -5. 15 = 0.5


16) L5)=
Bo = 6.25 - l·

4 0.5+ (-15)
MEAN SQUPRE ERgOR (MSE)
valvatan uie thet colutat
Mion Solarvd Etat CMSE) is or

all the data hant Whe


olilsd valu fu faile iunu don t
squasad t smawe bhat nugaluu and
lanl cach athee aut.

value as the i dat paint


de autal es obhud data paiat

EAN ABSOLUTE ERRoR (MAE)


coleulat
alatog meti sd t
ngrssian modl. MAE uoswes
bhe acay a poliid valus ata
pdetid and
absalut diee eaun tu
awsage
actucl valus.

gi punti tu alal valus


tu pruolictid valus.

Lay MAE value ioicots bettt maol tfamanc.


ot u i t s t the
sutlins
Dapndent
vaiabll

CastEaon'
Jlo, e,) - L

9ndypndont vaniabl

Bho, calulat
dyhndnt
[quaked rot CMSt).

Anb.
7

To fond linat

2
|6

36 30
6

lo 64 8o

144
dat To Zicuning
p dalastt
panamelirs
lo odela
an squanio mian tetuduce ducert gedint algoiton
meoygiy
Me chaluy y do
tu sing tiaind can iodlel
be
DESCENT
FoR GIRADIENT
LiNEAR
GRE (E
SS
2l75 =&.7
7:2 5 6
84 4.
2.89
5.3
O-16 3 2
Cy
75z +o =5
=/·5
0-?5 =
lminiizinng RMsE valu) onod achime tu but
tio moolel waLs
uuth kandam &, s oe valaus amd tn itromly
mmiinumm cot
up dat te valus , oching
but a daiabiu tat lnes he
gradunt s noltig
vatiatons in nft.
CosT tuNCTION
J9 9, ) = i=)

Vall

Larng

Chyl)-g)2cox-)

LINEAR MEGRES51ON
SiMFLE
NssuMPTIONS OF variakles
indebndnt and debendt
The

thet hags a luat Jauhion


then lmlat tugusto
will mot be
eccurale o d l
iabl defbensert
chang
he
Bagsien Linar Advantagzs
fpinoisls. fast
valu
Laings fhie slock
variabl.
Fore pliala bhauiet
a the prdit ond
sinfaet
vaiablcs
has inlaßerdt tu omaunt the thed
corelant
hie bbos
is aionee
th te 5), Variabl
moll. accwnak
a be noB weill
Vasiable
dos
olvalons fridlkinolnee
he :
inat Agnion cerputalanaly ind amd cam handd
Lange datauh ctuly . et con ke thand quicky
abplatons
Limat elatiuay sobst aulliss cofarud
hau
C srralks afoct on the onsall model peufasaane

comhoe machine larming


agauthn.

Disadantags Linan Mayusion


abses a linat elalonshil butusen

Aelatiansh i zot lro the meodel moy

high coelaton ktsn idl.


vaiakls ullicallirloily can inflat to vaianee y
he caliint s ad wetobl meoel folictons.
Linat nigusion aummes thad ta fats
Jatun ae alray

fomat that can ke


Jonclon lost Jecode Psudo
tauaton ugnusion Liaas The
z)/n Sua 2t -CSumy
3.
+
a) 2
the in9:) (i, pant data cach Fot
dalast :
dataat
0 bllm
Snitializa
Repasin Linas Psudecode
Jo
laring -rahirs asaned
vaiall.
Morebe lalaahs
lsar conply
6) aluott wmot

3. llatt t on
in te dataset )

lsudecode Jos Cradint Duent


manhalons &- loo0

whil t < mee itbnabins do

endl

Gubuited By i
(PN22/HseHpse
Hiuani (9keel MS MAT| 5?)
Subict Maciutny (Mc
DELHI TECHNOL06i CAL UNIVERSITY

MACHINE LEARNING
MSMA 218

Simple bogistie Begresion


Notes

Subri tted To : Submite d8y:


Rof Arjana qupta Suoati Suman
2KI2[MsCMAT/5|
Ritika Gupta
2K22/MScMAT |54
1 Loqistic Resvesion
a sugesvised machine leari'
Logistic neresion is binay cagtfation tak
algorithm thet ace omplshes'
probebility of an outtome,evnt,
by predicting the model deliers a
or oaseryatisn. The two pomm ble sutcoes:
to
dichotomous euteome limibed
o/1, or true/false
qesfno, o/1
the relationglip betten
Logial ngouin amalyes
more ndegendeut vapables and cashes
into discyete classes. 9 is extevely sed
data estimates
predicthie modelling, shere the model
in am instamce
the mathe matial probabilty of ohe ther
based
os not on a
to a spleike catgoy
ven data set ay independent vanables

for erample , o- represeuts a negatiie cas; ! -


vprents a positie cdass. Logstie resin
used in binary cluifestin
Problems where the otcome yanable reveals either
of the two categors (0 and ).
Exomphs?
1 Detmine the probability of heat attacks
PRssiblity of enyolling into wuversity
3 9demt tying pam emaila
2 Simple Logistie vexe sion Eguation
Simple agitit Regresin statistial test
wsed to predict å smale bnar yarnable wsing
a

one other vanable. This is done byy using the


sigmeid function to map predie tioni amd their
Probabilities.

The sigmsid functionandconverts ar 4 real


e value to
a range betuween 0 1. Such that, if the
Putput of the sigmoid funetion (estinated pro bablt)
is geater tham a predefned threshold on the
aph, the model preditts that the instance
belongs to that clas. 4 the estima ted probabib
is les than the predifined threshold, the madel
predicts that the initamce doe not belong to the
class

Logistie Regresion eqatim; |+ elbetbx)

where,
input value
8 predicted sutput
bo = biad or intercept term
bË = (oeffiient for input (x)

Common questions

Powered by AI

Logistic regression differs from linear regression in that it is used for binary classification problems where the outcome is categorical. While linear regression predicts a continuous outcome, logistic regression predicts the probability of a certain class or event occurring. It uses the logistic function (sigmoid function) to map predicted values to probabilities between 0 and 1, allowing for the classification of data into two distinct categories .

Simple linear regression primarily uses the dependent variable (response) and the independent variable (predictor). The goal is to model the relationship between these two by fitting a linear equation of the form Y = b0 + b1*X + e, where Y is the dependent variable, X is the independent variable, b0 is the y-intercept of the line, b1 is the slope, and e represents the error term. The relationship reflects the change in the dependent variable for a unit change in the independent variable .

The sigmoid function plays a crucial role in logistic regression by transforming linear predictor values (real numbers) into probabilities between 0 and 1, which align with the concept of binary classification. This allows predictions to be interpreted as probabilities of belonging to a particular class. The non-linear nature of the sigmoid function makes it critical for converting continuous input into a discrete output suitable for binary decisions .

Logistic regression is preferred over simple linear regression when the dependent variable is categorical and dichotomous, such as scenarios requiring binary classification. It is more appropriate in situations where the predictive outcome involves the probability of occurrence of an event, such as classification into true/false, success/failure, or positive/negative classes, where simple linear regression is unsuitable due to its continuous outcome prediction .

Gradient descent functions by iteratively adjusting the model parameters to minimize a cost function, typically the mean squared error in linear regression. It begins with a set of initial parameters and updates them by computing the gradient of the cost function with respect to these parameters. By moving in the direction opposite to the gradient (descending), it reduces the cost iteratively until convergence is reached or improvements become negligible, thereby finding the optimal model parameters .

Small sample sizes can undermine the reliability of both linear and logistic regression models by increasing the variability of the parameter estimates and reducing the power to detect significant relationships. In linear regression, it limits the ability to generalize findings to larger populations. In logistic regression, small samples may lead to overfitting, where the model performs well on training data but poorly on new, unseen data. Thus, ensuring adequate sample size is essential for model stability and generalization .

Multicollinearity refers to the correlation between independent variables in a regression model, which can inflate the variance of coefficient estimates and make the model unstable. It leads to difficulty in determining the effect of each predictor variable on the dependent variable. Multicollinearity can be detected using Variance Inflation Factor (VIF), where a VIF value greater than 10 indicates significant multicollinearity. It is also observable when adding or removing a variable causes large changes in model coefficients .

Mean Squared Error (MSE) is a measure of the average squared difference between predicted and actual values in regression analysis. It is calculated by averaging the squares of the differences between the observed and predicted values. It is significant because it provides a quantitative measure of the predictive accuracy of a model, with a lower MSE indicating a better fit to the data. It highlights the magnitude of prediction errors in the model .

Simple linear regression offers practical advantages such as interpretability, simplicity, and fast computation. It provides clear insights into the linear relationship between variables, facilitating easier communication of results. In scenarios where data shows a linear trend without complex patterns or interactions, it is advantageous due to its computational efficiency and minimal data preprocessing requirements .

The assumptions underlying simple linear regression include linearity, independence, homoscedasticity, and normality of errors. Linearity assumes a direct proportional relationship between the independent and dependent variables. Independence requires that the residuals be independent. Homoscedasticity refers to constant variance of the errors across all levels of the independent variable. Normality of errors implies that the residuals should be normally distributed. Violations can occur due to multicollinearity, presence of outliers, or using inappropriate models for the data structure .

You might also like