Data Analytics and Preprocessing Guide
Data Analytics and Preprocessing Guide
MODULE 2
Chapter 1: Descriptive Analytics I
3.2 The Nature of Data in Analytics
Data is the raw material for BI/Analytics/AI: without reliable data, models and
dashboards mislead.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Specialized:
Metadata: data about data (definitions, units, owner); critical for reuse.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Choose latency to match business need (e.g., fraud detection needs streaming).
By structure
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
By business role
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Goal: turn messy, heterogeneous raw data into high-signal, analysis-ready datasets.
Typical pipeline
3. Cleaning
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
4. Integration
6. Feature engineering
8. Imbalance handling
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Big data = data whose size/complexity/speed exceed the capabilities of traditional tools to
Why traditional RDBMS struggles: rigid schema, vertical scaling limits, single-
node I/O bottlenecks.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Core principles
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Analytical workflows
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
6. Ops: CI/CD for data & ML, drift/decay monitoring, A/B testing, feedback loops.
Pitfalls: schema drift, silent data failures, spurious correlations, training-serving skew,
privacy violations.
5. Dynamic Pricing & Revenue Mgmt (demand signals, competitors) → margin lift.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
10. ESG & Risk Reporting (energy, emissions, supplier data) → compliance &
reputation.
There are a number of technologies for processing and analyzing Big Data, but most have
some common characteristics (Kelly, 2012).
Namely, they take advantage of commodity hardware to enable scale-out and parallel-
processing techniques; employ nonrelational data storage capabilities to process
unstructured and semi structured data; and apply advanced analytics and data
visualization technology to Big Data to convey insights to end users.
The three Big Data technologies that stand out that most believe will transform the
business analytics and data management markets are Hadoop, MapReduce, and NoSQL.
Hadoop
It was designed to handle petabytes and exabytes of data distributed over multiple nodes
in parallel.
File systems such as HDFS are adept at storing large volumes of unstructured and semi
structured data as they do not require data to be organized into relational rows and
columns.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Each “part” is replicated multiple times and loaded into the file system so that if a node
fails, another node has a copy of the data contained on the failed node.
A Name Node acts as facilitator, communicating back to the client information such as
which nodes are available, where in the cluster certain data resides, and which nodes have
failed
The client submits a “Map” job—usually a query written in Java—to one of the nodes in
the cluster known as the Job Tracker.
The Job Tracker refers to the Name Node to determine which data it needs to access to
complete the job and where in the cluster that data is located.
Once determined, the Job Tracker submits the query to the relevant nodes.
Rather than bringing all the data back into a central location for processing, the
processing occurs at each node simultaneously, or in parallel.
When each node has finished processing its given job, it stores the results.
The client initiates a “Reduce” job through the Job Tracker in which results of the map
phase stored locally on individual nodes are aggregated to determine the “answer” to the
original query, and then are loaded onto another node in the cluster.
MapReduce
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
• Hadoop Distributed File System (HDFS): The default storage layer in any given
Hadoop cluster.
• Name Node: The node in a Hadoop cluster that provides the client information on
where in the cluster particular data is stored and if any nodes fail.
• Secondary Node: A backup to the Name Node, it periodically replicates and stores data
from the Name Node should it fail.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
• Job Tracker: The node in a Hadoop cluster that initiates and coordinates MapReduce
jobs or the processing of the data.
• Worker Nodes: The grunts of any Hadoop cluster, worker nodes store data and
takedirection to process it from the Job Tracker.
HBase: HBase is a nonrelational database that allows for low-latency, quick lookups in
Hadoop. It adds transactional capabilities to Hadoop, allowing users to conduct updates,
inserts, and deletes. eBay and Facebook use HBase heavily.
Flume: Flume is a framework for populating Hadoop with data. Agents are populated
throughout one’s IT infrastructure—inside Web servers, application servers, and mobile
devices, for example—to collect data and integrate it into Hadoop.
Oozie: Oozie is a workflow processing system that lets users define a series of jobs
written in multiple languages—such as MapReduce, Pig, and Hive—and then
intelligently link them to one another. Oozie allows users to specify, for example, that a
particular query is only to be initiated after specified previous jobs on which it relies for
data are completed.
Avro: Avro is a data serialization system that allows for encoding the schema of Hadoop
files. It is adept at parsing data and performing removed procedure calls.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Mahout: Mahout is a data mining library. It takes the most popular data mining
algorithms for performing clustering, regression testing, and statistical modeling and
implements them using the MapReduce model.
Sqoop: Sqoop is a connectivity tool for moving data from non-Hadoop data stores—such
as relational databases and data warehouses—into Hadoop. It allows users to specify the
target location inside of Hadoop and instructs Sqoop to move data from Oracle, Teradata,
or other relational databases to the target.
Spark vs Hadoop
NoSQL
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Stream Analytics
Applications
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Descriptive Statistics
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Measures of Dispersion
Quartiles are determined by first sorting the data and then splitting the sorted data into
four disjoint smaller data sets.
Quartiles are a useful measure of dispersion because they are much less affected by
outliers or a skewness in the data set than the equivalent measures in the whole data set.
Quartiles are often reported along with the median as the best choice of measure of
dispersion and central tendency, respectively, when dealing with skewed and/or data with
outliers.
difference between the third quartile (Q3) and the first quartile (Q1), telling us about the
The quartile-driven descriptive measures (both centrality and dispersion) are best
explained with a popular plot called a box plot (or box-and-whiskers plot).
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
The box-and-whiskers plot (or simply a box plot) is a graphical illustration of several
descriptive statistics about a given data set.
They can be either horizontal or vertical, but vertical is the most common representation,
especially in modern-day analytics software products.
Box plot is often used to illustrate both centrality and dispersion of a given data set (i.e.,
the distribution of the sample data) in an easy-to-understand graphical notation.
The box plot shows the centrality (median, and sometimes also mean) as well as the
dispersion (the density of the data within the middle half—drawn as a box between the
first and third quartile), the minimum and maximum ranges (shown as extended lines
from the box, looking like whiskers, that are calculated as 1.5 times the upper or lower
end of the quartile box) along with the outliers that are larger than the limits of the
whiskers.
A box plot also shows whether the data is symmetrically distributed with respect to the
mean, or it sways one way or another.
The relative position of the median versus mean and the lengths of the whiskers on both
side of the box give a good indication of the potential skewness in the data.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Shape of a Distribution
Distribution is the frequency of data points counted and plotted over a small number of
class labels or numerical ranges (i.e., bins).
In a graphical illustration of distribution, the y-axis shows the frequency (count or %),
and the x-axis shows the individual classes or bins in a rank-ordered fashion.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
As the dispersion of a data set increases, so does the standard deviation, and the shape of
the distribution looks wider.
There are two commonly used measures to calculate the shape characteristics of a
distribution: skewness and kurtosis. A histogram (i.e., frequency plot) is often used to
visually illustrate both skewness and kurtosis.
If the distribution sways left (i.e., the peak is on the left and the long tail is on the right
side, and the mean is greater than median), then it produces a positive skewness measure,
and if the distribution sways right (i.e., the peak is on the right and the long tail is on the
left side, and the mean is smaller than median), then it produces a negative skewness
measure.
Specifically, kurtosis measures the degree to which a distribution is more or less peaked
than a normal distribution.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Regression, especially linear regression, is perhaps the most widely known and used
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
explanatory (input) variables. Once identified, this relationship between the variables can
be formally represented as a linear/additive function/equation.
Regression aims to capture the functional relationship between and among the
characteristics of the real world and describe this relationship with a mathematical model,
which may then be used to discover and understand the complexities of reality—explore
and explain relationships or forecast future occurrences.
On the other hand, regression attempts to describe the dependence of a response variable
on one (or more) explanatory variables where it implicitly assumes that there is a one-
way causal effect from the explanatory variable(s) to the response variable, regardless of
whether the path of effect is direct or indirect.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Simple versus Multiple Regression: If the regression equation is built between one
response variable and one explanatory variable, then it is called simple regression. For
instance, the regression equation built to predict/explain the relationship between a height
of a person (explanatory variable) and the weight of a person (response variable) is a
good example of simple regression.
Multiple regression is the extension of simple regression where the explanatory variables
are more than one. For instance, in the previous example, if we were to include not only
the height of the person but also other personal characteristics (e.g., BMI, gender,
ethnicity) to predict the weight of a person, then we would be performing multiple
regression analysis.
In both cases, the relationship between the response variable and the explanatory
variable(s) are linear and additive in nature.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Linear Regression
To understand the relationship between two variables, the simplest and most intuitive
thing that one can do is to create a scatter plot, where the y-axis represents the values of
the response variable, and the x-axis represents the values of the explanatory variable. A
scatter plot would show the changes in the response variable as a function of the changes
in the explanatory variable.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
In reality, it tries to find the signature (i.e., algebraic representation) of a straight line
passing through right in between the plotted dots (representing the observation/historical
data) in such a way that the distance between the dots and the line is minimized (the
predicted values on the theoretical regression line).
One of the most commonly used method is called the ordinary least squares (OLS)
method. The OLS method aims to minimize the sum of squared residuals (squared
vertical distances between the observation and the regression point) and leads to a
mathematical expression for the estimated value of the regression line.
Equation: Y = β0 + β1X
In this equation, β0 is called the intercept and β1 is called the slope. Once OLS
determines the values of these two coefficients, the simple equation can be used to
forecast the values of y for given values of x. The sign and the value of β1 also reveal the
direction and the strengths of relationship between the two variables.
If the model is of a multiple linear regression type, then there would be more coefficients
to be determined, one for each additional explanatory variable.
In the simplest sense, a well-fitting regression model results in predicted values close to
the observed data values.
For the numerical assessment, three statistical measures are often used in evaluating the
fit of a regression model. R2 (R-squared), the overall F-test, and the root mean square
error (RMSE).
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
All three of these measures are based on the sums of the square errors (how far the data
are from the mean and how far the data are from the model’s predicted values).
Different combinations of these two values provide different information about how the
regression model compares to the mean model.
Of the three, R2 has the most useful and understandable meaning because of its intuitive
scale. The value of R2 ranges from zero to one (corresponding to the amount of
variability explained in percentage) with zero indicating that the relationship and the
prediction power of the proposed model is not good, and one indicating that the proposed
model is a perfect fit that produces exact predictions (which is almost never the case).
The improvement in the regression model can be achieved by adding mode explanatory
variables, taking some of the variables out of the model, or using different data
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Model Development
1. Collect data.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
3. Test significance.
Model Evaluation
Linearity. This assumption states that the relationship between the response variable and
the explanatory variables are linear. That is, the expected value of the response variable is
a straight-line function of each explanatory variable, while holding all other explanatory
variables fixed. Also, the slope of the line does not depend on the values of the other
variables. It also implies that the effects of different explanatory variables on the
expected value of the response variable are additive in nature.
Independence (of errors). This assumption states that the errors of the response variable
are uncorrelated with each other. This independence of the errors is weaker than actual
statistical independence, which is a stronger condition and is often not needed for linear
regression analysis.
Normality (of errors). This assumption states that the errors of the response variable are
normally distributed. That is, they are supposed to be totally random and should not
represent any non random patterns.
Constant variance (of errors). This assumption, also called homoscedasticity, states that
the response variables have the same variance in their error, regardless of the values of
the explanatory variables. In practice, this assumption is invalid if the response variable
varies over a wide enough range/scale.
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Multicollinearity. This assumption states that the explanatory variables are not correlated
(i.e., do not replicate the same but provide a different perspective of the information
needed for the model). Multicollinearity can be triggered by having two or more perfectly
correlated explanatory variables presented to the model (e.g., if the same explanatory
variable is mistakenly included in the model twice, one with a slight transformation of the
same variable). A correlation-based data assessment usually catches this error.
Logistic Regression
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555
ACHARYA INSTITUTE OF TECHNOLOGY
Affiliated to Visvesvaraya Technological University, Belagavi, Govt. of Karnataka.
Approved by AICTE, New Delhi and Accredited by NBA (AE, BT, CSE, ECE, ME and MTE)
Department of Artificial Intelligence & Machine Learning
and
Computer Science & Engineering (Data Science)
Acharya Dr. Sarvepalli Radhakrishnan Road, Soladevanahalli, ACHIT Nagar P. O., Bangalore-560 107
[Link] Ph.: 080 22555555