Exploratory Data Analysis - Notes
Exploratory Data Analysis - Notes
Faculty of Science
2 4 6
TYPES OF DISPERSION COORELATION AND
STATISTICS REGRESSION
What is statistics?
• Aggregation of facts
• Under this, the Statistics refers to numerical statement • Under this, Statistics refers to science in which we
of facts related to any field of enquiry such as data deal with techniques and methods for collecting,
related to income, expenditure, population etc. classifying, presenting, analyzing and interpreting the
data
• In the sense of numerical data or Statistical data.
• In the meaning of Statistical methods
• “Statistics are numerical statements of facts in any
department of enquiry placed in relation to each • “Statistics may be defined as the collection,
other”-Bowley presentation, analysis and interpretation of numerical
data.”-Croxton and Cowden
• “By statistics we mean quantitative data affected to a
marked extent by multiplicity of cause”- Yule and • “Statistics is the science which deals with the
Kendall collection, classification and tabulation of numerical
facts as a basis for the explanation, description and
comparison of phenomena- Lovitt
Application of Statistics
Epidemiology
Astro Statistics
Actuarial Science
Government
Environmental Statistics
Business and commerce Geo Statistics
STATISTICS Biostatistics
Agriculture Econometrics
• Data Visualization
• Measure of central tendency
• Measure of Dispersion
• Skewness
• Inferential Statistics: It deals with methods which describe the characteristics of population
or making decisions concerning population on the basis of sample results.
• Theory of Estimation
• Hypothesis Testing
Classification of Data
Basic
Classification
Geographical Qualitative
Chronological Quantitative
Types of Data
Characteristics of Nominal Scale
•A nominal scale variable is classified into two or more categories. In this measurement mechanism, the
answer should fall into either of the classes.
•It is qualitative. The numbers are used here to identify the objects.
•The numbers don’t define the object characteristics. The only permissible aspect of numbers in the
nominal scale is “counting.”
Example:
An example of a nominal scale measurement is given below:
What is your gender?
M- Male
F- Female
• Its Importance
• Easy to understand by layman
1. Who is my audience?
2. What insights do I want my readers to gain?
3. What should be the range of my axis?
4. Should I display values over time or among groups?
5. What information many categories do I need?
6. How many data points do I need for each category?
What information do we need to select a chart?
1. Comparing Data
Table 1. To compare individual or precise values, 2. The value involves various units of measure, 3. Displaying
quantitative information is more important than trends
[Link] display and compare the rank of values and focus on the extremes, 2. Short data category labels,
Column Chart
3. Items on the chart have less than seven categories and comparing with respect to time 4. Example: Revenue
per landing page, Sales by year, etc.
1. To show if the data values have attained a particular goal, 2. Items on the chart have over seven but less than
Bar Chart 15 categories, 3. Display negative numbers, 4. Long data category labels, 5. Example: Website visitors per
country, Customers won per role, etc.
[Link] compare many different items and show the composition of each item being compared
Stack Chart 2. To display visual aggregate all of the categories in a group when the size of individual categories is not
important, 3. Example: Sales by region, Sales by product, etc.
1. To compare multiple items or groups on various attributes, 2. To visualize comparisons of quality data,
Radar Chart 3. The number of attributes should be at least three but less than 10, 4. Example: Comparing the features of two or
three cars, Rainfall by month, etc.
2. Tracking data over a period of time
1. To show how a category changes over time, 2. To display a continuous dataset having a high number of data points,
Line Chart
3. Trend based visualizations, 4. Example: Web traffic over time, number of purchase returns by month, etc.
1. To track changes over time in two or more related groups that make up one whole category, [Link] emphasis is on the
Area Chart cumulative data volume rather than the data points, 3. Example: New vs. Returning website visitors, Sales by month for
two or more products, etc.
1. To show the presence or absence of a relationship between two variables, 2. Use for correlation and distribution
Scatter plot analysis, 3. Not more than two categories, 4. To display data distribution and clustering trends, 5. Example: Customer
satisfaction by response time, etc.
Bubble Chart 1. Same as scatter plot but with 3 or 4 categories, 2. Example: Products purchased by age and
gender, etc.
1. To show the relationship between a group and a matrix of two categories, 2. To provide rating
Heat Map
information To analyze a category across a matrix of data, 3. Example: Risk Matrix, Users by the
time of day and region, etc.
Inning
Year s Runs Balls Outs Avg SR HS 50 100 4s 6s Dot %
2008 5 159 239 5 31.8 66.5 54 1 0 21 1 71.5 Questions
2009 8 325 385 6 54.2 84.4 107 2 1 36 3 51.7
2010 23 995 1,169 21 47.4 85.1 118 7 3 90 4 48.3
2011 34 1,381 1,614 29 47.6 85.6 117 8 4 127 7 46.8
1. Year on year performance?
2012 17 1,026 1,094 15 68.4 93.8 183 3 5 92 7 43.6
2013 30 1,268 1,300 24 52.8 97.5 115 7 4 138 20 48.1 2. Which year have highest contributed to the
2014 20 1,054 1,058 18 58.6 99.6 139 5 4 94 20 43.6
total runs?
2015 20 623 773 17 36.6 80.6 138 1 2 44 8 49.0
2016 10 739 739 8 92.4 100.0 154 4 3 62 8 40.2
2017 26 1,460 1,473 19 76.8 99.1 131 7 6 136 22 42.8 3. Did performance improved over a period of
2018 14 1,202 1,172 9 133.6 102.6 160 3 6 123 13 41.7
time?
2019 25 1,377 1,429 23 59.9 96.4 123 7 5 133 8 41.8
2020 9 431 467 9 47.9 92.3 89 5 0 35 5 41.1
2021 3 129 149 3 43.0 86.6 66 2 0 10 1 41.6
2022 11 302 347 11 27.5 87.0 113 2 1 32 2 48.1
2023 6 338 252 5 67.6 134.1 166 0 2 32 10 30.2
Total 261 12,809 13,660 222 57.7 93.8 183 64 46 1,205 139 45.0
Source: [Link]
Solution to the previous dataset 1600
1600 1,460
1,381 1,377 1400
1400 1,268
1,202 1200
1200 1,054
995 1,026 1000
1000
800
RUNS
800 739
623 600
600
431 400
400 325 302 338
159 200
200 129
0
0 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023
2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023
Series2 Linear (Series2)
YEARS
2017 11%
Chart Title 2011 11%
8% 8%
2019 11% 1600
6% 2013 10%
8% 2% 1400
5% 3%
2018 9%
1200
2014 8%
9% 1%
2012 8% 1000
Runs
13% 2010 8% 800
1%
2016 6% 600
10% 3%
2015 5%
400
11%
2020 3%
3%
11% 2023 3% 200
11%
2009 3% 0
2017 2011 2019 2013 2018 2014 2012 2010 2022 2% 0 5 10 15 20 25 30 35 40
Innings
2016 2015 2020 2023 2009 2022 2008 2021 2008 1%
2021 1%
Measure of Central Tendency
• The data collected in any statistical investigation, by any methods of collection, known as
raw data. These raw data are organized systematically into different homogenous class,
which is known as classification of data.
1. Raw Data
2. Discrete Ungrouped Data
3. Discrete grouped Data (Inclusive Classes)
4. Continuous Grouped Data (Exclusive Classes)
Example:
• Raw Data
• It is observed that the mass of data have a tendency to concentrate around some data
point, generally this data point is the central value of the data is called Measures of Central
Tendency or Measures of Location.
• It is necessary to identify or calculate this typical value, called the central value or an
average to describe or project the characteristic of the entire data set.
• Croxton and Cowden have defined average as “An average value is a single value within the
range of the data that is used to represent all of the values in the series. Since an average
is somewhere within the range of the data, it is sometimes called a measure of central
value.”
Requisites of an Ideal measure of Central Tendency
Simple Mean
Median Deciles
Weighted Mean
Averages
Arithmetic Mean (A.M) Geometric mean (G.M) Harmonic Mean (H.M)
Definition Sum of observations divided by nth root of product of n non zero Reciprocal of A.M
no. of observations. Abbreviated observations
as A.M. and is denoted by 𝑥ҧ
Use General purpose 1)The concept of G.M. is used in the The harmonic mean is particularly
construction of Index number. useful for computation of average
2) Since G.M. ≤ A.M., therefore G.M. is rates and ratios. Such rates and
useful in those cases where smaller ratios are generally used to
observations are to be given importance. express relations between two
Such cases usually occur in social and different types of measuring units
economic areas of study. that can be expressed reciprocally
3) The G.M. of a data set is useful in
estimating the average rate of growth in the
initial value of an observation per unit
period. For example, it is useful in finding
the percentage increase in sales, profit,
production, population, and so on.
Formula
(Raw Data)
Arithmetic Mean Geometric mean Harmonic Mean
Ungrouped
Data ; where f = Frequency
Grouped
Data ; where f = Frequency and
mi =mid value of class
Merits 1) It is rigidly defined. 2) It is easy to 1) It is rigidly defined. 2) It is based 1) It is rigidly defined. 2) It is based on
calculate. 3) It is based upon all the on all the observations. 3) It is all observations. 3) It is suitable for
observations. 4) It is capable of further capable of further algebraic further mathematical treatment. 4) It
algebraic treatment. (i.e. possible to treatment is not affected much by sampling
find combined mean). 5. It is least fluctuations.
affected of sampling fluctuations.
Demerits 1) It is very much affected by extreme 1) It is difficult to understand and 1) Useful when small values are to be
values.. 2) It can’t be calculated for compute. 2) It cannot be given very high weightage. 2) difficult
open-end classes. 3) It can’t be located determined; if there are negative to compute and understand. 3) Can’t
graphically. 4) The mean cannot be values or any of the values is zero. be computed if positive and negative
calculated for qualitative characteristics 3) Not possible to locate graphically. values are present or one or more
such as intelligence, honesty, beauty, or zeros are present.
loyalty.
Application 1. The concept of G.M. is used in the construction of The harmonic mean is particularly useful
Index number. 2. Since G.M. ≤ A.M., therefore G.M. is for computation of average rates and
useful in those cases where smaller observations are ratios. Such rates and ratios are generally
to be given importance. Such cases usually occur in used to express relations between two
social and economic areas of study. 3. The G.M. of a different types of measuring units that
data set is useful in estimating the average rate of can be expressed reciprocally.
growth in the initial value of an observation per unit
period. For example, it is useful in finding the
percentage increase in sales, profit, production,
population, and so on.
Note 1) If any one of the observation is zero then G.M is 1) If any observation is Zero then H.M is
zero. 2) If any one of the observation is negative then not defined. 2) Harmonic mean is
G.M is imaginary especially useful in averaging rates and
ratios where time factor is variable and
the act being performed e.g., for finding
average speed of vehicles, typist etc.
1) The sum of the deviations of all observations from their arithmetic mean is always zero.
2) The sum of squared deviations of all the observations is minimum when deviation is taken from actual arithmetic
mean.
3) If we replace each individual observation in the data by the constant then mean is the constant itself.
5) If 𝑋1 and 𝑋2 be arithmetic mean of two groups of observations N1 and N2 then the combined mean of these two
groups can be computed by
Generalized formula
Weighted A.M
In the computation of simple arithmetic average assumption is that all the items in the distribution are of equal
importance. However, in practice, it is possible to come across situation where relative importance of all the items of
the distribution is not same. In such cases, due weightage is to be given to various item weighted mean is computed. For
example, if it is desired to have an idea of the change in the cost of living of a certain group of people, then the simple
arithmetic average of the prices of the commodities consumed by the people will not do, as all commodities are not
equally importance; e.g. items like wheat, rice pluses, fuels, housing, lighting etc. are more important than cigarettes,
confectionary, cosmetics, etc. Hence, different items should be assigned weights according to their relative importance
for the computation of mean, which will be weighted mean.
Remark
The weighted arithmetic mean should be used when
1) the importance of all the numerical values in the given data set is not equal.
2) the frequencies of various classes are widely varying.
3) there is a change either in the proportion of numerical values or in the proportion of their frequencies. • ratios,
percentages, or rates are being averaged
Positional Measures
Mode (A.M) Median (G.M) Quartiles (H.M)
Definition Mode is the value which occurs most The median is that value of variable which The values which divide the given
frequently in a set of observation and divides the distribution into two equal distribution into four equal parts
around which the other items of the parts, one part comprising all the values are known as Quartiles. There will
set cluster densely greater and the other, all values less than be three such points Q1 , Q2 and
median Q3, Such that Q1 Q2 ≤ Q3 .
Use Mode is especially useful in finding
the most popular size in studies
relating to marketing, trade, business
and industry. It is the appropriate
average to be used to find the ideal
size e.g., in business forecasting, in
the manufacture of shoes or
readymade garments, in sales, in
production, etc
Formula M0 = that value of variable which First arrange the observation in ascending Q1 = (N+1)/4, Q2 = 2((N+1)/4), Q3
(Raw Data) occur more frequently in data set. (increasing) order. = 3((N+1)/4)
Q2) Given the following frequency distribution of age 10th standard students of a particular school.
Age 13 14 15 16 17
Frequency 2 5 13 7 3
Q3) The following data shows distance covered by 100 persons to perform their routine jobs.
Definitio the difference between the largest The Quartile deviation (Q.D.) The mean deviation or the The standard deviation is
n and the smallest values is called is a measure of variability, average deviation is defined as defined as the positive square
as the range. based on dividing a data set the mean of the absolute root of the mean of the square
into quartiles. It is based on the deviations of observations deviation taken from
lower quartile Q1 and the from some suitable average arithmetic mean of the data. It
upper quartile Q3 . The which may be the arithmetic is denoted by σ (sigma)
difference Q3 -Q1 is called the mean, the median or the mode.
inter quartile range.
Absolute Range: L-S IQR = Q3 – Q1
Measure Where; L = Largest value of
observation. (Q3 – Q1)
S = Smallest value of Semi-IQR =
𝟐
observation.
Ungrouped
Data
Grouped
Data
• Literal meaning of skewness is ‘lack of symmetry’. Study of skewness is help to have an idea about the shape of the
curve which can be draw with the help of the given frequency distribution. The frequency curve of the distribution is
not a symmetric bell-shaped curve but it is stretched more to one side than other then it is called skewed distribution.
A frequency distribution for which the curve has longer tail towards the right is said to be positively skewed and if
the longer tail lies towards the left, it is said to be negatively skewed
Symmetric Distribution: For symmetric distribution curve falls at same rate from the highest peak. Thus frequency
curve has same tail from mean. For such curve Mean = Median = Mode.
Positively skewed distribution: For a positively skewed distribution curve rises rapidly, reaches the maximum and falls
slowly. In other words, if the frequency curve has longer tail to right the distribution is known as positively skewed
distribution and for a positively skewed distribution Mean > Median > Mode.
Negatively skewed distribution: A negatively skewed distribution curve rises slowly, reaches its maximum and falls
rapidly. In other words, if the frequency curve has longer tail to left the distribution is known as negatively skewed
distribution and for negatively skewed distribution Mean < Median < Mode.
Measures of Skewness: A measure which gives the extent of asymmetry is known as the measure of
‘skewness’. Measures of Skewness are categories in two ways.
Absolute measures Relative measures
• Sk = (Mean – Mode) • (ii) Relative Measures For comparing two or more
distributions for skewness compute relative measures of
• Sk = 3(Mean –Median) , when mode is not defined
skewness called coefficient of skewness which are pure
• Sk = Q3+ Q1 – 2 Median numbers independent of the units of measurement.
• Karl Pearson’s Coefficient of skewness: Coefficient of
• Absolute measures are not much practical because : skewness
• They involve the units of measurement, hence
cannot be used for comparative study of the
distribution measured in different units of • If mode is not uniquely defined then Coefficient of
measurements. skewness
• Even if the same units of measurements, one may
come across different distributions which have
identical absolute measures, but which vary widely
in the measures of central tendency and • If, Sk>0 = +ve Sk, Sk<0 -ve Sk, and Sk=0 symmetric
dispersion.
Bowley’s coefficient of skewness:
In case of open end distribution Bowely’s coefficient of skewness is used.
If,
Sk > 0 then distribution is positively skewed
Sk < 0 then distribution is negatively skewed
Sk = 0 then distribution is symmetrical distributed
Limits for Bowley’s Coefficient of Skewness is −1 𝑡𝑜 + 1
Theoretically the values of Sk varies between ±3, but for a moderately skewed distribution, value of Sk varies between
±1
Uses of Skewness
1. It helps in finding out the nature and degree of concentration whether it is in higher or the lower values.
2. The empirical (it include statistical data) relationship between Mean, Median and Mode is based on the assumption
of a moderately skewed distribution. The measure of skewness will show to what amount such empirical relationship
would holds good.
3. It helps in knowing the distribution is normal or not. Many statistical measures are based on the assumption of
normal distribution