0% found this document useful (0 votes)
3 views47 pages

Exploratory Data Analysis - Notes

The document outlines a course on Exploratory Data Analysis, detailing the topics covered, including data visualization, statistics, and types of statistics. It explains the origins and definitions of statistics, its applications, and various scales of measurement. Additionally, it discusses data visualization techniques and chart selection criteria for effective data representation.

Uploaded by

rathodskrillix
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views47 pages

Exploratory Data Analysis - Notes

The document outlines a course on Exploratory Data Analysis, detailing the topics covered, including data visualization, statistics, and types of statistics. It explains the origins and definitions of statistics, its applications, and various scales of measurement. Additionally, it discusses data visualization techniques and chart selection criteria for effective data representation.

Uploaded by

rathodskrillix
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Bachelor of Computer Application

Faculty of Science

BCA1320F01:Exploratory Data Analysis


2 credits

Ms. Priti Singh


Teaching Assistant
Department of Statistics, Faculty of Science
TOPICS OF PRESENTATION
INCLUDES…

INTRODUCTION TO DATA VISULIZATION,


STATISTICS CENTRAL TENDENCY SKEWNESS
1 3 5

2 4 6
TYPES OF DISPERSION COORELATION AND
STATISTICS REGRESSION
What is statistics?
• Aggregation of facts

• Summarizing lots of information and plotting

• Getting meaningful and useful insights from data

• Collecting, organizing, analyzing, and interpreting data.


Origin and general meaning
• Statistics has its origin in Latin word Status, Italian word Statista and German term
Statistik all of which mean “Political State”. In ancient times the beginning of Statistics
was made to meet the requirements of State primarily and hence the name.

• The science of collectiong, organizing, presenting, analyzing, and interpreting data to


assist in making more effective decisions

• Statistical analysis – used to manipulate, summarize, and investigate data, so that


useful decision-making information results.

• Meaning in Plural and Singular sense


Plural Sense Singular Sense

• Under this, the Statistics refers to numerical statement • Under this, Statistics refers to science in which we
of facts related to any field of enquiry such as data deal with techniques and methods for collecting,
related to income, expenditure, population etc. classifying, presenting, analyzing and interpreting the
data
• In the sense of numerical data or Statistical data.
• In the meaning of Statistical methods
• “Statistics are numerical statements of facts in any
department of enquiry placed in relation to each • “Statistics may be defined as the collection,
other”-Bowley presentation, analysis and interpretation of numerical
data.”-Croxton and Cowden
• “By statistics we mean quantitative data affected to a
marked extent by multiplicity of cause”- Yule and • “Statistics is the science which deals with the
Kendall collection, classification and tabulation of numerical
facts as a basis for the explanation, description and
comparison of phenomena- Lovitt
Application of Statistics
Epidemiology
Astro Statistics
Actuarial Science
Government
Environmental Statistics
Business and commerce Geo Statistics

STATISTICS Biostatistics
Agriculture Econometrics

Operation Research Jurimetrics


Machin Learning

Forensic Statistics Quantitative psychology


Quality Control
Type of Statistics
• Descriptive statistics: It describes the data and consists of methods and techniques to
explain characteristics of data. The methods can either be graphical or computational.

• Data Visualization
• Measure of central tendency
• Measure of Dispersion
• Skewness

• Inferential Statistics: It deals with methods which describe the characteristics of population
or making decisions concerning population on the basis of sample results.
• Theory of Estimation
• Hypothesis Testing
Classification of Data

Basic
Classification

Geographical Qualitative

Chronological Quantitative
Types of Data
Characteristics of Nominal Scale
•A nominal scale variable is classified into two or more categories. In this measurement mechanism, the
answer should fall into either of the classes.
•It is qualitative. The numbers are used here to identify the objects.
•The numbers don’t define the object characteristics. The only permissible aspect of numbers in the
nominal scale is “counting.”
Example:
An example of a nominal scale measurement is given below:
What is your gender?
M- Male
F- Female

Characteristics of the Ordinal Scale


•The ordinal scale shows the relative ranking of the variables
•It identifies and describes the magnitude of a variable
•Along with the information provided by the nominal scale, ordinal scales give the rankings of those
variables
•The interval properties are not known
•The surveyors can quickly analyse the degree of agreement concerning the identified order of variables
Example:
•Ranking of school students – 1st, 2nd, 3rd, etc.
•Ratings in restaurants
•Evaluating the frequency of occurrences
•Assessing the degree of agreement
Characteristics of Interval Scale:
•The interval scale is quantitative as it can quantify the difference between the values
•It allows calculating the mean and median of the variables
•To understand the difference between the variables, you can subtract the values between the variables
•The interval scale is the preferred scale in Statistics as it helps to assign any numerical values to arbitrary
assessment such as feelings, calendar types, etc.
Example:
•Likert Scale
•Net Promoter Score (NPS)
•Bipolar Matrix Table

Characteristics of Ratio Scale:


•Ratio scale has a feature of absolute zero
•It doesn’t have negative numbers, because of its zero-point feature
•It affords unique opportunities for statistical analysis. The variables can be orderly added, subtracted,
multiplied, divided. Mean, median, and mode can be calculated using the ratio scale.
•Ratio scale has unique and useful properties. One such feature is that it allows unit conversions like
kilogram – calories, gram – calories, etc.
Example:
An example of a ratio scale is:
What is your weight in Kgs?
Levels of Measurements
Text Visual

Eyes Ears Skin Smell Taste


What is Data Visualization?
• Data Visualization or Data Viz is any visual representation that helps in understanding,
organizing, and analyzing data. It helps your numbers tell a meaningful story and helps
your readers understand it.

• Its Importance
• Easy to understand by layman

• Long lasting impression

• Visualized Data is processed faster

• Identification of insights and unusual patterns

• Helps is taking quick decisions


What information do we need to select a chart?
By asking yourself the following questions:

1. Who is my audience?
2. What insights do I want my readers to gain?
3. What should be the range of my axis?
4. Should I display values over time or among groups?
5. What information many categories do I need?
6. How many data points do I need for each category?
What information do we need to select a chart?
1. Comparing Data
Table 1. To compare individual or precise values, 2. The value involves various units of measure, 3. Displaying
quantitative information is more important than trends

[Link] display and compare the rank of values and focus on the extremes, 2. Short data category labels,
Column Chart
3. Items on the chart have less than seven categories and comparing with respect to time 4. Example: Revenue
per landing page, Sales by year, etc.

1. To show if the data values have attained a particular goal, 2. Items on the chart have over seven but less than
Bar Chart 15 categories, 3. Display negative numbers, 4. Long data category labels, 5. Example: Website visitors per
country, Customers won per role, etc.

[Link] compare many different items and show the composition of each item being compared
Stack Chart 2. To display visual aggregate all of the categories in a group when the size of individual categories is not
important, 3. Example: Sales by region, Sales by product, etc.

1. To compare multiple items or groups on various attributes, 2. To visualize comparisons of quality data,
Radar Chart 3. The number of attributes should be at least three but less than 10, 4. Example: Comparing the features of two or
three cars, Rainfall by month, etc.
2. Tracking data over a period of time
1. To show how a category changes over time, 2. To display a continuous dataset having a high number of data points,
Line Chart
3. Trend based visualizations, 4. Example: Web traffic over time, number of purchase returns by month, etc.

1. To track changes over time in two or more related groups that make up one whole category, [Link] emphasis is on the
Area Chart cumulative data volume rather than the data points, 3. Example: New vs. Returning website visitors, Sales by month for
two or more products, etc.

3. Analyzing compositions and part-to-whole relationships


1. To display parts of a whole in percentages, 2. Show relative proportion, 3. All categories add up to 100,
Pie Chart 4. Example: Customer survey, Spending in a month, etc.

4. Studying the distribution of data


1. To show the distribution of data points among in a variable, 2. Example: students enrollement with respect to year,
Histogram
age group, etc.
Chart

1. To show the presence or absence of a relationship between two variables, 2. Use for correlation and distribution
Scatter plot analysis, 3. Not more than two categories, 4. To display data distribution and clustering trends, 5. Example: Customer
satisfaction by response time, etc.
Bubble Chart 1. Same as scatter plot but with 3 or 4 categories, 2. Example: Products purchased by age and
gender, etc.

1. To show the relationship between a group and a matrix of two categories, 2. To provide rating
Heat Map
information To analyze a category across a matrix of data, 3. Example: Risk Matrix, Users by the
time of day and region, etc.

Other charts that are frequently used are as follows:


1. Donut Chart

1. Same as Pie chart but with more rings can be added

2. Stem and Leaf plot


1. Helpful in identifying mode and
distribution of data. 2. Preferable if data
points are less then 200
Dataset Dataset: Virat Kholi Batting Performance in ODI matches

Inning
Year s Runs Balls Outs Avg SR HS 50 100 4s 6s Dot %
2008 5 159 239 5 31.8 66.5 54 1 0 21 1 71.5 Questions
2009 8 325 385 6 54.2 84.4 107 2 1 36 3 51.7
2010 23 995 1,169 21 47.4 85.1 118 7 3 90 4 48.3
2011 34 1,381 1,614 29 47.6 85.6 117 8 4 127 7 46.8
1. Year on year performance?
2012 17 1,026 1,094 15 68.4 93.8 183 3 5 92 7 43.6
2013 30 1,268 1,300 24 52.8 97.5 115 7 4 138 20 48.1 2. Which year have highest contributed to the
2014 20 1,054 1,058 18 58.6 99.6 139 5 4 94 20 43.6
total runs?
2015 20 623 773 17 36.6 80.6 138 1 2 44 8 49.0
2016 10 739 739 8 92.4 100.0 154 4 3 62 8 40.2
2017 26 1,460 1,473 19 76.8 99.1 131 7 6 136 22 42.8 3. Did performance improved over a period of
2018 14 1,202 1,172 9 133.6 102.6 160 3 6 123 13 41.7
time?
2019 25 1,377 1,429 23 59.9 96.4 123 7 5 133 8 41.8
2020 9 431 467 9 47.9 92.3 89 5 0 35 5 41.1
2021 3 129 149 3 43.0 86.6 66 2 0 10 1 41.6
2022 11 302 347 11 27.5 87.0 113 2 1 32 2 48.1
2023 6 338 252 5 67.6 134.1 166 0 2 32 10 30.2
Total 261 12,809 13,660 222 57.7 93.8 183 64 46 1,205 139 45.0

Source: [Link]
Solution to the previous dataset 1600
1600 1,460
1,381 1,377 1400
1400 1,268
1,202 1200
1200 1,054
995 1,026 1000
1000
800
RUNS

800 739
623 600
600
431 400
400 325 302 338
159 200
200 129
0
0 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023
2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023
Series2 Linear (Series2)
YEARS
2017 11%
Chart Title 2011 11%
8% 8%
2019 11% 1600
6% 2013 10%
8% 2% 1400
5% 3%
2018 9%
1200
2014 8%
9% 1%
2012 8% 1000

Runs
13% 2010 8% 800
1%
2016 6% 600
10% 3%
2015 5%
400
11%
2020 3%
3%
11% 2023 3% 200
11%
2009 3% 0
2017 2011 2019 2013 2018 2014 2012 2010 2022 2% 0 5 10 15 20 25 30 35 40
Innings
2016 2015 2020 2023 2009 2022 2008 2021 2008 1%
2021 1%
Measure of Central Tendency
• The data collected in any statistical investigation, by any methods of collection, known as
raw data. These raw data are organized systematically into different homogenous class,
which is known as classification of data.

• Classification of quantitative phenomenon is called frequency distribution. The


organization of the data pertaining to quantitative phenomenon involves the following four
stages:

1. Raw Data
2. Discrete Ungrouped Data
3. Discrete grouped Data (Inclusive Classes)
4. Continuous Grouped Data (Exclusive Classes)
Example:
• Raw Data

• Set or series of individuals X1, X2,………, Xn X1, X2,………, Xn

• Discrete ungrouped Data


X X1 X2 : Xn Total X 1 2 3 4 5 Total
• The data is of the form F f1 f2 : fn N F 3 6 10 4 2 25

• Discrete Grouped Data


Inclusive X1 X2 : Xn Total Inclusive 0-5 6-10 11- 16- Total
• Also Know as Inclusive Classes Classes 15 20
classes F 6 15 10 2 33
F f1 f2 : fn N
• The data is of the form

• Continuous Grouped Data


Exclusive X1 X2 : X n Total Exclusive 0-5 5-10 10- 15- Total
• Also Know as Exclusive Classes Classes 15 20
classes
• The data is of the form F f1 f2 : fn N F 6 15 10 2 33
• Since, one of the important aims of statistical analysis is to determine various numerical
measures which describe the inherent characteristics of a frequency distribution.

• It is observed that the mass of data have a tendency to concentrate around some data
point, generally this data point is the central value of the data is called Measures of Central
Tendency or Measures of Location.

• It is necessary to identify or calculate this typical value, called the central value or an
average to describe or project the characteristic of the entire data set.

• Croxton and Cowden have defined average as “An average value is a single value within the
range of the data that is used to represent all of the values in the series. Since an average
is somewhere within the range of the data, it is sometimes called a measure of central
value.”
Requisites of an Ideal measure of Central Tendency

1. It should be easy to understand and calculate

2. It should be based on all observations

3. It should not be affected much by extreme observations

4. It should be affected as little as possible by fluctuation of sampling

5. It should be rigidly defined

6. It should be suitable for further mathematical treatment


Measure of Central Tendency

Average Measures Positional Measure

Arithmetic Geometric Harmonic


Mode Quartiles Percentiles
Mean Mean Mean

Simple Mean
Median Deciles
Weighted Mean
Averages
Arithmetic Mean (A.M) Geometric mean (G.M) Harmonic Mean (H.M)

Definition Sum of observations divided by nth root of product of n non zero Reciprocal of A.M
no. of observations. Abbreviated observations
as A.M. and is denoted by 𝑥ҧ

Use General purpose 1)The concept of G.M. is used in the The harmonic mean is particularly
construction of Index number. useful for computation of average
2) Since G.M. ≤ A.M., therefore G.M. is rates and ratios. Such rates and
useful in those cases where smaller ratios are generally used to
observations are to be given importance. express relations between two
Such cases usually occur in social and different types of measuring units
economic areas of study. that can be expressed reciprocally
3) The G.M. of a data set is useful in
estimating the average rate of growth in the
initial value of an observation per unit
period. For example, it is useful in finding
the percentage increase in sales, profit,
production, population, and so on.
Formula
(Raw Data)
Arithmetic Mean Geometric mean Harmonic Mean

Ungrouped
Data ; where f = Frequency

Grouped
Data ; where f = Frequency and
mi =mid value of class

Merits 1) It is rigidly defined. 2) It is easy to 1) It is rigidly defined. 2) It is based 1) It is rigidly defined. 2) It is based on
calculate. 3) It is based upon all the on all the observations. 3) It is all observations. 3) It is suitable for
observations. 4) It is capable of further capable of further algebraic further mathematical treatment. 4) It
algebraic treatment. (i.e. possible to treatment is not affected much by sampling
find combined mean). 5. It is least fluctuations.
affected of sampling fluctuations.
Demerits 1) It is very much affected by extreme 1) It is difficult to understand and 1) Useful when small values are to be
values.. 2) It can’t be calculated for compute. 2) It cannot be given very high weightage. 2) difficult
open-end classes. 3) It can’t be located determined; if there are negative to compute and understand. 3) Can’t
graphically. 4) The mean cannot be values or any of the values is zero. be computed if positive and negative
calculated for qualitative characteristics 3) Not possible to locate graphically. values are present or one or more
such as intelligence, honesty, beauty, or zeros are present.
loyalty.
Application 1. The concept of G.M. is used in the construction of The harmonic mean is particularly useful
Index number. 2. Since G.M. ≤ A.M., therefore G.M. is for computation of average rates and
useful in those cases where smaller observations are ratios. Such rates and ratios are generally
to be given importance. Such cases usually occur in used to express relations between two
social and economic areas of study. 3. The G.M. of a different types of measuring units that
data set is useful in estimating the average rate of can be expressed reciprocally.
growth in the initial value of an observation per unit
period. For example, it is useful in finding the
percentage increase in sales, profit, production,
population, and so on.
Note 1) If any one of the observation is zero then G.M is 1) If any observation is Zero then H.M is
zero. 2) If any one of the observation is negative then not defined. 2) Harmonic mean is
G.M is imaginary especially useful in averaging rates and
ratios where time factor is variable and
the act being performed e.g., for finding
average speed of vehicles, typist etc.

Relation ship between Positional averages


1) A.M ≥ G.M ≥ H.M
2) G.M = Sqrt(A.M * H.M)
Properties of A.M

1) The sum of the deviations of all observations from their arithmetic mean is always zero.

2) The sum of squared deviations of all the observations is minimum when deviation is taken from actual arithmetic
mean.

3) If we replace each individual observation in the data by the constant then mean is the constant itself.

4) The arithmetic mean of first n natural numbers = (n+1)/2

5) If 𝑋1 and 𝑋2 be arithmetic mean of two groups of observations N1 and N2 then the combined mean of these two
groups can be computed by

Generalized formula
Weighted A.M

In the computation of simple arithmetic average assumption is that all the items in the distribution are of equal
importance. However, in practice, it is possible to come across situation where relative importance of all the items of
the distribution is not same. In such cases, due weightage is to be given to various item weighted mean is computed. For
example, if it is desired to have an idea of the change in the cost of living of a certain group of people, then the simple
arithmetic average of the prices of the commodities consumed by the people will not do, as all commodities are not
equally importance; e.g. items like wheat, rice pluses, fuels, housing, lighting etc. are more important than cigarettes,
confectionary, cosmetics, etc. Hence, different items should be assigned weights according to their relative importance
for the computation of mean, which will be weighted mean.

Remark
The weighted arithmetic mean should be used when
1) the importance of all the numerical values in the given data set is not equal.
2) the frequencies of various classes are widely varying.
3) there is a change either in the proportion of numerical values or in the proportion of their frequencies. • ratios,
percentages, or rates are being averaged
Positional Measures
Mode (A.M) Median (G.M) Quartiles (H.M)

Definition Mode is the value which occurs most The median is that value of variable which The values which divide the given
frequently in a set of observation and divides the distribution into two equal distribution into four equal parts
around which the other items of the parts, one part comprising all the values are known as Quartiles. There will
set cluster densely greater and the other, all values less than be three such points Q1 , Q2 and
median Q3, Such that Q1  Q2 ≤ Q3 .
Use Mode is especially useful in finding
the most popular size in studies
relating to marketing, trade, business
and industry. It is the appropriate
average to be used to find the ideal
size e.g., in business forecasting, in
the manufacture of shoes or
readymade garments, in sales, in
production, etc

Formula M0 = that value of variable which First arrange the observation in ascending Q1 = (N+1)/4, Q2 = 2((N+1)/4), Q3
(Raw Data) occur more frequently in data set. (increasing) order. = 3((N+1)/4)

Me = ((n+1)/2)th Obs. ; If n is odd


[(n/2)th Obs. + ((n/2)+1))th Obs.]
= ; If
2
n is even
Ungrouped M0 = that value of variable which First find the cumulative frequency
Data corresponds to highest frequency. (C.F) less than type. 𝑀𝑒 : that value of
variable which corresponds to C.F. just
greater than or equal to (N/2)
Grouped
Data

Where L: lower limit or boundary of


modal class 𝑓0: Frequency of class above
modal class 𝑓1: Frequency of modal class C = cumulative frequency of previous
f
𝑓2: Frequency of class below modal class class
𝑐: Class width of modal class
Merits 1) It is easy to calculate, easy to 1. It is easy to understand. 2. It is easy
understand. to calculate. 3. It is not affected by
2) It is not affected by extreme values. 3) extreme value. 4. It can be computed
It can be determined in open-end classes, in case of open end classes. 5. It is
at least modal class should be closed one most suitable average in a study of
4) It can be represented graphically qualitative data. 6. It can be located
(Histogram). graphically (Ogive Curve).
Demerits 1. It is not based on all the observations. 1. It is not capable for further algebraic
2. It is not capable of further algebraic treatment. 2. It is not based on all
treatment. 3. As compared to mean and observations. 3. It is affected more by
median, it is affected to a greater extent sampling fluctuations than the
by sampling fluctuations. arithmetic mean. 4. It requires
arranging the data before it can be
found, which tedious work is.
Numerical
Q1) The one-sided Vitcos fare (in Rs.) of five selected [Link]. Students are recorded as follows: 10, 5, 15, 8, and 12
Calculate arithmetic mean of the data.

Q2) Given the following frequency distribution of age 10th standard students of a particular school.

Age 13 14 15 16 17
Frequency 2 5 13 7 3

Q3) The following data shows distance covered by 100 persons to perform their routine jobs.

Distance (km) 0-10 10-20 20-30 30-40


Frequency 10 20 40 30
Measure of Dispersion
• One of the important characteristic of distribution is Central Tendency, gives one single value
that represents the entire data. Another important characteristic of distribution is to describe
the dispersion of data. The dispersion also means scatteredness, spread or variation of the
observations.
• The averages alone cannot adequately describe a set of observations, unless all the observations
are the same. It is necessary to describe the variability or dispersion of the observation. In two or
more distributions the central value may be the same but still there can be wide disparities in
the formulation of distribution.
• The extent to which the individual observations differ on an average from mean or any
other measure of central value is called measure of dispersion or measure of variation.
• As these measures give an average of the differences of the observations included in a group
from an average of these items, they are also known as “averages of second order”. Note that
Measures of central values are, therefore, called the “averages of first order”.
Definition of Dispersion:
• According to Spiegel – “The degree to which numerical data tend to spread about an
average value is called the variation or dispersion of the data.”
• Functions / Uses of measuring Dispersion
1. To determine the reliability of an Average.
2. To serve as a Basis for the Control of the Variability.
3. To Compare Two or More series with Regard to their Variability.
4. To Facilitate the Use of other Statistical Measures.

• Requisites of an Ideal Measure of Variation Good measures of dispersion should possess,


as far as possible, the following properties:
1. It should be simple to understand.
2. It should be easy to compute.
3. It should be rigidly defined.
4. It should be based on each and every observations of the distribution.
5. It should be capable to further algebraic treatment.
6. It should have sampling stability.
7. It should not be unduly affected by extreme observations
Measure of Dispersion

Relative Measures Absolute Measure

Coefficient of Range Range (R)

Coefficient of Quartile Deviation Quartile Deviation (Q.D)

Mean Deviation (M.D)


Coefficient of Mean Deviation
Standard Deviation (S.D)
Coefficient of Variation (C.V) and Variance
Relative Measure Absolute Measure
• A relative measure of dispersion is the ratio of an • Absolute Measures of dispersion are expressed in
absolute measure of dispersion to an appropriate the same statistical unit in which the original data
average. It is called a coefficient of dispersion, are given such as rupees, kilograms, kilometers
because “coefficient” means a pure number that is etc. These values may be used to compare the
independent of the unit of measurement. variations in two distributions provided the
• These measures are calculated for the comparison variables are expressed in the same units and of
of dispersion in two or more than two sets of the same average size.
observations. These measures are free of the units • They give the answers in the same units as the
in which the original data is measured. units of the original observations. When the
observations are in kilograms, the absolute
• In case the two sets of data are expressed in
measure is also in kilograms. If we have two sets
different units, however, such as quintals of sugar
of observations, we cannot always use the absolute
versus tones of sugarcane, or if the average size is
measures to compare their dispersion.
very different such as manager’s salary versus
workers’ salary, the absolute measures of
dispersion are not comparable. In such cases
measures of relative dispersion should be used.
Range (R) Quartile Deviation (Q.D) Mean Deviation (M.D) Standard Deviation (S.D)

Definitio the difference between the largest The Quartile deviation (Q.D.) The mean deviation or the The standard deviation is
n and the smallest values is called is a measure of variability, average deviation is defined as defined as the positive square
as the range. based on dividing a data set the mean of the absolute root of the mean of the square
into quartiles. It is based on the deviations of observations deviation taken from
lower quartile Q1 and the from some suitable average arithmetic mean of the data. It
upper quartile Q3 . The which may be the arithmetic is denoted by σ (sigma)
difference Q3 -Q1 is called the mean, the median or the mode.
inter quartile range.
Absolute Range: L-S IQR = Q3 – Q1
Measure Where; L = Largest value of
observation. (Q3 – Q1)
S = Smallest value of Semi-IQR =
𝟐
observation.

Relative Coefficient of Range: (𝑳−𝑺)


(𝑳+𝑺
Measure
Raw Data

Ungrouped
Data

Grouped
Data

Where m: mid value of classes


Merits 1. It is simplest and easiest 1. It is easy to calculate and simple 1. It is easy to calculate and simple 1. It is rigidly defined and based on
measure of dispersion. 2. It is to understand. 2. It is not affected to understand. 2. It is less affected all observations. 2. It is capable to
readily comprehensible and it by extreme values of the variable by the extreme values of variable. further algebraic treatment. 3. It is
requires very little calculations. 3. as it is concerned with the central 3. It is based on all observations in not affected by sampling
It is free of measure of central half portion of the distribution. 3. the distribution. fluctuations.
tendency. It is not at all affected by open end
class data.
Demerits 1. It is not suitable measure of 1. It ignores completely the 1. It ignores the negative deviation 1. It is difficult to understand and
dispersion as it affected by portions below the lower quartile and treats them as positive which is calculate. 2. It is very much affected
extreme vales. 2. It is not suitable and above upper quartile. 2. It is not justified mathematically. 2. It is by extreme values
measure of dispersion for the data not capable of further not a satisfactory measure when the
with open ended classes. 3. It is mathematical treatment. 3. It is deviations are taken from the mode.
very crude measure, as it is based greatly affected by the fluctuations 3. It is not suitable when the class
on two extreme values and not on in the sampling. intervals are open end type. 4. The
other values. mean deviation cannot be used in
statistical inference.
Application 1. It is useful in
studying the variations in
the prices of share and
stocks. 2. It is useful in
studying weather conditions
where minimum and
maximum temperature is
identified. 3. It is widely
used in industrial quality
control.
Note The difference Q3 -Q1 divided The mean deviation is 1. Standard deviation of
by 2 is called semi-inter- minimum when it is taken constant number is zero. That
quartile range or the quartile from median. σ 𝑋 − 𝐴 ≥ σ 𝑋 − is if 𝑋𝑖 = 𝑐 for all 𝑖 then S.D. X =
deviation. 𝑀𝑒 for ungrouped data σ 𝑓 𝑋 − 0. 2. S.D is independent of
𝐴 ≥ σ 𝑓 𝑋 − 𝑀𝑒 for frequency change of origin but not scale.
data Where, A is any constant
other than Median.
Skewness
• The voluminous raw data cannot be easily understood; hence, we calculate the measures of central tendencies and
obtain a representative figure. From the measures of variability, we can know that whether most of the items of the
data are close to or away from these central tendencies. But these statistical means and measures of variation are not
enough to draw sufficient description about the data. Another aspect of the data is to know its symmetry. The
symmetry of data is well studied by the knowledge of the "Skewness.

• Literal meaning of skewness is ‘lack of symmetry’. Study of skewness is help to have an idea about the shape of the
curve which can be draw with the help of the given frequency distribution. The frequency curve of the distribution is
not a symmetric bell-shaped curve but it is stretched more to one side than other then it is called skewed distribution.
A frequency distribution for which the curve has longer tail towards the right is said to be positively skewed and if
the longer tail lies towards the left, it is said to be negatively skewed
Symmetric Distribution: For symmetric distribution curve falls at same rate from the highest peak. Thus frequency
curve has same tail from mean. For such curve Mean = Median = Mode.

Positively skewed distribution: For a positively skewed distribution curve rises rapidly, reaches the maximum and falls
slowly. In other words, if the frequency curve has longer tail to right the distribution is known as positively skewed
distribution and for a positively skewed distribution Mean > Median > Mode.

Negatively skewed distribution: A negatively skewed distribution curve rises slowly, reaches its maximum and falls
rapidly. In other words, if the frequency curve has longer tail to left the distribution is known as negatively skewed
distribution and for negatively skewed distribution Mean < Median < Mode.
Measures of Skewness: A measure which gives the extent of asymmetry is known as the measure of
‘skewness’. Measures of Skewness are categories in two ways.
Absolute measures Relative measures
• Sk = (Mean – Mode) • (ii) Relative Measures For comparing two or more
distributions for skewness compute relative measures of
• Sk = 3(Mean –Median) , when mode is not defined
skewness called coefficient of skewness which are pure
• Sk = Q3+ Q1 – 2 Median numbers independent of the units of measurement.
• Karl Pearson’s Coefficient of skewness: Coefficient of
• Absolute measures are not much practical because : skewness
• They involve the units of measurement, hence
cannot be used for comparative study of the
distribution measured in different units of • If mode is not uniquely defined then Coefficient of
measurements. skewness
• Even if the same units of measurements, one may
come across different distributions which have
identical absolute measures, but which vary widely
in the measures of central tendency and • If, Sk>0 = +ve Sk, Sk<0 -ve Sk, and Sk=0 symmetric
dispersion.
Bowley’s coefficient of skewness:
In case of open end distribution Bowely’s coefficient of skewness is used.

If,
Sk > 0 then distribution is positively skewed
Sk < 0 then distribution is negatively skewed
Sk = 0 then distribution is symmetrical distributed
Limits for Bowley’s Coefficient of Skewness is −1 𝑡𝑜 + 1

Theoretically the values of Sk varies between ±3, but for a moderately skewed distribution, value of Sk varies between
±1
Uses of Skewness
1. It helps in finding out the nature and degree of concentration whether it is in higher or the lower values.
2. The empirical (it include statistical data) relationship between Mean, Median and Mode is based on the assumption
of a moderately skewed distribution. The measure of skewness will show to what amount such empirical relationship
would holds good.
3. It helps in knowing the distribution is normal or not. Many statistical measures are based on the assumption of
normal distribution

You might also like