Prof Shalabh Singh
Room No: 247
[Link]@ii
[Link]
Business Statistics
Introduction and Descriptive
Statistics
⚫ Using Statistics
⚫ Measures of Central Tendency
⚫ Measures of Variability
⚫ Skewness and Kurtosis
⚫ Methods of Displaying Data
⚫ Using the Computer
WHAT IS STATISTICS?
⚫ Statistics is a science that helps us make better
decisions in business and economics as well as in
other fields.
⚫ Statistics teaches us how to summarize, analyze,
and draw meaningful inferences from data that
then lead to improve decisions.
Using Statistics (Two Categories)
⚫ Descriptive Statistics ⚫ Inferential Statistics
✓ Collect ✓ Predict and forecast
✓ Organize values of population
✓ Summarize parameters
✓ Display ✓ Test hypotheses about
values of population
✓ Analyze
parameters
✓ Make decisions
Types of Data - Two Types
⚫ Qualitative - ⚫ Quantitative -
Categorical or Measurable or
Nominal: Countable:
Examples are- Examples are-
✓ Color ✓ Temperatures
✓ Gender ✓ Salaries
✓ Nationality ✓ Number of points
scored on a 100
point exam
• Population
– Collection of all the items or individuals about which you want
to draw a conclusion.
• Sample
– A portion of a population selected for analysis.
• Parameter
– A numerical measure that describes a characteristic of a
population.
• Statistic
– A numerical measure that describes a characteristic of a
sample.
Measures of location
Measures of Central Tendency
or Location
• Median ➢ Middle value when
sorted in order of
magnitude
➢ 50th percentile
• Mode ➢ Most frequently-
occurring value
• Mean ➢ Average
Measures of Central Tendency
• Data is clustered around a central point.
• Let the central point represent the data set.
• What is the most appropriate value for that central
point?
• First thing that comes to mind is : AVERAGE or MEAN
The Arithmetic Mean
• Sum of values of all observations (data points) divided by
number of elements in the data set
• Let xi = value of ith data point, and let n be the number of
data points
• Mean 𝑥ҧ = xi / n
Example
The magazine Forbes publishes
annually a list of the world’s
wealthiest individuals. For, 2007,
the net worth of the 20 richest
individuals, in $billions, is as
follows: (data is given on the next
slide). Also, the data has been
sorted in magnitude.
Example – Mean
Sorted
Billions Billions
33 18
26 18
n
24 18
538
21
19
18
19 x = xi = = 26.9
20
18
20
20 i =1 20
18 20
52 21
56 22
27 22
22 23
18 24
49 26
22 27
20 32
23 33
32 49
20 52
18 56
Sum = 538
Properties of Arithmetic Mean
• Every data set has a mean and it is unique
• Total value property: Mean times n equals the sum
of all observations.
• (Avg. daily production) (No. of days) = Total
Production
• Marginal addition
– Add a new item with xi > . How will it affect the mean?
– It pulls the average up.
– What if xi < . Average is pushed down.
Disadvantages of the arithmetic mean
• Outliers ??
– 3, 6, 4, 7, 5, 30 or 33, 42, 39, 45, 37, 28, 8
– Outliers may significantly affect the value of arithmetic
mean.
• Can not be computed for qualitative data
Median
• Middle value, middlemost or most central item
• Half the data items have value below the median,
and half above.
– Arrange n observations in increasing order.
– If n is odd, (n+1)/2th observation is the median.
– If n is even, median = average of (n/2)th and (n/2+1)th
observation.
1, 3, 4, 9, 12, 17, 22 n=7 (n+1)/2 = 4
Median = x4 = 9
1, 3, 4, 9, 12, 17
n=6 n/2 = 3 Median = (x3+x4)/2 = (4+9)/2 = 6.5
Example – Median
Sorted
Billions Billions
33 18
26 18 Median
24 18
21 18 50th Percentile
19 19
20 20
18 20 (20+1)50/100=10.5 22 + (.5)(0) = 22
18 20
52 21
56 22 Median
27 22
22
18
23
24
The median is the middle
49
22
26
27 value of data sorted in
20
23
32
33 order of magnitude. It is
32 49
20 52 the 50th percentile.
18 56
Advantages and disadvantages of the median
• Median, unlike the Mean, is not affected by outliers.
• Can be found for qualitative data also.
• Data must be arrayed first to calculate median
The Mode
• French word meaning fashion.
• Here it means most frequent: the value that is
repeated most often.
• Consider the following sorted data set.
– 26, 26, 27, 28, 28, 28, 28, 29, 29, 29, 30, 30, 31
– Which value occurs most frequently?
• Mode is not affected by the outliers.
• There may be no mode at all.
• There may be more than one modes.
Example - Mode
Mode = 18
The mode is the most frequently occurring value. It
is the value with the highest frequency.
Example - Mode
Mode = 18
The mode is the most frequently occurring value. It
is the value with the highest frequency.
Quartiles and Interquartile Range
⚫ The first quartile, Q1, (25th percentile) is
often called the lower quartile.
⚫ The second quartile, Q2, (50th
percentile) is often called the median
or the middle quartile.
⚫ The third quartile, Q3, (75th percentile)
is often called the upper quartile.
⚫ The interquartile range is the difference
between the first and the third quartiles.
Measures of dispersion
Measures of dispersion
• Is the Mean, Mode or Median an adequate
representation of data?
• It is measure of central tendency, which totally ignores
variability.
• Variability is a vital aspect of data.
• Enables us to judge the reliability of our measure of
central tendency
• To recognize the dispersion will help to tackle the
problem easily
– (quality control use, financial use)
• Comparing dispersions may be quite useful
Dispersion: Distance Measures
• Distance measures
– Range
– Interquartile range
• Range:
– The difference between the highest and the lowest observed
values
– Range = value of highest observation - value of lowest observation
• Disadvantages:
– Easy to find but usefulness is limited
– Ignores the nature of variation
– Heavily influenced by extreme values
– Changes drastically to one sample to the other in a given
population
– Open-ended distributions has no range
Measures of Variability
Dispersion: average deviation measures
• How much do the data points tend to deviate from
the mean?
• Can we define a measure for it ?
• Average deviation from some measure of central
tendency
– Variance
– Standard deviation
– Both deal with the average distance of any observation
from the mean.
Average Deviation?
• 1,2,3,4,5
• Mean : µ = 3
• Deviations? -2, -1, 0, 1, 2
• What is the average deviation? ZERO!
• Ignore the sign of deviation (absolute deviation)
• Average absolute deviation
= (∑|xi- µ |) / n = 6/5 = 1.2
Advantage:
It is a better measure of dispersion than the ranges
because it takes every observation into account.
Disadvantage:
It is mathematically inconvenient
Variance
• Absolute value function is mathematically
inconvenient. Is there an alternative?
• Square the deviations and then take average.
• Variance = 2 = [(xi-)2] / N
• Consider the data set: 1, 2, 3, 4, 5 , = 3
2 = (4+1+0+4+1) / 5 = 2.0
Standard deviation
• A significant change in variance to compute a useful
measure of variation
• Since we squared the deviations before taking average,
square root of this value is more meaningful.
• Moreover, for the variance the units are the squares of
the units.
• Standard deviation is the positive square root of the
variance.
• Standard Deviation = = ([(xi-)2] / N)
• This measure has nice mathematical properties, and is
universally accepted.
Variance and Standard Deviation
Population Variance Sample Variance
n
N
(x − )2 (x − x) 2
s = i =1
2
2 = i=1
N
(n − 1)
( x) ( )
2 2
N n
x
N n
i =1
−
x2 i =1 x − 2
= i=1 N =
i =1
n
N (n − 1)
= 2
s= s 2
X x- (x-)2
120 -22.5 506.25
125 -17.5 306.25
130 -12.5 156.25
135 -7.5 56.25
140 -2.5 6.25
145 2.5 6.25
150 7.5 56.25
155 12.5 156.25
160 17.5 306.25
165 22.5 506.25
= 142.5 2 = 206.25
= 14.3614
Example
Billions Sorted Billions
33 18
26 18
24 18
21 18
19 19
20 20
18 20
18 20
52 21
56 22
27 22
22 23
18 24
49 26
22 27
20 32
23 33
32 49
20 52
18 56
1-
31
Calculation of Sample Variance
x x−x (x − x) 2 x2
18 -8.9 79.21 324 n
18 -8.9 79.21 324 (x − x) 2
2657.8
18 -8.9 79.21 324 s2 = i =1
=
18 -8.9 79.21 324 (n − 1) (20 − 1)
19 -7.9 62.41 361 2657.8
= = 139.88421
20 -6.9 47.61 400 19
20 -6.9 47.61 400 2
20 -6.9 47.61 400
n x
x 2 − i =1
n
21 -5.9 34.81 441 n
22 -4.9 24.01 484 = i =1
22 -4.9 24.01 484 (n − 1)
23 -3.9 15.21 529 2
289444
24 -2.9 8.41 576 17130 − 538 17130 −
= 20 = 20
26 -0.9 0.81 676 (20 − 1) 19
27 0.1 0.01 729 17130 − 14472.2 2657.8
32 5.1 26.01 1024 = = = 139.88421
19 19
33 6.1 37.21 1089 2
49 22.1 488.41 2401 s= s = 139.88421 = 11.82
52 25.1 630.01 2704
56 29.1 846.81 3136
538 0 2657.8 17130
Data Summarization
How can we arrange/organize data??
• Arranging data using the data array and the frequency
distribution
– The Data Array
– The Frequency Distribution
• Data Array: Arranging values in increasing or decreasing
order.
– Lowest and highest values can be quickly noticed
– Can easily be divided into sections
– Distance between succeeding values can be observed
– Values appearing more frequently can be noticed easily
• The Frequency Distribution: A better way to arrange
data.
– Shows the number of observations from the data set that fall
into each of the classes
– tabulation of the frequencies of each value (or range of values)
Methods of Displaying Data
⚫ Bar Graphs
✓ Heights of rectangles represent group frequencies
⚫ Frequency Polygons
✓ Height of line represents frequency
⚫ Pie Charts
✓ Categories represented as percentages of total
Example
• 15, 12, 14, 17, 12, 16, 13, 17, 6, 13, 16, 24, 15, 22, 11, 16, 14,
23, 12, 15, 16, 12, 7, 27, 15, 14, 13, 14, 17, 10
Example
• 15, 12, 14, 17, 12, 16, 13, 17, 6, 13, 16, 24, 15, 22, 11, 16, 14,
23, 12, 15, 16, 12, 7, 27, 15, 14, 13, 14, 17, 10
• Data Array:
– 6, 7, 10, 11, 12, 12, 12, 12, 13, 13, 13, 14, 14, 14, 14, 15, 15, 15, 15, 16,
16, 16, 16, 17, 17, 17, 22, 23, 24, 27
• Frequency Distribution:
Class Frequency
6-10 3
11-15 16
16-20 7
21-25 3
26-30 1
Data Presentation
• There are a wide variety of ways to illustrate frequency distributions
including histograms, relative frequency histograms, density
histograms, and cumulative frequency distributions
• Histogram/Bar Chart
– Histogram consists of a series of rectangles whose widths are defined by
the limits of the classes, and whose heights are determined by the
frequency in each interval.
– Histogram depicts many attributes of the data, including location, spread,
and symmetry
Bar Chart
18
16
14
12
10
8 Frequency
6
4
2
0
6-10 11-15 16-20 21-25 26-30
Frequency Polygon
• Represent each class with class mark
• Convert the bar chart into a graph. How?
15
10
0
3 8 13 18 23 28 33
6-10 11-15 16-20 21-25 26-30
• Useful for comparing two distributions.
• It is easy to superimpose two or more polygons, unlike the bar charts.
Example
Solution