DESCRIPTIVE
STATISTICS
Suri Gurumurthi, Ph.D.
1
Populations and Samples
Population - all items of interest for a particular decision or investigation
All voters who are over the age of 65
All subscribers to Hulu
Sample - a subset of the population
A group of voters who are likely to vote, over the age of 65
A group of subscribers who liked “Handmaiden’s Tale”
The purpose of sampling is to obtain sufficient information to draw a valid inference about a
population.
4-2
Descriptive Analytics: Inference
Identify Draw from
goals sample
Start Population Sample
Draw Collect raw
conclusions data and
summarize
Make inferences
about population
Population Sample
parameters Statistics
Source: Bennett, Briggs, and Triola (2014), Statistical Reasoning for Everyday Life.
3
Average Measure (Location)
Arithmetic Mean
For a population of size N:
For a sample of n observations:
Excel function: =AVERAGE(data range)
Property of the mean:
Outliers can affect the value of the mean.
4-4
Median Measure (Location)
Median - middle value of the data when arranged from least to greatest
Middle value separating the greater and lesser halves of a data set
Median (1, 2, 2, 3, 4, 7, 9) = 3
Median (1, 2, 3, 4, 5, 6, 8, 9) = (4+5)/2 = 4.5
When to use Median, as opposed to Mean?
When the data are skewed
E.g. A company’s sales figures are skewed because of one month
Impacted by Covid-19
We don’t wish to consider the impact of outliers on our inference
Source: [Link]
4-5
Mode Measures (Location)
Mode - Most frequent value in a data set
Mode (1, 2, 2, 3, 4, 7, 9) = 2
Frequency of LA Rams Roster Salaries
40
35
30
25
20
15
10
0
500000 1000000 1500000 2000000 2500000 3000000 3500000 4000000 4500000 5000000 25000000 More
4-6
Mid-Range Measure (Location)
Player Position Base Salary
Aaron Donald DT $17,000,000
Natrez Patrick LB $595,588
Midrange = Average of greatest and least values
Midrange = Average(17000000, 595588) =$8,797,794
4-7
Measures of Dispersion
Dispersion refers to the degree of variation in the data.
Range is the difference between the maximum and minimum data
values.
Range = MAX – MIN
Range for LA Rams Salaries = 17000000- 595588 = $16,404,412
Samples from two populations with equal mean but
different dispersion.
Interquartile Range (IQR) is difference between the third and first
quartiles.
Variance is an average of the squared deviations form the mean (uses
all data values).
Standard Deviation is the square root of the variance.
4-8
Player Position Cap Hit
Aaron Donald DT $25,000,000
Robert Woods WR $8,175,000
Andrew Whitworth LT $6,666,666
Leonard Floyd ILB $6,666,666
Johnny Hekker P $4,687,000
Measures of Dispersion
Austin Blythe G $3,900,000
Troy Hill CB $4,443,750
Josh Reynolds WR $2,295,006
John Johnson FS $2,322,438
Samson Ebukam OLB $2,286,272
Tyler Higbee TE $9,125,000
Michael Brockers DE $3,833,333
Cooper Kupp WR $3,371,690
Gerald Everett TE $1,923,242
Jalen Ramsey CB $6,203,000
Austin Corbett G $1,163,490
Jake McQuaide LS $1,080,000
Jared Goff QB $28,842,682
Malcolm Brown RB $1,312,500
For a population: Rob Havenstein RT $6,240,000
Morgan Fox DE $825,000
Kenny Young ILB $750,000
Darious Williams CB $750,000
In Excel: =VAR.P(data) Ogbonnia Okoronkwo
Johnny Mundt
OLB
TE
$818,634
$750,000
VAR.P(CapHit) = 27,330,656,518,764.60 Sebastian Joseph
Brian Allen
DT
C
$788,809
$921,734
Taylor Rapp SS $1,062,203
Darrell Henderson RB $959,580
David Long CB $921,041
Nsimba Webster WR $675,000
Coleman Shelton C $675,000
Nick Scott FS $694,332
For a sample:
Troy Reeder ILB $679,000
Micah Kiser ILB $749,494
Justin Hollins OLB $675,000
Greg Gaines DT $831,653
Bobby Evans RT $880,545
David Edwards G $741,130
In Excel: =VAR.S(data) John Wolford
Sam Sloman
QB
K
$685,000
$628,873
VAR.S (CapHit) = 27,856,246,067,202.30 Jachai Polite OLB $610,000
Xavier Jones RB $613,000
Van Jefferson WR $1,020,206
Trishton Jackson WR $614,333
Brycen Hopkins TE $773,283
Jordan Fuller FS $652,678
Raymond Calais RB $610,000
Terrell Burgess SS $818,074
Eric Banks DE $610,666
Tremayne Anchrum G $628,873
4-9 Cam Akers RB $1,122,370
Natrez Patrick LB $595,588
Measures of Dispersion
For a population:
STDEV.P(data)
STDEV.P(CapHit) = $5,227,873.04
For a sample:
STDEV.S(data)
STDEV.S(CapHit) = $5,277,901.67
4-10
Co-efficient of Variation
Provides a relative measure of dispersion
Sometimes measured as a percentage.
Provides a relative measure of risk to return.
Return to risk = 1/CV
4-11
Chebyshev’s Theorem
k 1-1/k^2
1.1 17.4%
1.2 30.6%
1.3 40.8%
1.4 49.0%
For any SAMPLE, the proportion of values that lie within k (k > 1) standard 1.5 55.6%
deviations of the mean is at least 1 – 1/k2 1.6 60.9%
1.7 65.4%
Empirically though, for many samples or data sets that are bell-shaped 1.8 69.1%
1.9 72.3%
Approximately 68% of the observations fall within one standard deviation of the mean; 2 75.0%
that is, in the interval [Mean – Stdev, Mean + Stdev] 2.1 77.3%
Approximately 95% fall within [Mean – 2Stdev, Mean + 2Stdev] 2.2 79.3%
2.3 81.1%
Approximately 99.7% fall within [Mean – 3Stdev, Mean + 3Stdev] 2.4 82.6%
2.5 84.0%
2.6 85.2%
1-1/k^2 2.7 86.3%
2.8 87.2%
100.0% 2.9 88.1%
90.0%
3 88.9%
80.0%
3.1 89.6%
70.0%
3.2 90.2%
60.0%
3.3 90.8%
50.0%
3.4 91.3%
40.0%
3.5 91.8%
30.0%
3.6 92.3%
20.0%
3.7 92.7%
10.0%
3.8 93.1%
0.0%
1.1 1.3 1.5 1.7 1.9 2.1 2.3 2.5 2.7 2.9 3.1 3.3 3.5 3.7 3.9 3.9 93.4%
4 93.8%
4-12 1-1/k^2
The z-score Z-scores of select players Cap Hits on LA Rams
z-score =
z-score = STANDARDIZE(x, mean, standard
deviation)
z-score
-1 0 1 2 3 4 5 6
4-13
Outliers
The Mean and Range are sensitive to outliers. Player z-score of CapHit
How do we identify outliers? Aaron Donald 4.183807
Jared Goff 4.911877
With many samples:
z-scores greater than +3 or less than -3
extreme outliers are more than 3*IQR to the left of Q1 or right of Q3
mild outliers are between 1.5*IQR and 3*IQR to the left of Q1 or right of Q3
There is no standard definition of what constitutes an outlier.
4-14
Measures of Shape
Skewness describes lack of symmetry.
Coefficient of Skewness =
Skew of CapHit =SKEW(CapHit) = 3.8618 Cap Hit
CS is negative for left-skewed data.
CS is positive for right-skewed data.
|CS| > 1 suggests high degree of
skewness.
0.5 ≤ |CS| ≤ 1 suggests moderate
skewness.
|CS| < 0.5 suggests relative symmetry.
$0 $5,000,000 $10,000,000 $15,000,000 $20,000,000 $25,000,000 $30,000,000 $35,000,000
4-15
Measures of Shape
Kurtosis refers to peakedness or flatness.
Coefficient of Kurtosis =
Kurtosis of CapHit=KURT(CapHit)
CK < 3 indicates the data is somewhat flat with a wide degree of
dispersion.
CK > 3 indicates the data is somewhat peaked with less dispersion.
4-16
Statistical Thinking
Philosophy of learning and improvement of a system
Interconnected processes
variation exists in all systems and processes
better performance through understanding and reducing variation
Business Analytics provides insights into facts and relationships
Enables better decisions.
4-17