0% found this document useful (0 votes)
8 views17 pages

Descriptive Statistics Overview

Uploaded by

yschl001
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views17 pages

Descriptive Statistics Overview

Uploaded by

yschl001
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DESCRIPTIVE

STATISTICS
Suri Gurumurthi, Ph.D.

1
Populations and Samples

 Population - all items of interest for a particular decision or investigation


 All voters who are over the age of 65
 All subscribers to Hulu

 Sample - a subset of the population


 A group of voters who are likely to vote, over the age of 65
 A group of subscribers who liked “Handmaiden’s Tale”

 The purpose of sampling is to obtain sufficient information to draw a valid inference about a
population.

4-2
Descriptive Analytics: Inference
Identify Draw from
goals sample
Start Population Sample

Draw Collect raw


conclusions data and
summarize

Make inferences
about population
Population Sample
parameters Statistics

Source: Bennett, Briggs, and Triola (2014), Statistical Reasoning for Everyday Life.
3
Average Measure (Location)

Arithmetic Mean
 For a population of size N:

 For a sample of n observations:

 Excel function: =AVERAGE(data range)

 Property of the mean:

 Outliers can affect the value of the mean.

4-4
Median Measure (Location)
 Median - middle value of the data when arranged from least to greatest
 Middle value separating the greater and lesser halves of a data set

Median (1, 2, 2, 3, 4, 7, 9) = 3

Median (1, 2, 3, 4, 5, 6, 8, 9) = (4+5)/2 = 4.5

 When to use Median, as opposed to Mean?


 When the data are skewed
 E.g. A company’s sales figures are skewed because of one month
 Impacted by Covid-19
 We don’t wish to consider the impact of outliers on our inference

Source: [Link]

4-5
Mode Measures (Location)
 Mode - Most frequent value in a data set

Mode (1, 2, 2, 3, 4, 7, 9) = 2

Frequency of LA Rams Roster Salaries


40

35

30

25

20

15

10

0
500000 1000000 1500000 2000000 2500000 3000000 3500000 4000000 4500000 5000000 25000000 More
4-6
Mid-Range Measure (Location)

Player Position Base Salary

Aaron Donald DT $17,000,000


Natrez Patrick LB $595,588

 Midrange = Average of greatest and least values


 Midrange = Average(17000000, 595588) =$8,797,794

4-7
Measures of Dispersion

 Dispersion refers to the degree of variation in the data.

 Range is the difference between the maximum and minimum data


values.
 Range = MAX – MIN
 Range for LA Rams Salaries = 17000000- 595588 = $16,404,412
Samples from two populations with equal mean but
different dispersion.
 Interquartile Range (IQR) is difference between the third and first
quartiles.
 Variance is an average of the squared deviations form the mean (uses
all data values).
 Standard Deviation is the square root of the variance.

4-8
Player Position Cap Hit
Aaron Donald DT $25,000,000
Robert Woods WR $8,175,000
Andrew Whitworth LT $6,666,666
Leonard Floyd ILB $6,666,666
Johnny Hekker P $4,687,000

Measures of Dispersion
Austin Blythe G $3,900,000
Troy Hill CB $4,443,750
Josh Reynolds WR $2,295,006
John Johnson FS $2,322,438
Samson Ebukam OLB $2,286,272
Tyler Higbee TE $9,125,000
Michael Brockers DE $3,833,333
Cooper Kupp WR $3,371,690
Gerald Everett TE $1,923,242
Jalen Ramsey CB $6,203,000
Austin Corbett G $1,163,490
Jake McQuaide LS $1,080,000
Jared Goff QB $28,842,682
Malcolm Brown RB $1,312,500
 For a population: Rob Havenstein RT $6,240,000
Morgan Fox DE $825,000
Kenny Young ILB $750,000
Darious Williams CB $750,000

 In Excel: =VAR.P(data) Ogbonnia Okoronkwo


Johnny Mundt
OLB
TE
$818,634
$750,000
 VAR.P(CapHit) = 27,330,656,518,764.60 Sebastian Joseph
Brian Allen
DT
C
$788,809
$921,734
Taylor Rapp SS $1,062,203
Darrell Henderson RB $959,580
David Long CB $921,041
Nsimba Webster WR $675,000
Coleman Shelton C $675,000
Nick Scott FS $694,332

 For a sample:
Troy Reeder ILB $679,000
Micah Kiser ILB $749,494
Justin Hollins OLB $675,000
Greg Gaines DT $831,653
Bobby Evans RT $880,545
David Edwards G $741,130
 In Excel: =VAR.S(data) John Wolford
Sam Sloman
QB
K
$685,000
$628,873
 VAR.S (CapHit) = 27,856,246,067,202.30 Jachai Polite OLB $610,000
Xavier Jones RB $613,000
Van Jefferson WR $1,020,206
Trishton Jackson WR $614,333
Brycen Hopkins TE $773,283
Jordan Fuller FS $652,678
Raymond Calais RB $610,000
Terrell Burgess SS $818,074
Eric Banks DE $610,666
Tremayne Anchrum G $628,873
4-9 Cam Akers RB $1,122,370
Natrez Patrick LB $595,588
Measures of Dispersion

 For a population:

 STDEV.P(data)
 STDEV.P(CapHit) = $5,227,873.04

 For a sample:

 STDEV.S(data)
 STDEV.S(CapHit) = $5,277,901.67

4-10
Co-efficient of Variation

 Provides a relative measure of dispersion

 Sometimes measured as a percentage.


 Provides a relative measure of risk to return.
 Return to risk = 1/CV

4-11
Chebyshev’s Theorem
k 1-1/k^2
1.1 17.4%
1.2 30.6%
1.3 40.8%
1.4 49.0%
 For any SAMPLE, the proportion of values that lie within k (k > 1) standard 1.5 55.6%
deviations of the mean is at least 1 – 1/k2 1.6 60.9%
1.7 65.4%
 Empirically though, for many samples or data sets that are bell-shaped 1.8 69.1%
1.9 72.3%
 Approximately 68% of the observations fall within one standard deviation of the mean; 2 75.0%
that is, in the interval [Mean – Stdev, Mean + Stdev] 2.1 77.3%
 Approximately 95% fall within [Mean – 2Stdev, Mean + 2Stdev] 2.2 79.3%
2.3 81.1%
 Approximately 99.7% fall within [Mean – 3Stdev, Mean + 3Stdev] 2.4 82.6%
2.5 84.0%
2.6 85.2%
1-1/k^2 2.7 86.3%
2.8 87.2%
100.0% 2.9 88.1%
90.0%
3 88.9%
80.0%
3.1 89.6%
70.0%
3.2 90.2%
60.0%
3.3 90.8%
50.0%
3.4 91.3%
40.0%
3.5 91.8%
30.0%
3.6 92.3%
20.0%
3.7 92.7%
10.0%
3.8 93.1%
0.0%
1.1 1.3 1.5 1.7 1.9 2.1 2.3 2.5 2.7 2.9 3.1 3.3 3.5 3.7 3.9 3.9 93.4%
4 93.8%
4-12 1-1/k^2
The z-score Z-scores of select players Cap Hits on LA Rams

 z-score =

 z-score = STANDARDIZE(x, mean, standard


deviation)

z-score

-1 0 1 2 3 4 5 6

4-13
Outliers

 The Mean and Range are sensitive to outliers. Player z-score of CapHit
 How do we identify outliers? Aaron Donald 4.183807
Jared Goff 4.911877
 With many samples:
 z-scores greater than +3 or less than -3
 extreme outliers are more than 3*IQR to the left of Q1 or right of Q3
 mild outliers are between 1.5*IQR and 3*IQR to the left of Q1 or right of Q3
 There is no standard definition of what constitutes an outlier.

4-14
Measures of Shape
 Skewness describes lack of symmetry.

 Coefficient of Skewness =

 Skew of CapHit =SKEW(CapHit) = 3.8618 Cap Hit

 CS is negative for left-skewed data.


 CS is positive for right-skewed data.
 |CS| > 1 suggests high degree of
skewness.
 0.5 ≤ |CS| ≤ 1 suggests moderate
skewness.
 |CS| < 0.5 suggests relative symmetry.
$0 $5,000,000 $10,000,000 $15,000,000 $20,000,000 $25,000,000 $30,000,000 $35,000,000

4-15
Measures of Shape
 Kurtosis refers to peakedness or flatness.

 Coefficient of Kurtosis =

 Kurtosis of CapHit=KURT(CapHit)

 CK < 3 indicates the data is somewhat flat with a wide degree of


dispersion.
 CK > 3 indicates the data is somewhat peaked with less dispersion.

4-16
Statistical Thinking

 Philosophy of learning and improvement of a system


 Interconnected processes
 variation exists in all systems and processes
 better performance through understanding and reducing variation

 Business Analytics provides insights into facts and relationships


 Enables better decisions.

4-17

You might also like