0% found this document useful (0 votes)
3 views36 pages

Descriptive Statistics Overview

The document discusses descriptive statistics for continuous and categorical features, including measures of central tendency such as mean and median, as well as measures of variation like range, variance, and standard deviation. It also covers the calculation of percentiles and inter-quartile range, and highlights the importance of understanding populations versus samples in statistical analysis. Examples are provided to illustrate these concepts, particularly in the context of basketball squad data.

Uploaded by

ninabhatt30
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views36 pages

Descriptive Statistics Overview

The document discusses descriptive statistics for continuous and categorical features, including measures of central tendency such as mean and median, as well as measures of variation like range, variance, and standard deviation. It also covers the calculation of percentiles and inter-quartile range, and highlights the importance of understanding populations versus samples in statistical analysis. Examples are provided to illustrate these concepts, particularly in the context of basketball squad data.

Uploaded by

ninabhatt30
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Descriptive Statistics Data Visualization Summary

Descriptive Statistics
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

The arithmetic mean (or sample mean or just mean) of a


set of n values for a feature a, a1 , a2 . . . an , is denoted by
the symbol a, and is calculated as:

n
1X
a= ai
n
i=1
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149

150 163 145 140 157 151 140 149

Figure: The members of a school basketball squad. The dashed grey


line shows the arithmetic mean of the players’ heights.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149

150 163 145 140 157 151 140 149

Figure: The members of a school basketball squad. The dashed grey


line shows the arithmetic mean of the players’ heights.

1
H EIGHT = × (150 + 163 + 145 + 140 + 157 + 151 + 140 + 149)
8
= 149.375
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

The arithmetic mean is one measure of the central


tendency of a sample (for our purposes a sample is just a
set of values for a feature in an ABT).
Any measure of central tendency is, however, just an
approximation.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
Suppose our basketball squad manage to sign a ringer
measuring in at 229cm

150 163 145 140 157 151 140 149 229

The arithmetic mean for the full group is 158.235cm and no


longer represents the central tendency of the group.
An unusually large or small value like this is referred to as
an outlier - the arithmetic mean is very sensitive to
outliers.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

The median of a set of values can be calculated by


ordering the values from lowest to highest and selecting
the middle value.
If there is an even number of values in the sample then the
median is obtained by calculating the arithmetic mean of
the middle two values.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example

140 140 145 149 150 151 157 163 229

Figure: The members of the school basketball squad ordered by


height, the dashed grey line shows the median.

ID 4 7 3 8 1 6 5 2 9
Height 140 140 145 149 150 151 157 163 229
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

We also measure the variation in our data.


In essence, most of statistics, and in turn analytics, is
about describing and understanding variation.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

The simplest measure of variation is the range:

range = max(a) − min(a)


Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
What is the range of the heights of the two basketball squads?

ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149

ID 1 2 3 4 5 6 7 8
Height 192 102 145 165 126 154 123 188
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
What is the range of the heights of the two basketball squads?

range = 163 − 140 = 23

range = 192 − 102 = 90


Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

The variance of a sample measures the average


difference between each value in a sample and the mean
of that sample.
The variance of the n values of a feature a, a1 , a2 . . . an , is
denoted var (a) and is calculated as:
n
X
(ai − a)2
i=1
var (a) =
n−1
Variance measures average diff between each value and
mean

We square each diff, as some diffs can be positive and other::


can be negative.

Denominator is n-1, so that sample variance is an unbiased


estimator of population variance.

We say that an estimator is unbiased, if sample variance on


an average equals population variance.

If Denominator is n, then we have a biased estimator that on


an average underestimates population variance.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
What is the variance of the heights of the two basketball
squads?

ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149

ID 1 2 3 4 5 6 7 8
Height 192 102 145 165 126 154 123 188
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example

(150 − 149.375)2 + (163 − 149.375)2 + . . . + (149 − 149.375)2


var (H EIGHT) =
8−1
= 63.125

(192 − 149.375)2 + (102 − 149.375)2 + . . . + (188 − 149.375)2


var (H EIGHT) =
8−1
= 1, 011.41071
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

The standard deviation, sd, of a sample is calculated by


taking the square root of the variance of the sample:

p
sd(a) = var (a) (1)
v
u n
uX
u
u (ai − a)2
t i=1
= (2)
n−1
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
What is the standard deviation of the heights of the two
basketball squads?

ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149

ID 1 2 3 4 5 6 7 8
Height 192 102 145 165 126 154 123 188
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example

sd(H EIGHT) = 63.125
= 7.9451 . . .

p
sd(H EIGHT) = 1, 011.41071
= 31.8026 . . .
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Percentiles are another useful measure of the variation of


i
the values for a feature: a proportion of 100 of the values in
a sample take values equal to or lower than the i th
percentile of that sample.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

To calculate the i th percentile of the n values of a feature a,


a1 , a2 . . . an :
First order the values in ascending order and then multiply
i
n by 100 to determine the index.
If the index is a whole number we take the value at that
position in the ordered list of values as the i th percentile.
If index is not a whole number then we interpolate the
value for the i th percentile as:

i th percentile = (1 − index_f ) × aindex_w + index_f × aindex_w+1

where index_w is the whole part of index, index_f is the


fractional part of index and aindex_w is the value in the
ordered list at position index_w.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
ID 2 7 5 3 6 4 8 1
Height 102 123 126 145 154 165 188 192

102122 126 145 154 165 188 192

What is the 25th percentile of the heights of the basketball


squad?
What is the 80th percentile of the heights of the basketball
squad?
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
To calculate the 25th percentile we first calculate index as
25 th
100 × 8 = 2. So, the 25 percentile is the second value in
the ordered list which is 123.
To calculate the 80th percentile we first calculate index as
80
100 × 8 = 6.4. Because index is not a whole number we
set index_w to the whole part of index, 6, and index_f to
the fractional part, 0.4. Then we can calculate the 80th
percentile as:
(1 − 0.4) × 165 + 0.4 × 188 = 174.2
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

We can use percentiles to describe another measure of


variation know as the inter-quartile range.
The inter-quartile range is calculated as the difference
between the 25th percentile and the 75th percentile.1

1
These percentiles are also known the lower quartile (or 1st quartile) and
upper quartile (or 3rd quartile) hence the name inter-quartile range.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Continuous Features

Example
For the heights of the first basketball team the inter-quartile
range is 151 − 140 = 11, while for the second team it is
165 − 123 = 42.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Categorical Features

For categorical features we are interested primarily in


frequency counts and proportions.
The frequency count of each level of a categorical feature is
calculated by counting the number of times that level
appears in the sample.
The proportion for each level is calculated by dividing the
frequency count for that level by the total sample size.
Frequencies and proportions are typically presented in a
frequency table.
The mode is a measure of the central tendency of a
categorical feature and is simply the most frequent level.
We often also calculate a second mode which is just the
second most common level of a feature.
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Categorical Features

Table: A dataset showing the positions and weekly training expenses


of a school basketball squad.
Training Training
ID Position Expenses ID Position Expenses
1 center 56.75 11 center 550.00
2 guard 1,800.11 12 center 223.89
3 guard 1,341.03 13 center 103.23
4 forward 749.50 14 forward 758.22
5 guard 1,150.00 15 forward 430.79
6 forward 928.30 16 forward 675.11
7 center 250.90 17 guard 1,657.20
8 guard 806.15 18 guard 1,405.18
9 guard 1,209.02 19 guard 760.51
10 forward 405.72 20 forward 985.41
Descriptive Statistics Data Visualization Summary

Descriptive Statistics for Categorical Features

Table: A frequency table for the P OSITION feature from the


professional basketball squad dataset in Table 4 [34] .
Level Count Proportion
guard 8 40%
forward 7 35%
center 5 25%
Descriptive Statistics Data Visualization Summary

Populations & Samples

In statistics it is very important to understand the difference


between a population and a sample.
The term population is used in statistics to represent all
possible measurements or outcomes that are of interest to
us in a particular study or piece of analysis.
The term sample refers to the subset of the population that
is selected for analysis.
The margin of error reported in poll results takes into
account the fact that the result is based on a sample from
a much larger population.
Descriptive Statistics Data Visualization Summary

Populations & Samples

Table: A number of poll results from the run up to the 2012 US


Presidential election.

Margin Sample
Poll Obama Romney Other Date of Error Size
Pew Research 50 47 3 04-Nov ±2.2 2, 709
Gallup 49 50 1 04-Nov ±2.0 2, 700
ABC News/Wash Pos 50 47 3 04-Nov ±2.5 2, 345
CNN/Opinion Research 49 49 2 04-Nov ±3.5 963
Pew Research 50 47 3 03-Nov ±2.2 2, 709
ABC News/Wash Post 49 48 3 03-Nov ±2.5 2, 069
ABC News/Wash Post 49 49 2 30-Oct ±3.0 1, 288

You might also like