Descriptive Statistics Data Visualization Summary
Descriptive Statistics
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
The arithmetic mean (or sample mean or just mean) of a
set of n values for a feature a, a1 , a2 . . . an , is denoted by
the symbol a, and is calculated as:
n
1X
a= ai
n
i=1
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149
150 163 145 140 157 151 140 149
Figure: The members of a school basketball squad. The dashed grey
line shows the arithmetic mean of the players’ heights.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149
150 163 145 140 157 151 140 149
Figure: The members of a school basketball squad. The dashed grey
line shows the arithmetic mean of the players’ heights.
1
H EIGHT = × (150 + 163 + 145 + 140 + 157 + 151 + 140 + 149)
8
= 149.375
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
The arithmetic mean is one measure of the central
tendency of a sample (for our purposes a sample is just a
set of values for a feature in an ABT).
Any measure of central tendency is, however, just an
approximation.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
Suppose our basketball squad manage to sign a ringer
measuring in at 229cm
150 163 145 140 157 151 140 149 229
The arithmetic mean for the full group is 158.235cm and no
longer represents the central tendency of the group.
An unusually large or small value like this is referred to as
an outlier - the arithmetic mean is very sensitive to
outliers.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
The median of a set of values can be calculated by
ordering the values from lowest to highest and selecting
the middle value.
If there is an even number of values in the sample then the
median is obtained by calculating the arithmetic mean of
the middle two values.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
140 140 145 149 150 151 157 163 229
Figure: The members of the school basketball squad ordered by
height, the dashed grey line shows the median.
ID 4 7 3 8 1 6 5 2 9
Height 140 140 145 149 150 151 157 163 229
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
We also measure the variation in our data.
In essence, most of statistics, and in turn analytics, is
about describing and understanding variation.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
The simplest measure of variation is the range:
range = max(a) − min(a)
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
What is the range of the heights of the two basketball squads?
ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149
ID 1 2 3 4 5 6 7 8
Height 192 102 145 165 126 154 123 188
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
What is the range of the heights of the two basketball squads?
range = 163 − 140 = 23
range = 192 − 102 = 90
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
The variance of a sample measures the average
difference between each value in a sample and the mean
of that sample.
The variance of the n values of a feature a, a1 , a2 . . . an , is
denoted var (a) and is calculated as:
n
X
(ai − a)2
i=1
var (a) =
n−1
Variance measures average diff between each value and
mean
We square each diff, as some diffs can be positive and other::
can be negative.
Denominator is n-1, so that sample variance is an unbiased
estimator of population variance.
We say that an estimator is unbiased, if sample variance on
an average equals population variance.
If Denominator is n, then we have a biased estimator that on
an average underestimates population variance.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
What is the variance of the heights of the two basketball
squads?
ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149
ID 1 2 3 4 5 6 7 8
Height 192 102 145 165 126 154 123 188
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
(150 − 149.375)2 + (163 − 149.375)2 + . . . + (149 − 149.375)2
var (H EIGHT) =
8−1
= 63.125
(192 − 149.375)2 + (102 − 149.375)2 + . . . + (188 − 149.375)2
var (H EIGHT) =
8−1
= 1, 011.41071
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
The standard deviation, sd, of a sample is calculated by
taking the square root of the variance of the sample:
p
sd(a) = var (a) (1)
v
u n
uX
u
u (ai − a)2
t i=1
= (2)
n−1
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
What is the standard deviation of the heights of the two
basketball squads?
ID 1 2 3 4 5 6 7 8
Height 150 163 145 140 157 151 140 149
ID 1 2 3 4 5 6 7 8
Height 192 102 145 165 126 154 123 188
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
√
sd(H EIGHT) = 63.125
= 7.9451 . . .
p
sd(H EIGHT) = 1, 011.41071
= 31.8026 . . .
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Percentiles are another useful measure of the variation of
i
the values for a feature: a proportion of 100 of the values in
a sample take values equal to or lower than the i th
percentile of that sample.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
To calculate the i th percentile of the n values of a feature a,
a1 , a2 . . . an :
First order the values in ascending order and then multiply
i
n by 100 to determine the index.
If the index is a whole number we take the value at that
position in the ordered list of values as the i th percentile.
If index is not a whole number then we interpolate the
value for the i th percentile as:
i th percentile = (1 − index_f ) × aindex_w + index_f × aindex_w+1
where index_w is the whole part of index, index_f is the
fractional part of index and aindex_w is the value in the
ordered list at position index_w.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
ID 2 7 5 3 6 4 8 1
Height 102 123 126 145 154 165 188 192
102122 126 145 154 165 188 192
What is the 25th percentile of the heights of the basketball
squad?
What is the 80th percentile of the heights of the basketball
squad?
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
To calculate the 25th percentile we first calculate index as
25 th
100 × 8 = 2. So, the 25 percentile is the second value in
the ordered list which is 123.
To calculate the 80th percentile we first calculate index as
80
100 × 8 = 6.4. Because index is not a whole number we
set index_w to the whole part of index, 6, and index_f to
the fractional part, 0.4. Then we can calculate the 80th
percentile as:
(1 − 0.4) × 165 + 0.4 × 188 = 174.2
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
We can use percentiles to describe another measure of
variation know as the inter-quartile range.
The inter-quartile range is calculated as the difference
between the 25th percentile and the 75th percentile.1
1
These percentiles are also known the lower quartile (or 1st quartile) and
upper quartile (or 3rd quartile) hence the name inter-quartile range.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Continuous Features
Example
For the heights of the first basketball team the inter-quartile
range is 151 − 140 = 11, while for the second team it is
165 − 123 = 42.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Categorical Features
For categorical features we are interested primarily in
frequency counts and proportions.
The frequency count of each level of a categorical feature is
calculated by counting the number of times that level
appears in the sample.
The proportion for each level is calculated by dividing the
frequency count for that level by the total sample size.
Frequencies and proportions are typically presented in a
frequency table.
The mode is a measure of the central tendency of a
categorical feature and is simply the most frequent level.
We often also calculate a second mode which is just the
second most common level of a feature.
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Categorical Features
Table: A dataset showing the positions and weekly training expenses
of a school basketball squad.
Training Training
ID Position Expenses ID Position Expenses
1 center 56.75 11 center 550.00
2 guard 1,800.11 12 center 223.89
3 guard 1,341.03 13 center 103.23
4 forward 749.50 14 forward 758.22
5 guard 1,150.00 15 forward 430.79
6 forward 928.30 16 forward 675.11
7 center 250.90 17 guard 1,657.20
8 guard 806.15 18 guard 1,405.18
9 guard 1,209.02 19 guard 760.51
10 forward 405.72 20 forward 985.41
Descriptive Statistics Data Visualization Summary
Descriptive Statistics for Categorical Features
Table: A frequency table for the P OSITION feature from the
professional basketball squad dataset in Table 4 [34] .
Level Count Proportion
guard 8 40%
forward 7 35%
center 5 25%
Descriptive Statistics Data Visualization Summary
Populations & Samples
In statistics it is very important to understand the difference
between a population and a sample.
The term population is used in statistics to represent all
possible measurements or outcomes that are of interest to
us in a particular study or piece of analysis.
The term sample refers to the subset of the population that
is selected for analysis.
The margin of error reported in poll results takes into
account the fact that the result is based on a sample from
a much larger population.
Descriptive Statistics Data Visualization Summary
Populations & Samples
Table: A number of poll results from the run up to the 2012 US
Presidential election.
Margin Sample
Poll Obama Romney Other Date of Error Size
Pew Research 50 47 3 04-Nov ±2.2 2, 709
Gallup 49 50 1 04-Nov ±2.0 2, 700
ABC News/Wash Pos 50 47 3 04-Nov ±2.5 2, 345
CNN/Opinion Research 49 49 2 04-Nov ±3.5 963
Pew Research 50 47 3 03-Nov ±2.2 2, 709
ABC News/Wash Post 49 48 3 03-Nov ±2.5 2, 069
ABC News/Wash Post 49 49 2 30-Oct ±3.0 1, 288