Chap 2: Description of Populations
and Samples (STAT 350)
Chapter 2
Review of Chap. 1
Introduction to Statistics in life sciences
– coverage: chap 1 + additional materials
1.1. Statistics and variation in data: (motivation)
One group, two groups
1.2. What does Statistics do? (its role in science)
Types of Evidence
1.3. Random Sampling.
2.1, Variable & measurement scale
Variable
A characteristic of a person or a thing
assigned a number or a category
A variable with categories
A variable with numeric values
Categorical/
Quantitative
Qualitative
Ex: Age of a person
Ex: Gender: Male, Female
No of cars in a mall
Blood type: A,B,O,AB
parking lot
Continuous Discrete
Ordinal Nominal Measured on a continuous scale
Meaningful ranking NO ranking of categories
Ex: Age of students: 20-30years 3
Ex: Survey status: Ex: Gender, Blood type Measured on a discrete scale
Fair, Good,
Ex: No of cars in the mall:
Excellent
0,1,2,3 and so on.
1
A Brain Teaser!
A survey of 50 STAT 350 students took
information on the student’s age, gender, year
in college, no of cars owned (!) and the grade
expected in the course.
Question:
– Identify the variables and the type of variables in
the survey
– Sample size and the observational unit in the
survey
Did you figure out ? or look for answer in the next slide…….
4
Solution
– Observational unit : STAT 350 students
– Sample Size: 50 (ie, n=50)
– Variables
Age: Continuous variable (Ex: 20-27 years)
Gender: Nominal variable (M,F)
Year in College : Ordinal Variable (Freshman,
Sophomore, Junior, Senior)
No of Cars owned: Discrete variable (0,1,2,3…)
Grade expected: Ordinal variable (A,B,C,D,E,F)
Two Main Aspect of Statistics
Descriptive Statistics
– Exploratory Data Analysis
Organize or summarize the data
Graphical display of important patterns and variation
Statistical Inference
– Confirmatory Data Analysis
Draw conclusions about a population from a sample
Assess strength of evidence for competing hypotheses.
Make comparison
Made prediction
2
Objective of Exploratory Data Analysis
Why we need to organize our data?
Condense data present it in a manageable form
Look at detailed distributions (shape, center,
spread) before computing summary measures or
performing statistics
Gain insight into the nature of the data set
Look for errors, abnormal points
Note: Think about this when you do your
experiment in the biology lab! You are collecting
7 data on some variable, then explore it…
Outline of Chap 2
How to organize your data?
2.1 Frequency Plots: (Section 2.2 in Textbook)
– Bar chart (categorical data)
– Dotplot (quantitative data)
– Stem-and-Leaf Diagrams
– Histogram & interpretation
2.2 Summarizing your data: (Section 2.3,2,4,2.5,2.6)
– Shapes of Frequency Distribution
– Measure of Center
– Boxplot , Percentiles
– Measure of Dispersion (how spread is your data)
Bar Chart Example 2.1 and 2.5 in textbook
(for Categorical Variable)
• Identify the variable, variable type, sample
size and the observational unit
• Bar chart
X -axis: Categories of the variable
Y-axis:
• Frequency = count of occurrence
• Relative frequency = Frequency / n
• Percent frequency = Relative frequency x
100
9 • Sample size could be large
3
Dot plot Quantitative variable w/
small sample size
How to draw
1. Draw a line covering the range of the
data
2. Put a dot above the number line for
each observation.
3. If two or more observations on the
same value, stack the dots in plot on
top of each other
Did you identify the variable here?
10 Was the number of surviving piglets – Discrete variable?
Histogram & interpretation
Histogram is like a bar chart, except that it displays a
quantitative variable
– Natural order and scale for the X-axis (variable of interest)
does matter.
– For continuous data, it is common to group data before plot
histogram
Example 2.6 Serum CK (page 14 – 15)
Data:
36 people’s CK concentration in blood
121, 82, 100, 151, ………123, 70, 48, 95, 42
11
Grouped frequency table for
Serum CK data (Page: 32)
How to get the frequency table
• Step 1, Obtain the range for the data. i.e.
(Largest – smallest) = 203 – 25 = 178
• Step 2, Decide on the number of classes,
usually between 5 to 20 depending on the
sample size. Here, they have used 10
classes.
• Step 3,Obtain Class width = Range / No of Things to remember:
classes. Here, 178 – 10 = 17.8, rounded to • Classes should be non-overlapping and
nearest 10’s, which is 20. of equal width
• Step 4, Write the frequency for each class. • Make sure the classes include the
minimum (smallest) and the maximum
12 (Largest) observations.
4
The histogram for serum CK data
How to draw the histogram
X-axis: Quantitative variable (Serum CK)
Divide the axis in to intervals corresponding to the classes
Y-axis: Frequency
Continuously, draw rectangles, whose heights (areas) correspond to
frequency of the class. The rectangles should be of same width and should touch
each other
13
Two ways to view Histogram
1, Tops of the bars sketch out the
shape of the distribution
Density Curve
– Smooth curve approximating
the histogram, which is
based on relative frequency.
2, The area of each bar is
proportional to its frequency.
Similarly, the area of several
bars express the number of
observations in those classes
e.g. shaded area of these two bar =
42% of total area, then, 42% of
total observations are in the
range of 60 – 100.
(Total areas under the density curve
= 1)
14
Problem:
Construct a frequency Histogram
The following sample data set lists the weight (in
pounds) of 30 people. Construct a frequency
distribution that has seven classes.
90 130 400 200 350 70 325 250 150 250
275 270 150 130 59 200 160 450 300 130
220 100 200 400 200 250 95 180 170 150
15
5
Solution Lower Upper
Class limit limit
width = 60 59 118
Step 1: Find the range (Max value – Min value): 119 178
450 – 59 = 391
Step 2: No of classes is given: 7 179 238
Step 3: Find the Class width = Range/No of
classes 239 298
391/7 = 55.86 Round up to 60
Step 4: Find lower limit and upper limit for
299 358
each class.
Use minimum value as the starting point, add
359 418
class width to it (59 + 60 = 119) 419 479
Find the remaining lower limits.
Upper limit for class 1 will be one less than
lower limit of 2nd class
Find remaining upper limits
Comments:
Here, the number of classes to be used is given to be 7.
If not, you can choose the number of classes as at least square root of n.
16 For example: if n= 38,square root of 38 = 6.16. So you could use at least 6 classes.
Class Frequency, f
59 – 118 5
Solution 119 – 178 8
179 – 238 6
239 – 298 5
Step 4:Once you have the
classes set, count the 299 – 358 3
number of observations
(frequency) in each class 359 – 418 2
419 – 479 1
Comments:
– Check whether the 9
frequency adds up to 8
the sample size 7
– Check whether the 6
Frequency
minimum and maximum 5
value are included in the 4
class 3
– Check whether the 2
classes are non 1
overlapping
0
59-118 119-178 179-238 239-298 299-358 359-418 419-479
Weight (pounds)
Skewed to the right.
17
2.2 Summarize your data
2.2.1 - Shape
18
6
Shape of Density Curve
19
Section 2.2 - Descriptive Statistics
Important concepts + language
20
Sample Mean
The mean of the sample (Sample mean) is
the sum of the observations divided by the
number of observations.
21
7
Example 2.3.3
Finding the sample mean
Weight gain of Lambs
– Data: 11, 13, 19, 2, 10, 1
22
Sample Median
23
Sample Median – cont.
Calculation Procedure:
1. Order the observations in the increasing trend
2. In ordered observations,
If N is odd, the middle value.
Median =
If N is even, the midway between
two middle value
24
8
Compare Mean and Median
25
Visualize Mean vs Median
For symmetric distribution,
mean=median
For left skewed distribution,
mean < median
Left skewed means a long left tail
Same as saying Negatively skewed
26
Visualize Mean vs Median
For right skewed distribution, mean > median
Right skewed means a long right tail
Same as saying positively skewed
27
9
Quartiles
Section 2.4, Page 45
28
How to calculate quartiles
Example 2.4.1, Page 46
Data: Systolic blood pressure of seven men
151,124,132,170,146,124,113
First, Arrange data in ascending order
Odd number of observations, median = middle value
113, 124, 124, 132, 146, 151, 170
Q1 Q3
Q2 = Median
Lower part of the data = 113,124,124: Median of these three
values = First quartile or Q1 = 124
Upper part of the data = 146,151,170: Median of these three
values = Third quartile or Q3 = 151
Inter quartile range (IQR) = 151-124 = 27
29
Example 2.4.2, Page 46.
Notice that number of observations is even
30
10
Box plots
Modified Box plots - popular one to use.
Box plot is a visual representation of five-number
summary (highlights important data features)
– Minimum value
– First quartile Q1
– Median Q2
– Third quartile Q3
– Maximum value
How to draw box plot is put in appendix.
Modified Box Plots visual five number summary
and outliers.
31
Modified Box plot
Is a box plot in which outliers are graphed as
separate points
Question is how to know whether an
observation is an outlier or not?
Outlier Calculation
– Calculate Lower and Upper fence
Lower Fence = Q1 – 1.5 * IQR
Upper fence = Q3 + 1.5 * IQR
– An outlier is a data point that falls outside of the
fences
32
Drawing a Box plot Example 2.4.5
Data: Radish growth in Light of 14 radish seedlings
3 5 5 7 7 8 9 10 10 10 10 14 20 21
Five (9 +10)/2
number Q1 Q3 =10 Max =21
Min
Summary =3 =7
Q2 = Median=9.5
3 5 7 9 11 1 1 1 1 2
3 5 7 9 1
Growth of seedlings in Light
Obviously you can see that the maximum value of 21 looks to be an outlier
because of long whisker
33
11
Modified Box plot for Example 2.4.5
Let us do the outlier calculation
– Lower fence = 7-1.5 * 3 = 2.5 (No data point is less than
lower fence)
– Upper fence = 10+1.5 * 3 = 14.5 (17 and 21 are greater
than upper fence)
– We have 17 and 21 as outliers and are marked as dots
on the modified boxplot
– The upper whisker would be drawn from Q3 to the largest
data point that is not an outlier (upper whisker till 14)
34
Modified Boxplot - Step -1
35
36
12
How to draw whiskers
when you have outliers?
37
What can you say about this box plot?
A box plot gives a quick visual summary of the
distribution of the data
The line inside the box = median = tells you
the center of the data
Spread of the total distribution, from range of
data, whiskers, and the box ( Q1 to Q3) tells
you the spread of the middle half of the
distribution
38
Parallel Boxplot
39
13
40
2.6 Measure of Dispersion
Range = maximum – minimum
IQR = Q3 – Q1
Compare “range” vs. IQR:
– Range is very sensitive to extreme values
– IQR is more robust to extreme values
Standard Deviation (SD)
41
Sample SD
Formula & Calculation Steps.
Denominator is (n-1) !
Degree of freedom =
n-1, due to constrain:
Sum(all di) = 0
42
14
Example (SD)
43
Example 2.6.2, Page 60
76-73 =
72-73 =
65-73 =
70-73 =
82-73 =
44
Coefficient of Variation
45
15
Visualizing
Measures of Dispersion
Range: min max, cover 100% of data
– Intuitive, but poor for extreme tails of dist
IQR: spread of middle 50% of the data
– In histogram, Q1, Q2, Q3 divide the area under the
bars into 4 equal parts.
– Intuitive, resistant to extreme tails
– More challenging to estimate in statistics
SD:
– Less intuitive, may subject to extreme tails of the
data
– Nice statistical property
46
Visualizing SD
99.7% within 3 standard deviations
95% within 2 standard deviations
68% within 1
standard deviation
34 34
% %
2.35% 2.35%
13.5% 13.5%
47 x 3s x 2s x s x xs x 2s x 3s
Visualize SD - cont.
In another words,
Note (mean, SD) is the foundation of Classical
Statistics.
48
16
Univariate summaries
Section 2.5
So far, you learnt about univariate summary
measures..
What is univariate summary?
– Graphical or numeric summary of a single variable
Univariate summary of numeric data
– Histogram, boxplot, sample mean and median
Univariate summary of categorical data
– Bar chart, frequency and relative frequency
49
Bivariate graphical summaries
To understand relationship between two
categorical variables
Stacked bar chart
Bivariate frequency table
50
Numeric – categorical relationship
Side by side Box plots
• Notice categorical variables
on the x-axis
• Continuous variable on the
•y-axis
51
17
Numeric-numeric relationship
Scatter plot of two continuous variables,
Liver Selenium and Tooth Selenium
52
Statistical Inference
Section 2.8
So far, you saw the use of Descriptive
statistics to understand the data, which is
nothing but a random sample from a large
population
The process of drawing conclusions about a
population, based on observations in a sample
from that population, is called Statistical
Inference
53
Describing a Population
Characteristics of biological populations are not
known exactly
Why? : Because we always collect only
samples
Sample characteristic (Statistic) is
considered an estimate of the corresponding
population characteristic (Parameter)
Very Important to understand
what is a parameter and a statistic
54
18
Some Parameters and Statistic
Comments:
• Proportion is a measure used in the case of categorical
variables see Example 2.8.4, Page 76
• P-hat (sample proportion) is the estimate of the
population proportion
55
Practice Problem!
Covering all concepts from Chapter 2
56
57
19
58
Take Home for Chapter 2
Variables and Variable types
– Categorical (Nominal,Ordinal)
– Quantitative –Numeric ( Discrete, Continuous)
How to organize your data and display it?
– Bar chart (Categorical data)
– Dot plot, Frequency table and Histogram (Quantitative
data)
– What can we learn from a histogram
Skewed (right, left)
Symmetric
Modality (Unimodal, bimodal)
How to summarize your data? Descriptive
Statistics
59
Take Home for Chapter 2 – cont.
Descriptive Statistics
Measure of Center
– Parameter vs. statistics, sample mean, sample median, and the
visualization of their relationships.
Modified Boxplot , Quartiles
– Quartiles (Q1, Q2, Q3), five-number summary, lower fence,
upper fence, modified boxplot, parallel boxplot
Measure of Dispersion (how spread is your data)
– Range, IQR, sample SD/sample variance, coefficient of variation
– visualizing those measures of dispersion (range, IQR, SD) and
compare them.
– Empirical Rule (68-95-99.7 rule)
Statistical Inference
60 – What is a Population, Sample, Parameter, Statistic
20
Homework -1
Problems from Chapter 2
2.1.1, 2.1.4
2.2.7, Also calculate, mean, median, quartiles,
IQR, standard deviation and construct a box-
plot and provide comments on the distribution
of the data
2.4.6, 2.S.3, 2.S.9 (Page 80)
61
Appendix
62
How to draw a boxplot
1. Find the five-number summary of the data set.
2. Construct a horizontal scale that spans the
range of the data.
3. Plot the five numbers above the horizontal scale.
4. Draw a box above the horizontal scale from Q1 to
Q3 and draw a vertical line in the box at Q2.
5. Draw whiskers from the box to the minimum and
maximum entries.
Box
Whisker Whisker
Minimum Maximu
entry Q1 Median, Q3 m entry
Q2 63
21
Drawing a Box plot
Example 2.4.5
Data: Radish growth in Light of 14 radish seedlings
3 5 5 7 7 8 9 10 10 10 10 14 20 21
(9 +10)/2
Five
number Q1 Q3 =10 Max =21
Min
Summary =3 =7
Q2 = Median=9.5
3 5 7 9 11 1 1 1 1 2
3 5 7 9 1
Growth of seedlings in Light
Obviously you can see that the maximum value of 21 looks to be an outlier
because of long whisker
64
What can you say about this box plot?
3 5 7 9 11 1 1 1 1 2
3 5 7 9 1
Growth of seedlings in Light
A box plot gives a quick visual summary of the
distribution of the data
The line inside the box = median = tells you the
center of the data
Spread of the total distribution, from min to max,
the box ( Q1 to Q3) tells you the spread of the
middle half of the distribution
Long upper whiskers tell you that the data is right
skewed (long right tails)
65
Stem-and-Leaf Diagrams
Simple ungrouped frequency distribution is limited
(basically, just another way to display raw data)
Grouping the data to show grouped frequency
distribution is necessary, especially for continuous
data.
Steps:
1. Select one or more leading digit for the stem values.
The tail digits become the leaves.
2. List possible stem values in a vertical column,
3. record the leaf for every observation beside the
corresponding stem value in order (small to large)
4. Indicate the units for stems and leaves some place in
the display.
66
22
67
Back-to-back Stem and Leaf plot
- for two group comparison (Eg. 2.9, book)
68
23