0% found this document useful (0 votes)
4 views23 pages

Chapter 2 Slides

Chapter 2 focuses on the description of populations and samples in statistics, covering key concepts such as variables, measurement scales, and types of data. It introduces descriptive statistics and statistical inference, emphasizing the importance of exploratory data analysis and various methods for organizing and summarizing data. The chapter also discusses graphical representations like histograms and box plots, alongside measures of central tendency and dispersion.

Uploaded by

margarn1
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views23 pages

Chapter 2 Slides

Chapter 2 focuses on the description of populations and samples in statistics, covering key concepts such as variables, measurement scales, and types of data. It introduces descriptive statistics and statistical inference, emphasizing the importance of exploratory data analysis and various methods for organizing and summarizing data. The chapter also discusses graphical representations like histograms and box plots, alongside measures of central tendency and dispersion.

Uploaded by

margarn1
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chap 2: Description of Populations

and Samples (STAT 350)

Chapter 2

Review of Chap. 1

 Introduction to Statistics in life sciences


– coverage: chap 1 + additional materials

1.1. Statistics and variation in data: (motivation)


One group, two groups

1.2. What does Statistics do? (its role in science)


Types of Evidence

1.3. Random Sampling.

2.1, Variable & measurement scale


Variable
A characteristic of a person or a thing
assigned a number or a category

A variable with categories


A variable with numeric values
Categorical/
Quantitative
Qualitative
Ex: Age of a person
Ex: Gender: Male, Female
No of cars in a mall
Blood type: A,B,O,AB
parking lot

Continuous Discrete
Ordinal Nominal Measured on a continuous scale
Meaningful ranking NO ranking of categories
Ex: Age of students: 20-30years 3
Ex: Survey status: Ex: Gender, Blood type Measured on a discrete scale
Fair, Good,
Ex: No of cars in the mall:
Excellent
0,1,2,3 and so on.

1
A Brain Teaser!

 A survey of 50 STAT 350 students took


information on the student’s age, gender, year
in college, no of cars owned (!) and the grade
expected in the course.
 Question:
– Identify the variables and the type of variables in
the survey
– Sample size and the observational unit in the
survey

Did you figure out ? or look for answer in the next slide…….
4

 Solution
– Observational unit : STAT 350 students
– Sample Size: 50 (ie, n=50)
– Variables
 Age: Continuous variable (Ex: 20-27 years)
 Gender: Nominal variable (M,F)
 Year in College : Ordinal Variable (Freshman,
Sophomore, Junior, Senior)
 No of Cars owned: Discrete variable (0,1,2,3…)
 Grade expected: Ordinal variable (A,B,C,D,E,F)

Two Main Aspect of Statistics

 Descriptive Statistics
– Exploratory Data Analysis
 Organize or summarize the data
 Graphical display of important patterns and variation

 Statistical Inference
– Confirmatory Data Analysis
 Draw conclusions about a population from a sample
 Assess strength of evidence for competing hypotheses.
 Make comparison
 Made prediction

2
Objective of Exploratory Data Analysis

Why we need to organize our data?


 Condense data present it in a manageable form
 Look at detailed distributions (shape, center,
spread) before computing summary measures or
performing statistics
 Gain insight into the nature of the data set
 Look for errors, abnormal points
 Note: Think about this when you do your
experiment in the biology lab! You are collecting
7 data on some variable, then explore it…

Outline of Chap 2
How to organize your data?

2.1 Frequency Plots: (Section 2.2 in Textbook)


– Bar chart (categorical data)
– Dotplot (quantitative data)
– Stem-and-Leaf Diagrams
– Histogram & interpretation
2.2 Summarizing your data: (Section 2.3,2,4,2.5,2.6)
– Shapes of Frequency Distribution
– Measure of Center
– Boxplot , Percentiles
– Measure of Dispersion (how spread is your data)

Bar Chart Example 2.1 and 2.5 in textbook


(for Categorical Variable)

• Identify the variable, variable type, sample


size and the observational unit
• Bar chart
X -axis: Categories of the variable
Y-axis:
• Frequency = count of occurrence
• Relative frequency = Frequency / n
• Percent frequency = Relative frequency x
100
9 • Sample size could be large

3
Dot plot  Quantitative variable w/
small sample size

How to draw
1. Draw a line covering the range of the
data
2. Put a dot above the number line for
each observation.
3. If two or more observations on the
same value, stack the dots in plot on
top of each other
Did you identify the variable here?
10 Was the number of surviving piglets – Discrete variable?

Histogram & interpretation

 Histogram is like a bar chart, except that it displays a


quantitative variable
– Natural order and scale for the X-axis (variable of interest)
does matter.
– For continuous data, it is common to group data before plot
histogram
Example 2.6 Serum CK (page 14 – 15)
 Data:
36 people’s CK concentration in blood
121, 82, 100, 151, ………123, 70, 48, 95, 42

11

Grouped frequency table for


Serum CK data (Page: 32)

How to get the frequency table


• Step 1, Obtain the range for the data. i.e.
(Largest – smallest) = 203 – 25 = 178
• Step 2, Decide on the number of classes,
usually between 5 to 20 depending on the
sample size. Here, they have used 10
classes.
• Step 3,Obtain Class width = Range / No of Things to remember:
classes. Here, 178 – 10 = 17.8, rounded to • Classes should be non-overlapping and
nearest 10’s, which is 20. of equal width
• Step 4, Write the frequency for each class. • Make sure the classes include the
minimum (smallest) and the maximum
12 (Largest) observations.

4
The histogram for serum CK data
 How to draw the histogram

X-axis: Quantitative variable (Serum CK)


 Divide the axis in to intervals corresponding to the classes
Y-axis: Frequency
 Continuously, draw rectangles, whose heights (areas) correspond to
frequency of the class. The rectangles should be of same width and should touch
each other

13

Two ways to view Histogram


1, Tops of the bars sketch out the
shape of the distribution 
 Density Curve
– Smooth curve approximating
the histogram, which is
based on relative frequency.
2, The area of each bar is
proportional to its frequency.
Similarly, the area of several
bars express the number of
observations in those classes
e.g. shaded area of these two bar =
42% of total area, then, 42% of
total observations are in the
range of 60 – 100.
(Total areas under the density curve
= 1)

14

Problem:
Construct a frequency Histogram

The following sample data set lists the weight (in


pounds) of 30 people. Construct a frequency
distribution that has seven classes.
90 130 400 200 350 70 325 250 150 250
275 270 150 130 59 200 160 450 300 130
220 100 200 400 200 250 95 180 170 150

15

5
Solution Lower Upper
Class limit limit
width = 60 59 118
Step 1: Find the range (Max value – Min value): 119 178
450 – 59 = 391
Step 2: No of classes is given: 7 179 238
Step 3: Find the Class width = Range/No of
classes 239 298
391/7 = 55.86 Round up to 60
Step 4: Find lower limit and upper limit for
299 358
each class.
 Use minimum value as the starting point, add
359 418
class width to it (59 + 60 = 119) 419 479
 Find the remaining lower limits.
 Upper limit for class 1 will be one less than
lower limit of 2nd class
 Find remaining upper limits

Comments:
Here, the number of classes to be used is given to be 7.
If not, you can choose the number of classes as at least square root of n.
16 For example: if n= 38,square root of 38 = 6.16. So you could use at least 6 classes.

Class Frequency, f
59 – 118 5
Solution 119 – 178 8
179 – 238 6
239 – 298 5
 Step 4:Once you have the
classes set, count the 299 – 358 3
number of observations
(frequency) in each class 359 – 418 2
419 – 479 1
 Comments:
– Check whether the 9

frequency adds up to 8
the sample size 7

– Check whether the 6


Frequency

minimum and maximum 5


value are included in the 4
class 3
– Check whether the 2
classes are non 1
overlapping
0
59-118 119-178 179-238 239-298 299-358 359-418 419-479
Weight (pounds)
 Skewed to the right.
17

2.2 Summarize your data


2.2.1 - Shape

18

6
Shape of Density Curve

19

Section 2.2 - Descriptive Statistics


Important concepts + language

20

Sample Mean

 The mean of the sample (Sample mean) is


the sum of the observations divided by the
number of observations.

21

7
Example 2.3.3
Finding the sample mean

 Weight gain of Lambs


– Data: 11, 13, 19, 2, 10, 1

22

Sample Median

23

Sample Median – cont.

Calculation Procedure:
1. Order the observations in the increasing trend
2. In ordered observations,

 If N is odd, the middle value.


Median =
 If N is even, the midway between
two middle value

24

8
Compare Mean and Median

25

Visualize Mean vs Median

For symmetric distribution,


mean=median

For left skewed distribution,


mean < median
Left skewed means a long left tail
Same as saying Negatively skewed

26

Visualize Mean vs Median

For right skewed distribution, mean > median


Right skewed means a long right tail
Same as saying positively skewed

27

9
Quartiles
Section 2.4, Page 45

28

How to calculate quartiles


Example 2.4.1, Page 46
Data: Systolic blood pressure of seven men
151,124,132,170,146,124,113
 First, Arrange data in ascending order
 Odd number of observations, median = middle value
113, 124, 124, 132, 146, 151, 170

Q1 Q3
Q2 = Median

 Lower part of the data = 113,124,124: Median of these three


values = First quartile or Q1 = 124
 Upper part of the data = 146,151,170: Median of these three
values = Third quartile or Q3 = 151
 Inter quartile range (IQR) = 151-124 = 27
29

Example 2.4.2, Page 46.


Notice that number of observations is even

30

10
Box plots
 Modified Box plots - popular one to use.
 Box plot is a visual representation of five-number
summary (highlights important data features)
– Minimum value
– First quartile Q1
– Median Q2
– Third quartile Q3
– Maximum value
 How to draw box plot is put in appendix.
 Modified Box Plots visual five number summary
and outliers.
31

Modified Box plot

 Is a box plot in which outliers are graphed as


separate points
 Question is how to know whether an
observation is an outlier or not?
 Outlier Calculation
– Calculate Lower and Upper fence
 Lower Fence = Q1 – 1.5 * IQR
 Upper fence = Q3 + 1.5 * IQR
– An outlier is a data point that falls outside of the
fences

32

Drawing a Box plot Example 2.4.5


Data: Radish growth in Light of 14 radish seedlings

3 5 5 7 7 8 9 10 10 10 10 14 20 21

Five (9 +10)/2
number Q1 Q3 =10 Max =21
Min
Summary =3 =7
Q2 = Median=9.5

3 5 7 9 11 1 1 1 1 2
3 5 7 9 1
Growth of seedlings in Light

Obviously you can see that the maximum value of 21 looks to be an outlier
because of long whisker
33

11
Modified Box plot for Example 2.4.5
 Let us do the outlier calculation
– Lower fence = 7-1.5 * 3 = 2.5 (No data point is less than
lower fence)
– Upper fence = 10+1.5 * 3 = 14.5 (17 and 21 are greater
than upper fence)
– We have 17 and 21 as outliers and are marked as dots
on the modified boxplot
– The upper whisker would be drawn from Q3 to the largest
data point that is not an outlier (upper whisker till 14)

34

Modified Boxplot - Step -1

35

36

12
How to draw whiskers
when you have outliers?

37

What can you say about this box plot?

 A box plot gives a quick visual summary of the


distribution of the data
 The line inside the box = median = tells you
the center of the data
 Spread of the total distribution, from range of
data, whiskers, and the box ( Q1 to Q3) tells
you the spread of the middle half of the
distribution

38

Parallel Boxplot

39

13
40

2.6 Measure of Dispersion

 Range = maximum – minimum


 IQR = Q3 – Q1
 Compare “range” vs. IQR:
– Range is very sensitive to extreme values
– IQR is more robust to extreme values
 Standard Deviation (SD)

41

Sample SD
Formula & Calculation Steps.

Denominator is (n-1) !
Degree of freedom =
n-1, due to constrain:
Sum(all di) = 0

42

14
Example (SD)

43

Example 2.6.2, Page 60

76-73 =
72-73 =
65-73 =
70-73 =
82-73 =

44

Coefficient of Variation

45

15
Visualizing
Measures of Dispersion

 Range: min  max, cover 100% of data


– Intuitive, but poor for extreme tails of dist
 IQR: spread of middle 50% of the data
– In histogram, Q1, Q2, Q3 divide the area under the
bars into 4 equal parts.
– Intuitive, resistant to extreme tails
– More challenging to estimate in statistics
 SD:
– Less intuitive, may subject to extreme tails of the
data
– Nice statistical property
46

Visualizing SD

99.7% within 3 standard deviations

95% within 2 standard deviations


68% within 1
standard deviation

34 34
% %

2.35% 2.35%
13.5% 13.5%

47 x  3s x  2s x s x xs x  2s x  3s

Visualize SD - cont.

 In another words,

Note  (mean, SD) is the foundation of Classical


Statistics.
48

16
Univariate summaries
 Section 2.5

 So far, you learnt about univariate summary


measures..
 What is univariate summary?
– Graphical or numeric summary of a single variable
 Univariate summary of numeric data
– Histogram, boxplot, sample mean and median
 Univariate summary of categorical data
– Bar chart, frequency and relative frequency

49

Bivariate graphical summaries


 To understand relationship between two
categorical variables

Stacked bar chart

Bivariate frequency table

50

Numeric – categorical relationship

Side by side Box plots

• Notice categorical variables


on the x-axis
• Continuous variable on the
•y-axis

51

17
Numeric-numeric relationship

Scatter plot of two continuous variables,


Liver Selenium and Tooth Selenium

52

Statistical Inference
Section 2.8

 So far, you saw the use of Descriptive


statistics to understand the data, which is
nothing but a random sample from a large
population
 The process of drawing conclusions about a
population, based on observations in a sample
from that population, is called Statistical
Inference

53

Describing a Population

 Characteristics of biological populations are not


known exactly
 Why? : Because we always collect only
samples
 Sample characteristic (Statistic) is
considered an estimate of the corresponding
population characteristic (Parameter)

Very Important to understand


 what is a parameter and a statistic
54

18
Some Parameters and Statistic

Comments:
• Proportion is a measure used in the case of categorical
variables  see Example 2.8.4, Page 76
• P-hat (sample proportion) is the estimate of the
population proportion

55

Practice Problem!
Covering all concepts from Chapter 2

56

57

19
58

Take Home for Chapter 2

 Variables and Variable types


– Categorical (Nominal,Ordinal)
– Quantitative –Numeric ( Discrete, Continuous)
 How to organize your data and display it?
– Bar chart (Categorical data)
– Dot plot, Frequency table and Histogram (Quantitative
data)
– What can we learn from a histogram
 Skewed (right, left)
 Symmetric
 Modality (Unimodal, bimodal)
 How to summarize your data? Descriptive
Statistics
59

Take Home for Chapter 2 – cont.


Descriptive Statistics

 Measure of Center
– Parameter vs. statistics, sample mean, sample median, and the
visualization of their relationships.
 Modified Boxplot , Quartiles
– Quartiles (Q1, Q2, Q3), five-number summary, lower fence,
upper fence, modified boxplot, parallel boxplot
 Measure of Dispersion (how spread is your data)
– Range, IQR, sample SD/sample variance, coefficient of variation
– visualizing those measures of dispersion (range, IQR, SD) and
compare them.
– Empirical Rule (68-95-99.7 rule)
 Statistical Inference
60 – What is a Population, Sample, Parameter, Statistic

20
Homework -1
Problems from Chapter 2

 2.1.1, 2.1.4
 2.2.7, Also calculate, mean, median, quartiles,
IQR, standard deviation and construct a box-
plot and provide comments on the distribution
of the data
 2.4.6, 2.S.3, 2.S.9 (Page 80)

61

Appendix

62

How to draw a boxplot


1. Find the five-number summary of the data set.
2. Construct a horizontal scale that spans the
range of the data.
3. Plot the five numbers above the horizontal scale.
4. Draw a box above the horizontal scale from Q1 to
Q3 and draw a vertical line in the box at Q2.
5. Draw whiskers from the box to the minimum and
maximum entries.

Box
Whisker Whisker

Minimum Maximu
entry Q1 Median, Q3 m entry
Q2 63

21
Drawing a Box plot
Example 2.4.5
Data: Radish growth in Light of 14 radish seedlings

3 5 5 7 7 8 9 10 10 10 10 14 20 21

(9 +10)/2
Five
number Q1 Q3 =10 Max =21
Min
Summary =3 =7
Q2 = Median=9.5

3 5 7 9 11 1 1 1 1 2
3 5 7 9 1
Growth of seedlings in Light

Obviously you can see that the maximum value of 21 looks to be an outlier
because of long whisker
64

What can you say about this box plot?

3 5 7 9 11 1 1 1 1 2
3 5 7 9 1
Growth of seedlings in Light

 A box plot gives a quick visual summary of the


distribution of the data
 The line inside the box = median = tells you the
center of the data
 Spread of the total distribution, from min to max,
the box ( Q1 to Q3) tells you the spread of the
middle half of the distribution
 Long upper whiskers tell you that the data is right
skewed (long right tails)
65

Stem-and-Leaf Diagrams

 Simple ungrouped frequency distribution is limited


(basically, just another way to display raw data)
 Grouping the data to show grouped frequency
distribution is necessary, especially for continuous
data.
Steps:
1. Select one or more leading digit for the stem values.
The tail digits become the leaves.
2. List possible stem values in a vertical column,
3. record the leaf for every observation beside the
corresponding stem value in order (small to large)
4. Indicate the units for stems and leaves some place in
the display.
66

22
67

Back-to-back Stem and Leaf plot


- for two group comparison (Eg. 2.9, book)

68

23

You might also like