Biology 1413-151-152-451
Lecture # 03B,
Statistics & Relative Risk, Boxes & Whiskers
continued
••••• •
•••
••••••••
•••••••
••
•••
•
•
•••
•
•
••
•
•
•
•
Data Points plotted by frequency
with that bacterial count
Number of participants
10
Bell
Curve
5
800
900
600
500
1200
1100
1500
1000
Number of bacterial colonies
Example of a Box-and-whiskers graph
(comparable with three Bell Curves)
1500
1400 frequency
1300
1200
1100
1000
900
800
700
600
500
400 Control
300 Group or
200 Baseline
100 Test Group #2
Test Group #1
0
n = 50
n = 25 n = 25
Example of a Box-and-whiskers graph
This is what’s called a Bell Curve,
1000
or a Normal Curve, (here it’s green
with a red top) and shows how the
distribution of data points are
700 clustered by frequency– how
similar each data point is to the
500 arithmetic mean or the “average”
And 95% of the data points fall
inside the green area of the Curve.
95% of results are in this grey-box or
green-colored bell curve
Each Grey Box is like a green-colored “bell curve”
turned on it’s side, and 95% of the “measurements” in Outliers, or Outliers, or
sample falling inside the box. Only the “outliers” or just “due to just “due to
chance” chance”
“just due to chance data” fall at the Edges of the Curve
(like the Whiskers on the Box-and-Whiskers graph)
The 95% Confidence Interval
As long as the mean (average) of one sample
does not overlap with the confidence interval
for the other sample (or vise versa),
then the average difference between the
two means (two averages) is
statistically significant.
Box-and-whiskers Graphs
Whiskers show range of values in the data from this
sample or group
Cross-hatch shows average (mean) for the data
from sample or group, not always “even” or “normal”
Shaded box encloses 95% confidence interval
The whole sample (or all your data) has a 95%
likelihood of being within the box on the graph.
Bell Curve with
Bell Curve with Normal Bell Curve with
Skewed Distribution Skewed
Distribution of Data Points Distribution
of Data Points of Data Points
Graphs of Data Point Distributions
Fall into “Bell Shapes”, but not all “normal” or
even. Can be asymmetrical.
The 95% Confidence Interval
If the mean (average) of one sample
does overlap with the confidence interval
for the other sample (or vise versa),
then the average difference between
the two means (two averages) is
not statistically significant.
See next slide for the soap example.
Sample Box-and-whiskers graph
1
1000 ……………………………...…
b 3
700 ……………………………….
2
500 ………………
Groups 1 and 2, and 1 and 3 are significantly different (statistically), whereas notice
Groups 2 and 3 may “show a trend” but are not significantly different (statistically). See
how the average value (mean) of Group 2 overlaps with Confidence interval (grey box)
for Group 3 (see arrow “a”) and vise versa (see arrow “b”).
The 95% Confidence Interval
If the overlap between mean (average) of
one sample with the confidence interval of
the other sample is small,
then there is a suggestion of a difference, and would
need to investigate further to find out if really
a meaningful finding.
So “washing with any soap was better than before
washing”, but whether subjects used anti-bacterial
soap versus regular soap did not make a statistically
significant difference in number of bacterial colonies–
only a suggestion of a difference.
The Process of Science:
Understanding Statistics
Statistical significance versus biological
significance or practical significance or social
significance….
Not always the same level of meaning
Secondary news sources often do not distinguish
between statistical significance and more casual
use of the word significance (“meaningful
or important” or “draws audiences and makes $$”)
Relative Risk versus Absolute Risk
Risk Assessment
Absolute risk of a disease is your
risk of developing the disease over
a time period.
We all have absolute risks of developing various diseases
such as heart disease, cancer, stroke, etc.
The same absolute risk can be expressed in different
ways:
example: Say you have a 1 in 10 risk of developing
a certain disease in your lifetime.
This can also be said to be a 10% risk (use percentages),
or a 0.1 risk (using decimals)
Risk Assessment
Relative risk is used to compare
the risk in two different groups of
people. For example, the groups
could be smokers and non-smokers.
All sorts of groups are compared to others in medical
research to see if belonging to a group increases or
decreases your risk of developing certain diseases.
For example, research has shown that smokers have a
higher risk of many different diseases compared to
(or “relative to”) non-smokers.
Risk Assessment
Absolute Risk versus Relative Risk
Say the absolute risk of a non-smoker getting breast
cancer over a lifetime is 13% (1 in 8).
Then, if smokers have an increased relative risk of
24% compared to non-smokers:
that 24% refers to the original 13%, so take 24% of 13%.
If you smoke, then the absolute risk increase of getting
breast cancer over your lifetime jumps: calculate 24% of
13 = 3.1. Now, add that 3.1 to original 13 = 16.1%.
You’ve raised your absolute risk from 1/8 (for non-smokers)
to a risk of 1/6 (for smokers),
but NOT twenty-four times more, nor 24% more
absolute risk than in non-smokers.
Risk Assessment
Absolute Risk versus Relative Risk
When you see a headline in the media like “Crossword
Puzzles decrease your odds of getting dementia by 25%,” that
statement is referring to relative risk.
That reduction in risk is usually smaller than you imagine it is.
Say the absolute risk of getting dementia is 20 percent in
people who don’t do crossword puzzles.
The 25% reduction applies to the original 20%.
Calculate 25% of 20%, and the reduction is equal to 5%, so
Absolute risk in non-crossword workers is 20% (1 in 5)
Absolute risk in crossword workers is 15% (1 in 7)
Risk Assessment
Relative Risk cannot be understood without
knowing the Absolute Risk.
The Process of Science:
Understanding Statistics
Statistical significance versus biological
significance or practical significance or social
significance….
Not always the same level of meaning
Secondary news sources often do not distinguish
between statistical significance and more casual
use of the word significance (“meaningful
or important”)
Is it biologically significant? Does it matter in life, for how many?
Is the cost worth the benefit? Does it help only a few?
Should it be available to everyone? Who pays the cost?