Problem Solving, I
IMAN KAMAL
PROFESSOR – COMMUNITY MEDICINE
DTQM- AMERICAN UNIVERSITY IN CAIRO
DGSHH- CLAUDE BERNARD - FRANCE
- *This distribution is skewed to the (left – right “positive”)
- The mean is (less than – greater) than the median
- This distribution is skewed to the (left – right)
- The median is (less than – greater) than the
mean
- In the following standard normal curve, the mean= ...SD=...
- The mean, mode, and median are (equal- not equal)
*Measures of Central Tendency
Suppose these are the scores of biostatistics quiz for five students:
2,3,3,4&5
Calculate:
1- mean=……
2- median=…..
3- mode=……
This data seems to be symmetric or skewed and why?
Measures of Central Tendency
Suppose these are the scores of biostatistics quiz for five students:
2,3,3,4&5
Calculate:
1- mean=3.4……
2- median=3…..(odd number: n+1/2) (even number of observations: n/2 & n/2 +1) Position
3- mode=…3…
This data seems to be symmetric or skewed and why?
Symmetric because the mean is nearly equals to the median
Measures of Central Tendency
3, 3,5, 9, 11
Calculate:
1- mean=……
2- median=…..
3- mode=……
This data seems to be symmetric or skewed “right or left” and why?
Measures of Central Tendency
3, 3,5, 9, 11
Calculate:
1- mean=…6.2…
2- median=…5..
3- mode=……3
This data seems to be symmetric or skewed “right or left” and why?
Positively “right” skewed data
Mean> median
Mean is the only measure of central tendency sensitive to outlier
Measures of Dispersion
5, 3,9, 3, 11
Calculate:
1- Range=….
2- variance= 13.2
3- standard deviation=…..
*Measures of Dispersion
5, 3,9, 3, 11
Calculate:
1- Range=6….
2- variance= 13.2 (sum of squares of mean difference / n-1 “degrees of freedom” )
3- standard deviation=…√13.2= 3.63..
SD = 9 …….Variance =81
Variance = 100………..SD=10
*Box Plot
3,2,3,5,7,4,4,6,8,5,6
Q2
Q1
Q3
IQR
Box Plot
2,3,3, 4,4,5, 5, 6, 6 ,7, 8 “N=11”
Q2 “Median n+1/2” 11+1/2 = 12/2 = 6th = 5
We have 5 values to the right of this median and 5 values to the left
Q1 “5+1/2 = 3rd = 3
Q3 “5+1/2 = 3rd = 6
IQR Q3-Q1 = 6-3 = 3
Empirical Rule
Within one standard deviation (μ±SD)=……% of data
Within two standard deviations (μ±2SD)=……% of data
Within three standard deviations (μ±3SD)=……% of data
Empirical Rule
Within one standard deviation (μ±SD)=…68…% of data
Within two standard deviations (μ±2SD)=…95…% of data
Within three standard deviations (μ±3SD)=…99.9…% of data
*Empirical Rule
Example if mean = 140 SD= 10
Within one standard deviation (μ±SD) (130-150)=…68…% of data
Within two standard deviations (μ±2SD) (120 – 160)=…95…% of data
Within three standard deviations (μ±3SD) 110 – 170) =…99.9…% of data
Types of data “Numerical – Nominal – Ordinal”
◦List all the languages you can speak:………
◦How many friends do you have on Facebook? (0-100, 101-
200, 201-300…):…………
◦Please type in the exact number of friends you have on
Facebook. ………….
◦What is your height in inches? …………
◦How many siblings do you have?............
Types of data “Numerical – Nominal – Ordinal”
◦List all the languages you can speak:Nominal………
◦How many friends do you have on Facebook? (0-100, 101-
200, 201-300…):Ordinal…………
◦Please type in the exact number of friends you have on
Facebook. …Numerical “discrete” ……….
◦What is your height in inches?
…Numerical…”continuous”……
◦How many siblings do you have?.....Numerical
“discrete”.......
*Types of data “Numerical – Nominal – Ordinal”
◦Chlorine concentration in water:………
◦Marital status “Single- Married”:…………
◦Cancer Stages “I, II, III”. ………….
◦Body mass Index…………
◦Blood group “A- B – AB – O”............
◦Height in centimeter ……..
*Types of data “Numerical – Nominal – Ordinal”
◦Chlorine concentration in water:…Numerical
“continuous”……
◦Marital status “Single- Married”:Nominal…………
◦Cancer Stages “I, II, III”. ……Ordinal…….
◦Body mass Index…Numerical “continuous”………
◦Blood group “A- B – AB – O”...Nominal.........
◦Height in centimeter Numerical “continuous”.
GRAPH
- In order to graph numerical data, we have…………………..
- In order to graph nominal data, we
have………………….&………………
- In order to graph ordinal data, we have……………………
*GRAPH
- In order to graph numerical data, we
have…Histogram………………..
- In order to graph nominal data, we have Bar graph &Pie Chart”%”
- In order to graph ordinal data, we have……Bar graph
Confidence Interval
Its width determine the variability of the data
Example
1- CI (140-200) 60
2- CI (140-160) 20
Which CI represent less variability?
*Confidence Interval
1- CI (140-200) 60
2- CI (140-160) 20
2
The narrower the CI, the less variability in the data & large sample size
*Confidence Interval
If the confidence Intervals overlapped = no significant difference
between the group
Example
1- CI (140-200)
2- CI (140-160)
3- CI (130 – 150)
*Confidence Interval
If the confidence Intervals not overlapped = significant difference
between the group
Example
1- CI (140-160)
2- CI (165-180)
3- CI (185 – 200)
*Comparing Two means
- If we are to compare mean hemoglobin of male and female: the test
statistic here will be ………………
- If we compare mean total cholesterol between nurses and physicians
the test statistic here will be ………………
Comparing Two means
- If we are to compare mean hemoglobin of male and female: the test
statistic here will be …Independent t test……………
-If we compare mean total cholesterol between nurses and physicians
the test statistic here will be … Independent t test ……………
*Comparing Two means
- If we are to compare mean hemoglobin of group of children at
certain time, with mean hemoglobin of the same children after 6
months: the test statistic here will be ………………
- If we are planning to measure mean blood cholesterol level of
participants in May 2021 then we will give the same participants
special diet for eight months and then measure their mean blood
cholesterol again the test statistics specific here to compare the two
means is the ……………….
*Comparing Two means
- If we are to compare mean hemoglobin of group of children at
certain time, with mean hemoglobin of the same children after 6
months: the test statistic here will be …Paired t test……………
- If we are planning to measure mean blood cholesterol level of
participants in May 2021 then we will give the same participants
special diet for eight months and then measure their mean blood
cholesterol again the test statistics specific here to compare the two
means is the …… Paired t test ………….
Comparing Two Means
Serum iron level of two samples of children “healthy &
cystic fibrosis”:
Is this observed difference in sample means 18.9 & 11.9 – is
the result of chance variation or should we conclude that the
discrepancy is due to true difference in population means?
Group I “Healthy” Group II “Cystic F”
n n1=9 n2=13
x‾ x1‾=18.9umol/l x2‾=11.9umol/l
S s1=5.9 umol/l s2=6.3 umol/l
30
Testing Hypothesis
1- Set H0: μ1=μ2
2- Set Ha: μ1≠μ2
3- Set α=0.05
4- t test=2.63
5- p value<0.05
6- Decision:…………
7- Conclusion:………………….
31
Testing Hypothesis
1- Set H0: μ1=μ2
2- Set Ha: μ1≠μ2
3- Set α=0.05
4- t test=2.63
5- p value<0.05
6- Decision:…reject the null hypothesis………
7- Conclusion:…there is evidence for association
32
Comparing Two Means
Effect of antihypertensive drug on persons over 60 suffer
from isolated systolic blood pressure “sys>160 and dias <90”
GI active drug µ1……..GII placebo µ2 “ for one year”
33
GI “Active Drug” GII “Placebo”
N n1=2308 n2=2293
X x1=142.5 mm Hg x2=156.5 mm Hg
S s1=15.7 mm Hg s2=17.3 mm Hg
34
Testing Hypothesis
1- Set H0: μ1=μ2
2- Set Ha: μ1≠μ2
3- Set α=0.05
4- t test=-28.74
5- p value<0.05
6- Decision:…………
7- Conclusion:………………….
35
Testing Hypothesis
1- Set H0: μ1=μ2
2- Set Ha: μ1≠μ2
3- Set α=0.05
4- t test=-28.74
5- p value<0.05
6- Decision:…reject the null hypothesis………
7- Conclusion:……there is evidence for association
36
If Testing Hypothesis gave us the following
1- Set H0: μ1=μ2
2- Set Ha: μ1≠μ2
3- Set α=0.05
4- t test=-0.8
5- p value>0.05
6- Decision:…fail to reject the null hypothesis………
7- Conclusion:……there is no evidence for association
37
*Comparing More Than Two Means
Exercise stresses the bones, and this causes them to get
stronger.
There were three groups of rats: GI control with no jumping,
GII a low-jump condition (the jump height was 30
centimeters), and GIII a high-jump condition (60 centimeters).
After 8 weeks of 10 jumps per day, 5 days per week, the bone
density of the rats (expressed in mg/cm3 ) was measured.
- The test statistic here for comparing more than two means is
………………
38
Comparing More Than Two Means
Exercise stresses the bones, and this causes them to get
stronger.
There were three groups of rats: GI control with no jumping,
GII a low-jump condition (the jump height was 30
centimeters), and GIII a high-jump condition (60 centimeters).
After 8 weeks of 10 jumps per day, 5 days per week, the bone
density of the rats (expressed in mg/cm3 ) was measured.
- The test statistic here for comparing more than two means is
…One way ANOVA……………
39
Testing Hypothesis
1- Set H0: μ1=μ2 = μ3
2- Set Ha: at least one mean is different
3- Set α=0.05
4- …F… test=7.98
5- p value=0.0019
6- Decision:…………
7- Conclusion:………………….
40
Testing Hypothesis
1- Set H0: μ1=μ2 = μ3
2- Set Ha: at least one mean is different
3- Set α=0.05
4- …F… test=7.98
5- p value=0.0019 “significant association” / p>0.05 “0.06 –
0.3 – 0.5 …”…….non significant association
6- Decision:…reject the null hypothesis………
7- Conclusion:…there is evidence for association
41
*Testing the association between two
categorical variables “two proportions”
EFM EFM- Total
CS 358 229 587
CS- 2492 2745 5237
Total 2850 2974 5824
Testing Hypothesis
1- Set H0: There is no association between electro-fetal
monitoring & cesarean section
2- Set Ha: There is association between electro-fetal monitoring
& cesarean section
3- Set α=0.05
4- …… test=67
5- p value<0.05
6- Decision:…………
7- Conclusion:………………….
43
Testing Hypothesis
1- Set H0: There is no association between electro-fetal
monitoring & cesarean section
2- Set Ha: There is association between electro-fetal monitoring
& cesarean section
3- Set α=0.05
4- …Chi square test… test=67
5- p value<0.05 significant association “0.01 – 0.02 – 0.03 – 0.04”
6- Decision:…reject the null hypothesis………
7- Conclusion:…there is association……………….
44
*Chi square test examples
Smoking and coronary heart disease
Disease and immunization
Radiation and breast cancer
Correlation between Systolic Blood Pressure &
Body Mass Index “Name this figure……………”
300
250
200
150
100
10 20 30 40 50 60
Body mass index, exam 1
Systolic blood pressure (mmHg), exam 1 Fitted values
There is (negative / positive) correlation
46
Correlation between Systolic Blood Pressure & Body Mass
Index “Name this figure……scatter plot………” it
illustrates the relation between two numerical variables
300
250
200
150
100
10 20 30 40 50 60
Body mass index, exam 1
Systolic blood pressure (mmHg), exam 1 Fitted values
There is (negative / positive) correlation
47
Testing the association between two
numerical variables
r=0.87
This means (strong – weak) and (positive – negative) linear
correlation between systolic blood pressure and body mass index
r: correlation …………..
Testing the association between two
numerical variables
r=0.87 p =0.01
This means (strong – weak) and (positive – negative) linear
significant correlation between systolic blood pressure and body
mass index
r: correlation ………coefficient…..
*Testing the association between two
numerical variables
r= - 0.87 p=0.01
This means (strong – weak) and (positive – negative) linear
significant correlation between serum hemoglobin level and lead
level
Example:
SBP & BMI
Height & weight
TC & SBP
*Simple Linear Regression
sysbp1 Coef. Std. Err. t P>|t| [95% Conf. Interval]
bmi1 1.78852 .0775178 23.07 0.000 1.636546 1.940493
_cons 86.66684 2.028606 42.72 0.000 82.68975 90.64392
Regression coefficient=1.78……….
It means for each unit increase in BMI the systolic blood pressure will be
Increased by 1.78 mmHg / YES we can use bmi as predictor for systolic
Blood pressure because the p value is less than 0.05……………………….
51
*Simple Linear Regression
. regress sysbp1 age1 cigpday1
Source SS df MS Number of obs = 4402
F( 2, 4399) = 417.35
Model 351403.295 2 175701.648 Prob > F = 0.0000
Residual 1851948.47 4399 420.993059 R-squared = 0.1595
Adj R-squared = 0.1591
Total 2203351.76 4401 500.64798 Root MSE = 20.518
sysbp1 Coef. Std. Err. t P>|t| [95% Conf. Interval]
age1 1.016436 .0362736 28.02 0.000 .9453215 1.087551
cigpday1 -.0407954 .0264093 -1.54 0.122 -.0925709 .0109802
_cons 82.50898 1.896165 43.51 0.000 78.79154 86.22642
Regression coefficient=-0.04……….
It means for each pack increase in smoking the systolic blood pressure will be
decreased by 0.04 mmHg / NO we can not use cigarette packs as predictor for systolic
Blood pressure because the p value “0.122” is greater than 0.05……………………….
Sampling
Two main types of sampling techniques:
A- Non-probability Sampling:
1- Convenient sample
2- snow balling sample
3- purposive sample
B- Probability Sampling:
1- Simple random sample
2- Systematic random sample
3- stratified sample
4- Cluster sample
Sampling
- If we have a list of total 500 students and we withdrew a sample of
20 students, by random number generator we selected the sample: the
type of this sample is (non-probability – probability) (simple
random sample – systematic random sample – stratified random
sample – cluster sample)
Sampling
- If we have a list of total 500 students and we withdrew a sample of
20 students, by random number generator we selected the sample: the
type of this sample is (non-probability – probability) (simple
random sample – systematic random sample – stratified random
sample – cluster sample)
*Sampling
- If we aimed to test the association between gender and stress we
sampled 300 females and 700 males, and we knew that the ration of
distribution of female to male was 30:70. the type of this sample is
(non-probability – probability) (simple random sample –
systematic random sample – stratified random sample – cluster
sample)
Sampling
- If we aimed to test the association between gender and stress we
sampled 300 females and 700 males, and we knew that the ration of
distribution of female to male was 30:70. the type of this sample is
(non-probability – probability) (simple random sample –
systematic random sample – stratified random sample – cluster
sample)