Which of the following is essential for converting data into information?
Context
Epidemiology
Research
Mathematics
Computers
Which of the following is the INDEPENDENT variable in the following research question? What is the rate
of lung cancer in 40-year-old males who smoke versus 40-year-old non-smokers, over a ten-year period?
Smoking
Lung cancer
40-year-olds
Males
Skin cancer
-----------
Country of origin is an example of which type of variable?
Ratio
Interval
Ordinal
Continuous
Categorical
Which of the following is a continuous variable?
Hair color
Blood pressure
Age group
Vaccination status
Ethnicity
-----------
What is meant by the term “dichotomization” in epidemiology?
Sampling a small group of the population to represent a larger sample of the population
"All or none" thinking in dialectical behavioral therapy
Conversion of a continuous variable into two groups
The categorization of variables into dependent and independent
Dividing a continuous variable into multiple groups
Which of the following is a dichotomous variable?
Household income
Number of siblings
Education level
Results of a screening test
Incidence of a disease
A study is developed to examine associations between weight and nutritional status in a sample of
American children. Which of the following is the “reference population” for this study?
Vaccination status of children in the USA
All children in the USA
Nutritional status of children in the USA
Children with obesity in the USA
The children whose data is included in the study
A study is developed to examine the effect of social settings on the acceptance of HIV testing. Which of
the following represents a “sampling frame” for this study?
List of patients at a medical office
The respondents who fill out the survey
All the people who refused to be tested for HIV in the population
All the people who have been tested for HIV in the population
All the people who have HIV in a population
------------
Which of the following statements about the null hypothesis is NOT correct?
Rejecting or failing to reject the null hypothesis relies on statistical tests.
Failing to reject the null hypothesis implies that there is no relationship between the variables
that are tested.
It is a statement that there is no relationship between variables being tested.
Accepting the null hypothesis indicates that there is a relationship between the variables that
are tested.
Rejecting the null hypothesis implies that there is a relationship between the variables being
tested.
Elevated erythrocyte sedimentation rate (ESR) is associated with lymphoma. A group of investigators
intends to evaluate this association. Which of the following is the best null hypothesis for this study?
A high ESR level carries an increased risk of lymphoma.
A high ESR level is related to the occurrence of lymphoma.
A high ESR level has no association with lymphoma.
A high ESR level reduces the risk of lymphoma.
Lymphoma can be predicted by a high ESR level.
Which of the following is determined prior to data collection by the research team, in order to decide
between rejecting the null hypothesis and failing to reject the null hypothesis?
Sample
Categorical variables
Alpha
P-value
Continuous variables
In a study investigating the association between high school test scores and the number of older siblings
of students, a p-value of 0.04 is found. If alpha is set to 0.05, which of the following statements is
CORRECT?
The null hypothesis could not be rejected.
There is no association between high school test scores and the number of older siblings.
The p-value of 0.04 proves that the results are not significant.
There is a significant association between high school test scores and the number of older
siblings.
More information is needed to conclude.
What statistical test would you use to determine whether the mean age of two groups of patients is the
same or different?
Regression
Chi-square
Correlation
ANOVA
T-test
What type of variable is required to perform a t-test?
Interval
Dichotomous
Continuous
Ratio
Categorical
Which of the following tests is used to determine whether two categorical variables are associated?
Chi-square
ANOVA
Correlation
Regression
T-test
Which of the following best describes the term confidence interval?
A range in which a value is probably a reality.
A value that lets one know the disease is more likely with a positive test result.
A range in which a value has no meaning.
A specific number that has a meaning.
The probability that rejecting the null hypothesis was done incorrectly,
STATISTICAL BIASES
Definition of Bias
Which of the following statements regarding the definition of bias is INCORRECT?
The error must be systematic and result in an incorrect association or estimate.
Bias cannot be measured using statistics because it comes from the research process itself.
It is impossible to determine the presence of bias after the data has been collected.
Bias is entirely or mostly avoidable.
It is easier to remove bias during design and implementation than during data analysis.
Which of the following was NOT discussed in the lecture as a potential outcome of bias?
Cause us to overestimate the size of a real relationship
Create errors that can be measured through statistical analysis
Cause us to underestimate the size of a real relationship
Mask an association between two variables that are related
Create a spurious relationship between two variables
What type of bias is exhibited when study participants are lost to follow-up?
Observer bias
Selection bias
Recall bias
Confounding variables
Lead-time bias
Which of the following is associated with bias?
Unavoidable error
Random error
Systematic error
Correct conclusions
Standard error of the mean
Selection Bias
Which of the following is the most important consideration in terms of study design to avoid selection
bias?
Sampling scheme
Informed consent
The method used for data collection
Total population
Time-frame of study
Based on the case-control study about pancreatic cancer and coffee consumption discussed in the
lecture, which of the following decisions by the researchers lead to selection bias?
Control subjects were less likely to enjoy the taste of coffee.
Control subjects were chosen from hospitalized patients.
Control subjects did not include those with pancreatitis.
Control subjects were chosen from patients with GI disorders.
Control subjects were more likely to leave the hospital before study completion.
In a study examining the relationship between socioeconomic status and health outcomes, subjects
were recruited by using a flyer and were asked to attend a session at 11 AM on a weekday. Which of the
following changes would most effectively decrease the risk for selection bias in the sample?
Change the study type from retrospective to prospective
Offer more enticing compensation for study participation
Change the time of the meeting to avoid conflicting with working hours
Recruit people at the hospital in person rather than with a flyer
Distribute the flyer to more areas
Healthy Worker Effect
What is the assumption that explains the “Healthy worker effect”?
Healthy workers are less likely to participate in studies.
Those who are working tend to have more money for health-related expenses.
Those who attend work on a given day are more likely to be healthy.
People who are working are less likely to seek the services of physicians.
Those who work tend to be a healthy subset of the general population.
What type of study is most likely to be affected by the "Healthy Worker Effect"?
Cohort studies
Mortality studies
Incidence studies
Prevalence studies
Randomized controlled trials
Loss to Follow-up Bias
Which of the following is true of loss to follow-up bias?
Those who leave the study are likely to be the same as the people who stay.
It is a subtype of measurement bias.
It is known as Berkson`s bias.
It does not influence the outcome of the study significantly.
The group that leaves the study can cause over or underestimation of the results.
--------
Berkson Bias
In a case-control study of the relationship between dietary intake of fruits and vegetables and the risk of
developing hypertension, the control subjects are samples from participants at a health fair. What bias
or error is likely to result?
Detection bias
Non-response bias
Information bias
Berkson’s bias
Protopathic bias
What is the most common cause of bias in medical research?
Measurement bias
Information bias
Selection bias
Confounding bias
Recall bias
-----
Information Bias
Which of the following is NOT a type of information bias?
Differential misclassification
Non-differential misclassification
Recall
Non-response
Interviewer
What is the definition of information bias?
A bias that results from a systematic error in measurement.
A bias that results from a person answering a question or survey inaccurately.
A type of bias that results from selected study subjects not being representatives of the study
population.
A bias that results when the early diagnosis of a disease falsely makes it look like people survive
longer.
A phenomenon that results when the effect of the main exposure is mixed with the effect of
extraneous factors.
Misclassification Bias
What characteristic of differential misclassification bias defines it relative to non-differential
misclassification bias?
It leads to underestimating the relationship between exposure and outcome.
The effect of the bias favors the null hypothesis.
The records for only a certain subset of the sample are recorded incorrectly.
It leads to overestimating the relationship between exposure and outcome.
It skews the results in one direction.
Which of the following statements about non-differential misclassification bias is INCORRECT?
The bias does not differ between study groups.
Each group or category of variable has the same probability of being misclassified for all study
subjects.
Rejecting the null hypothesis when it is true is a likely consequence of non-differential
misclassification bias.
If the data is collected correctly then it is avoidable.
The bias is inherent in the data collection methodology.
What type of study is most vulnerable to recall bias?
Case-control study
Cohort study
Prospective survey study
Retrospective survey study
Randomized control trial
Which of the following strategies would be the LEAST effective in reducing interviewer bias?
Using a computer to communicate with the subject with the interviewer in another room
Avoiding questions that violate social norms on survey questionnaires
Randomly assigning subjects to different interviewers
Training interviewers to use neutral body language and tone of voice
Using a structured process to record observations/responses
Response Bias
What is the term used to describe when scientists only report or publish positive results?
Response bias
Confirmation bias
Reporting bias
Misclassification bias
Observer bias
Which of the following describes a systematic review that only lists studies that agree with its thesis?
Reporting bias
Response bias
Confirmation bias
Observer bias
Misclassification bias
Which of the following is an example of response bias?
Patients in the treatment group spend more time in specialized hospital units.
The respondents answer questions on a survey untruthfully to portray themselves in the best
light.
A researcher only cites studies that align with her findings.
Selecting control subjects for a case-control study from hospitalized patients.
A researcher believes in the efficacy of the treatment, so he is more likely to document the
positive outcomes.
Detection Bias
A study is conducted to examine the association between thrombocytopenia and the diagnosis of a
rare form of cancer. The study leads to an increase in screening and early detection of this type of
cancer in the subjects. Which type of bias does the vignette portray?
Detection bias
Confirmation bias
Response bias
Reporting bias
Misclassification bias
------------------------------
Hawthorne and Rosenthal Effect
A study seeks to measure how often store clerks ask for proof of age from customers trying to buy
cigarettes. The investigator introduces herself to the proprietors of several corner stores, explains the
purpose of the study, then watches to see whether the proprietor asks for identification from incoming
customers. Which of the following changes to the study design would be most helpful in avoiding the
Hawthorne effect?
Use retrospective data from cameras in the store rather than having an interviewer present while the
sales are taking place.
Assign stores randomly amongst a variety of interviewers known to the clerks, and take data
from random site visits.
Ask the store clerks to note every instance they ask for proof of age over a specific period of
time.
Ask the store clerks to retrospectively self-report the number of instances they ask for proof of
age.
How a study subject responds may be affected by a researcher’s expectations of the study subject.
Which of the following is NOT one of the terms used to describe this phenomenon?
Golem effect
Pygmalion effect
Hawthorne effect
Rosenthal effect
Self-fulfilling prophecy
Confounding Bias
A study finds an association between birth order and the presence of Down's Syndrome, specifically that
3rd, 4th, or later order children are more likely to have Down's Syndrome than are children who are
born first. What is the likely confounder in this example?
Gender
Smoking
Socioeconomic status
Maternal age
Any of the traditional confounders
How do we identify a characteristic as a confounder during the statistical analysis of a study?
It is impossible to determine if a characteristic is a confounder by statistical analysis. It must be
determined before study design.
Stratify the sample by the characteristic and if there is no difference in direction between the
groups then it is a confounder.
Stratify the sample by the characteristics and if the confounder influences both the dependent
and independent variables, it is a confounder.
Consult other studies of the same topic to determine if the characteristic is in the causal
pathway.
Determine if the characteristic is in the causal pathway, then it is not a confounder.
---------------
Effect Modification
Which of the following statements about effect modifiers is INCORRECT?
Effect modifiers are described as an “interaction term” in regression analysis.
Effect modifiers modify the nature or direction of a real association.
Effect modifiers are not a type of bias.
Effect modifiers may create the illusion of an association when none exists.
Effect modifiers do not mask the association between two variables that are associated.
Many people who only receive their news from social networking sites often find that the information
they receive tends to affirm the views of their general social circle rather than presenting an objective
view of the news cycle. What type of bias is this?
Clustering illusion
Confirmation bias
Response bias
Reporting bias
Hindsight bias
What type of bias is responsible for the tendency for people to see patterns in horoscopes?
Hindsight bias
Clustering illusion
Information bias
Confirmation bias
Confounding
DESCRIPTIVE EPIDEMIOLOGY
Descriptive Epidemiology: Introduction
Which of the following correctly defines descriptive epidemiology?
Measuring associations between causal factors
Characterizing the amount and distribution of health and disease within a population
Characterizing the amount of health and disease within a population
Characterizing the distribution of health and disease within a population
Characterizing the amount of health within a population
Which of the following best describes the term incidence?
A way of depicting an age distribution of a population.
The percentage of people who have a positive screening test for a disease.
The proportion of new cases in a population at risk.
The proportion of a population that has died over a defined period.
The proportion of all present cases of a disease in a population.
From Jan 2008 to Dec 2008, the town of Cityville had 1000 citizens. At the start of the year, there were
150 known cases of Mystery Disease. Over the course of the year, 50 additional cases of Mystery
Disease were detected. What was the cumulative incidence of Mystery Disease in Cityville at the end of
December?
50 cases per 1000 people per year
150 cases per 1000 people a year
200 cases per 1000 people per year
10 cases per 100 people per year
100 cases per 1000 people per year
From Jan 2008 to Dec 2008, the town of Cityville had 1000 citizens. At the start of the year, there were
150 known cases of Mystery Disease. Over the course of the year, 50 additional cases of Mystery
Disease were detected. What was the prevalence of Mystery Disease in Cityville at the end of 2008?
There is not enough evidence to calculate.
50/1000 or 5%
150/1000 or 15%
200/1000 or 20%
100/1000 or 10%
If the population of a region is 292,287,454 and there were about 173,770 new cases of lung cancer in
2017 and about 160,440 people die of lung cancer. What is the incidence rate of lung cancer in the
region for this year?
2.49 x 10-4
4.49 x 10-4
6.49 x 10-4
5.95 x 10-4
3.49 x 10-4
-------------
Mortality Rate
The population of Cityville in 2008 was 1000 people, that same year, a total of 200 people had Mystery
Disease, of which 20 died. What is the CFR of the Mystery Disease?
1/100 or 1%
1/200 or 0.5%
20/200 or 10%
200/1000 or 20%
20/1000 or 2%
Which of the following correctly defines the term, child mortality rate?
The number of maternal deaths due to childbearing per 100.000 live births.
The number of deaths of children younger than one year old per 1,000 live births.
The number of deaths of children younger than eighteen years old per 1,000 live births.
The number of deaths of children younger than five years old per 1,000 live births.
The total number of deaths per year per 1,000 people.
What is the meaning of having a standardized mortality rate > 1.0 in a study population?
The study population is experiencing a higher than expected death rate.
It tells us little because we need to know the standard ratio first.
There is no significant change in the mortality rate.
The prevalence of the death rate is increasing.
The study population is experiencing a lower than expected death rate.
----------
Population Pyramid
What does the horizontal axis in a population pyramid represent?
The number of people
The number of young adults
The number of women
The number of healthy people
The number of people with a disease
-------------------------
DATA
Measurement – Data
What is the term used to describe the determination of the dimensions of something using a standard
unit?
Data
Variable
S.I Unit
Logic
Measurement
Which component of the variable is the statistical component?
Logical component
Conceptual component
Observational component
Moderating component
Dependent component
Which of the following is an example of a continuous variable?
Gender
Age group
Citizenship
Height
Race
------------
Levels of Measurement – Data
What level of measurement is frequently found on surveys and is essentially a label?
Interval
Nominal
Ratio
Ordinal
Continuous variable
A medical intake questionnaire asks, "On a scale of 1 to 10, how much pain are you in?" What level of
measurement is denoted by this question?
Ratio
Interval
Nominal
Ordinal
Which level of measurement identifies numbers where the distance between values is assumed equal?
The mathematical level of measurement
Ratio level of measurement
The nominal level of measurement
Interval level of measurement
Which level of measurement has a true zero and allows one to count, rank, add, subtract, multiply and
divide?
Ordinal level of measurement
Mathematical level of measurement
Nominal level of measurement
Interval level of measurement
Ratio level of measurement
Which level of measurement allows numbers to be assigned to objects or events and represents the
rank order of the entities assessed?
Ratio level of measurement
Interval level of measurement
Ordinal level of measurement
Nominal level of measurement
Mathematical level of measurement
Distribution – Data
What is the significance of the p-value?
It is a representation of relative risk reduction.
It is a representation of absolute risk reduction.
It is a numerical representation of the mean of a data set.
It represents the true odds ratio.
It allows one to accept or reject the null hypothesis.
What is the vertical axis (Y) used for in graphing frequency distributions?
To display frequency
To display categories
To display scores
To show percentile
To display population
Type I and II Errors
Which of the following would be an accurate description of type I error?
The jury correctly rejects the null hypothesis and correctly identifies a person to be guilty.
The jury rejects the null hypothesis and considers a person to be guilty when in reality, he is
innocent.
None of the above are correct.
The jury does not reject the null hypothesis and considers a person to be innocent when in
reality, he is guilty.
The jury doesn't reject the null hypothesis and correctly identifies a person to be innocent.
How does a Type 1 error occur?
A researcher fails to reject a null hypothesis that is really false.
A researcher fails to reject a true null hypothesis.
A researcher incorrectly rejects a true null hypothesis.
A researcher fails to select study subjects that are representatives of the study population.
A systemic error in which the collection, analysis, or interpretation leads to conclusions that are
different from the truth.