Research Methods
Chapter 1
The empirical approach
Definition: an approach to the study and explanation of
psychological phenomena that emphasizes objective
observation and the experimental method as the source of
information about the phenomena under consideration
• Conclusions are based on direct evidence, such as
observations or measurements.
• Hypotheses must be falsifiable (must be true or false).
• Variables must be operationalized (make an operational
definition).
• Avoids the pitfall of illusory correlation (the appearance of a
relationship that in reality does not exist).
• Conclusions undergo peer review.
* Intuition is still useful for developing hypotheses in the first
place!
Goals of scientific psychology
Definition: the body of psychological facts, theories, and
techniques that have been developed and validated through
the use of the scientific method.
• Describe behavior
– (e.g., people are more likely to be persuaded by a
speaker they find attractive.)
• Predict behavior
– (e.g., short-term memory decreases as we get older.)
• Determine the cause(s) of behavior
– Remember, correlation does not imply causation! And
sometimes correlations are completely spurious.
• [Link]
• [Link]
• Explain behavior
– Why does a given behavior take place?
Basic vs. applied research
• Basic research: Attempts to answer fundamental questions
about the nature of behavior.
• Applied research: Attempts to apply research findings to
some specific real-world issue(s).
– e.g., evaluating a clinical intervention
– e.g., program evaluation
Research Methods
Chapter 2
Getting started
Hypotheses and specific predictions
How does one come up with an idea for a psychological study?
- Your Own Interests
- Observations
- Past Research
- Theories (esp. two competing theories)
- Applied practical issues
Anatomy of an APA-format research article
• Title page
• Abstract (Do it Last)
• Introduction (State past research, use of this paper, then state thesis
last)
• Method
- Participants
- Materials
- Procedure (Separate Materials and Procedure)
• Results
• Discussion
• References (After Main Paper)
• Appendix (After Main Paper)
• Figures
CHECK PAGE 391 IN TEXTBOOK FOR SAMPLE
PAPER
Literature searches (for the “Introduction” section
of paper)
• Use the PsycInfo database, not the web!
• Log in through the NSU Library databases page
• You may wish to limit your search by using only peer-reviewed journal
articles
• Sometimes you can get articles online; but often you have to actually
go to the library and xerox the journal articles
Research Methods
Chapter 3
Current Ethical Guidelines
• Developed from the Belmont Report (1979), which led to the modern-
day Institutional Review Board (IRB) and the American Psychological
Association (APA) Ethics Code.
– [Link]
– [Link]
• Assessment of risks and benefits
• Informed consent
• Debriefing
Deception & withholding information
• Is deception always unethical?
• Alternatives to deception:
– Role-playing
– Simulation studies
– Non-deceptive studies
Institutional Review Board (IRB) Approval
• Check out the NSU IRB at:
IRB Forms & Templates | Institutional Review Board ([Link])
and click on the “manuals and forms” link.
• You will almost always need the following:
– Informed consent forms
– Description of study, ensuring no harm or violations of privacy
will come to participants
– Inducements to participate: not too large
– Debriefing form
A few final ethics issues
• Humane treatment of animals
• Plagiarism
• Fabrication/fraud
– The case of Diederik Stapel
• [Link]
[Link]
• [Link]
Research Methods
Chapter 4
Empirical Studies
• Testable (i.e., falsifiable) hypotheses
• Operational definitions of variables
• Clear procedure
• Studies must be replicable.
• Minimize observer/experimenter bias
– (Rosenthal: The self-fulfilling prophecy: a belief or
expectation that helps to bring about its own fulfillment, as, for
example, when a person expects nervousness to impair their
performance in a job interview or when a teacher’s
preconceptions about a student’s ability influence the child’s
achievement for better or worse)
[Link] Method
Descriptive
Requires interjudge reliability (the extent to which independent
evaluators produce similar ratings in judging the same abilities or
characteristics in the same target person or object). Either 80% or 90%
of the time.
A. (Ethnography): the descriptive study of cultures or societies based on
direct observation and some degree of participation. not done often within
psychology, more within anthropology.
B. Participant observation: a quasi-experimental research method in
which a trained investigator studies a preexisting group by joining it as a
member, while avoiding a conspicuous role that would alter the group
processes and bias the data
C. Archival analysis: the use of books, journals, historical documents, and
other existing records or data available in storage in scientific research
2. Correlation Method
• Often used in surveys.
• Measured using a correlation coefficient
– For example: (r = .75)
• Correlation does not (always) imply causation!
[Link] Method
• Requires independent and dependent variables.
• Requires random assignment to conditions.
• You can infer causality! Yay!
• Between-participants vs. within-participants designs
Three Validities
• Construct validity: Does the operational definition of a variable actually
reflect the true theoretical meaning of the variable?
• Internal validity: Can you draw conclusions about causal relationships
from the data?
• External validity: Can the results be generalized to other populations
and settings in the “real world”?
• The internal vs. external validity dilemma
• Suppose you wanted to find out whether playing violent video games
makes teenagers more aggressive. How would you find this out using
observational, correlational, and experimental methods? What are the
advantages and disadvantages of each method?
Research Methods
Chapter 5
Reliability of measures
• Test-retest reliability: do it multiple times to confirm the reliability of
the study
• Internal consistency reliability: the degree of interrelationship or
homogeneity among the items on a test, such that they are consistent
with one another and measuring the same thing
– Split-half reliability: a measure of the internal consistency of
surveys, psychological tests, questionnaires, and other
instruments or techniques that assess participant responses on
particular constructs
– Cronbach’s alpha: a measure of the average strength of
association between all possible pairs of items contained within a
set of items (allows you to find bad items on a test)
• Interrater reliability: the extent to which independent evaluators
produce similar ratings in judging the same abilities or characteristics
in the same target person or object
* Remember, reliability is important – but it doesn’t say anything about a
measure’s accuracy! For that, you need validity…
Construct validity of measures
• Face validity
• Criterion-oriented validity
– Predictive validity
– Concurrent validity
– Convergent vs. discriminant validity
* Beware of reactivity!
* Most personality scales in psych. literature have already established
reliability and validity. Take advantage of this!
Types of measurement scales
Nominal scales: a sequence of numbers that do not indicate order,
magnitude, or a true zero point but rather identify items as belonging
to mutually exclusive categories
Ordinal scales: a sequence of numbers that do not indicate magnitude
or a true zero point but rather reflect a rank ordering on the attribute
being measured
Interval scales: a scale marked in equal intervals so that the difference
between any two consecutive values on the scale is equivalent
regardless of the two values selected. Interval scales lack a true,
meaningful zero point, which is what distinguishes them from ratio
scales
Ratio scales: a measurement scale having a true zero (i.e., zero on the
scale indicates an absence of the measured attribute) and a constant
ratio of values. Thus, on a ratio scale an increase from 3 to 4 (for
example) is the same as an increase from 7 to 8
Research Methods
Chapter 6
Quantitative vs. qualitative approaches
• Quantitative:
– Focuses on specific behaviors that can be measured.
– Numerical values assigned to responses
– Data analyzed using statistics
• Qualitative:
– In-depth study of a behavior or group, without attempting to
measure outcomes
– Describe or capture themes that emerge from the data
– Conclusions based on researcher’s interpretation
Naturalistic Observation (also called
ethnography)
• Researchers immerse themselves in a natural setting to make
observations.
• Primarily qualitative
• Non-participant vs. participant observers
• Limitations:
– Useful for developing theories, but not for testing hypotheses
– Must revise conclusions in light of negative cases
Systematic Observation
• Scientific observations of specific behaviors in a particular setting
• Requires testable, empirical hypothesis.
• Requires a coding system with high inter-judge reliability.
• Be careful of reactivity!
Case Studies
• Detailed description of a single individual
– (e.g., a description of a patient by a clinical psychologist)
• Psychobiography: case study in which an individual’s entire life is
analyzed. (Often used with historical figures)
– [Link]
• Useful for describing a rare or unique case; otherwise, larger samples
= better
Archival Research
• Statistical records
• Survey archives
• Written records
• Content analysis of documents
– [Link]
pennebak#explorations-into-language
Research Methods
Chapter 7
The purpose of surveys
• A “snapshot” of how people think & behave at a given point in time.
• Can also be administered in several “waves” to see the pattern of
change in thoughts & behaviors over time.
• Always be careful of social desirability, or “faking good.”
• Many survey instruments can be found online:
– [Link]
Constructing Questions to ask
• First, determine your research objectives.
• Surveys can be used to ask about:
– Attitudes and beliefs
– Facts and demographics
– Behaviors
• The biggest problem is that people are often not objective at
evaluating themselves!
Potential problems with item wording
• Unclear wording
– Unfamiliar or vague terms
– Ungrammatical sentence structure
– Questions that are too long
– Misleading information
• Double-barreled questions
• Loaded questions
• Negative wording
• Yea-saying and nay-saying
Response Alternatives
• Open-ended vs. closed-ended
• Number of response alternatives
• Types of rating scales
– Likert-type scales
– Graphic rating scale
– Nonverbal scale (for the kids!)
Finalizing the questionnaire
• Format: attractive and professional
• Use scale format consistently throughout (e.g., don’t change from 5-
point scales to 4-point scales to 7-point scales).
• Ask demographic questions LAST so that they don’t influence other
answers.
• Do an informal “pilot study”: show the survey to a few people.
Administering surveys
• Written questionnaires
– Groups vs. individuals
– Mail surveys (inexpensive, but low response rate)
– Internet surveys
• Interviews
– Beware of interviewer bias!
– Face-to-face, telephone, and focus groups
Sampling
• Sample size (see Chap. 7, Table 1)
– Confidence intervals: a range of values for a
population parameter that is estimated from a sample with a
preset, fixed probability (known as the confidence level) that the
range will contain the true value of the parameter
– [Link]
• Representative samples: the selection of study units (e.g., participants,
homes, schools) from a larger group (population) in an unbiased way,
such that the sample obtained accurately reflects the total population.
For example, a researcher conducting a study of university admissions
would need to ensure they used a representative random
sample of schools—in other words, each school would have an equal
probability of being chosen for inclusion, and the group as a whole
would provide an appropriate mix of different school characteristics
(e.g., private or public, student body size, cost, proportion of students
admitted, geographic location).
– (see Chap. 7, Table 2)
– [Link]
• Beware of low response rates!
• Reasons for using convenience samples
Research Methods
Chapter 8
Why experiments are cool
As long as the experiment is designed properly (i.e., no confounding
variables), one can infer that only the independent variable causes an
effect on the dependent variable
Posttest-only design
R = random sampling
Pretest-posttest design
• Pretest is given before the experimental manipulation is introduced to
make sure groups are equivalent at the beginning of the experiment.
• Good idea when you have small sample size, or danger of mortality.
• Big risk: Could sensitize participants to what is being studied!
Two types of study designs
1. Independent groups design (also called between-participants design)
2. Repeated measures design (also called within-participants design)
Repeated measures design
• Advantages
Fewer participants
Eliminates the effects of outliers (i.e., “weirdos”)
• Disadvantages
Order effects:
- Practice effect
- Fatigue effect
- Contrast effect
* Can be beaten by COUNTERBALANCING!
Counterbalancing
R = random sampling
Latin squares design
• Designed to reduce order effects, when you can’t run all possible
orders
– (e.g., what if you have 5 conditions, or 120 different orders?
That’s way too many, dude. Use a Latin squares design instead!)
• You could also just wait a long time between the conditions to avoid
order effects.
Matched-pairs design
• Goal is to match participants on the dependent measure
• Ensures that the groups are equivalent prior to introduction of
independent variable manipulation
• Analysis of Covariance (ANCOVA)
Research Methods
Chapter 9
Thoughts on selecting participants
• When it is important to describe population accurately (e.g., with
political polls), you must use probability sampling.
• When you’re just interested in testing hypotheses about behavior
(which probably describes most of your ideas), you can probably just
use convenience sampling.
Manipulating the I.V.
• Straightforward manipulations
VS.
• Staged manipulations
– Frequently employ a confederate
• Strength of the manipulation
• Cost, time factors
Measuring the D.V.
• Types of measures
– Self-report
– Behavioral
– Physiological
• Sensitivity (i.e., the scale being used)
– Beware of ceiling and floor effects!
• Mulitiple D.V.’s = good (usually).
• (Be aware of cost, ethics)
Additional Controls
• Be careful of demand characteristics!
– In questionnaires, can include filler items.
• Placebo groups
– Balanced placebo design
• Expectancy effects (also called experimenter bias)
– Single-blind and double-blind experiments
Additional Considerations
• Research proposals
• Pilot studies
• Manipulation checks
• Debriefing
• Analyzing and interpreting results
• Publicizing research results:
– Posters at psychology conferences
– Peer-reviewed journal articles
Research Methods
Chapter 10
Increasing the number of levels of an
independent variable
• Can provide more information about the relationship than a two level
design.
• Check page 226 for another example of this graph design.
• Solid line indicates two conditions.
• Dotted line addresses the amount of reward needed to get a high-
performance level. Thus, it gives more information.
When you have more conditions, you need more participants
Downside: minimizing statistical power
• Curvilinear relationship (Page 227):
Answer for this probably that it depends but if you left out lvl3 then it
shows the dependent is a positive effect of the independent variable
If no lvl1, then the dependent would be a negative effect of the
independent variable
If no lvl2, then there would be zero effect
Therefore, having all three levels show the best result of the study
Increasing the number of IV’s: Factorial
Design
Example of a 2 x 2 factorial design:
You can have multiple types of factorial design
The numbers refer to the conditions of the independent variables
An interaction is a combination of variables
Three questions when 2 x 2: What is the main effect of IV 1? What is
the main effect of IV 2? Does the effect of IV 1 depend on the 2 nd IV?
Interpretation of Factorial Designs
o Is there a significant main effect for independent variable A?
o Is there a significant main effect for independent variable B?
o Is there a significant interaction between IV “A” and IV “B” ?
IV x PV (Participant Variable) designs
TRICK FOR FINDING INTERACTIONS: PUT A DOT ON THE TOP OF
EACH BAR AND DO A DOTTED LINE TO EACH BAR OF THE SAME
COLOR (if the two lines are parallel, then that means no
interaction; no parallel lines mean interactions)
What are the main effects and interactions in
this example?
1. No main effects
2. No main effects
3. No interactions
What are the main effects and interactions in
this example?
1. No main effects (Average of 5 and 5)
2. Yes main effects
3. No interactions
What are the main effects and interactions in
this example?
1. Yes main effects
2. Yes main effects
3. No interactions (each does four points better, thus it doesn’t depend)
What are the main effects and interactions in
this example?
1. Yes main effects
2. Yes main effects
3. Yes interactions
What are the main effects and interactions in
this example?
1. No main effects
2. No main effects
3. Yes interactions
“Mixed” factorial designs
• When one IV uses an independent groups design, and the other IV uses
a repeated measures design
* For example:
• IV #1 = watch violent vs. non-violent TV shows
(repeated measures)
• PV #2 = low vs. high self-esteem
(independent groups participant variable)
• DV = aggressive behavior
• CHECK PAGE 236
Increasing the number of IVs in a factorial
design
• (e.g., imagine a 2 X 3 X 2 factorial design. How many conditions are
there?)
Research Methods Chapter
11
Single case experimental design
• Reversal design (ABA design)
– A (baseline) -> B (treatment) -> A (baseline)
• Multiple baseline design
• Page 249 Figure 2 Graph
Program Evaluation
• Research on programs that are proposed and implemented to achieve
some positive effect on a group of individuals
– (e.g., does DARE reduce drug use?)
Quasi-Experimental Designs
One-Group Posttest-Only Design:
(Page 251)
Nonequivalent Control Group Design:
Nonequivalent Control Group Pretest-Posttest Design:
(Page 256)
• Interrupted Time Series Design
Examines the dependent variable over an extended period of
time, both before and after the IV is implemented
Interpretation problems (possible regression to the mean)
• Control Series Design
Improves interrupted time series design by finding an
appropriate “control group”
Involves finding a similar population that did not receive a
particular manipulation
Limited because this is not a true “control group”
Developmental Research Designs
Cross-Sectional Method – persons of different ages measured at the
same point in time
Longitudinal Method – same group is observed at different times (as
they age)
Sequential Method – combination of cross-sectional and longitudinal
designs
Research Methods Chapter
12
Review: Scales of Measurement
• Nominal
No numerical, quantitative properties
Levels represent different categories or groups
• Ordinal – minimal quantitative distinctions
Levels ordered from lowest to highest
• Interval – quantitative properties
Intervals between levels are equal in size
No absolute zero
Can be summarized using means
• Ratio – detailed quantitative properties
Equal intervals
Absolute zero
• Correlational studies use mainly interval scales, as it is the most
common scale
Review: Analyzing Results
• Correlational designs
– (correlating individual scores)
• Experimental and quasi-experimental designs
– (comparing group means or group percentages)
Frequency distribution
• First step – draw a picture! (Graph your data.)
– pie charts (nominal variable), bar graphs (useful for
experiments), frequency polygons, histograms
Descriptive Statistics
• Measures of central tendency
– Mean
– Median (Useful in ordinal scales & interval-ratio scales where
heavy outliers exist)
– Mode
• Measures of variability
(refers to the amount of spread in the distribution of scores)
– Standard deviation (s) or (SD)
– Variance (s²)
– Range
Review: Correlation Coefficients
• Pearson correlation coefficient (r)
• r can range from -1.00 to +1.00
Effect Size
• Refers to the strength of association between variables
• Some effect size statistics:
– r (size of correlation)
– r² (%-age of variance in Variable A accounted for by Variable B)
– z-score, t, F (in experiments)
-----------------------------------------------------------------
• Also important: statistical significance (p < .05)
Regression Equations
• Used to predict a person’s score on one variable (Y) when that person’s
score on another variable (X) is already known.
• Formula: Y = a + bX
Y = Score we wish to predict
X = Score that is known
a = constant (y-intercept)
b = slope of line
Multiple correlation
• Used to combine a number of predictor variables to increase the
accuracy of prediction of a given criterion or outcome variable
• Effect size measure: R
Partial Correlation and the third variable
problem
• A partial correlation is a correlation between two variables of interest,
with the influence of the third variable removed from, or “partialed out
of” the original correlation.
Structural equation modeling
• Testing an expected pattern of relationships among a bunch of
variables.
• For example:
Research Methods Chapter
13
Inferential Statistics
• Allows researchers to decide if differences in the sample data reflect
true differences in the population.
– (e.g., if you see that Group A got a score of 43 and Group B got a
score of 45, how do you know if this is a real difference, rather
than just random fluctuation around the population mean?)
Null and Research Hypotheses
• Null hypothesis (Ho): Population means are equal.
• Research hypothesis (H1): Population means are not equal.
• When there is a very low probability that the mean differences were
due to random error (i.e., the results were statistically significant), we
can reject Ho.
Sampling Distributions
• (e.g., the case of ESP: see Table 2, Page 303. About 20% of the time,
answers would be correct by chance.)
– Note that there is random fluctuation around this mean!
T-Tests
• Used to examine whether two groups are significantly different from
each other.
– (If they are significantly different, you can reject H o! If not, you
fail to reject Ho.)
group difference
– t=
within-group variability
• If you know the means, standard deviations, and N’s, you can calculate
a t-value. (See formula on p. 306.)
• Then calculate the degrees of freedom.
• Then look up critical t-value (see:
T-Distribution Table of Critical Values - Statistics By Jim).
– Decide on a 1-tailed or 2-tailed test!
F-test or analysis of variance (ANOVA)
• Used when your IV has more than two levels OR when you have more
than one IV.
• Ratio of between-group variance to within-group variance.
Effect size and statistical significance
• Effect size: How big is the effect? (Usually measured by r or d
statistics.)
• Statistical significance: “p < .05” means you can be more than 95%
sure that your findings are real, and not due to random fluctuations.
But what if you’re wrong? That would be a Type I error.
Type I and Type II errors
• Type I error: When you reject Ho (null hypothesis), but Ho is actually
true.
– This happens when you get statistically significant results, but
you should not have.
• Type II error: When you fail to reject Ho, but H1 (alternative hypothesis)
is actually true.
– This happens when you do not get statistically significant results,
but you should have.
Power Analysis
• Smaller effect sizes require larger samples to be significant.
• Power is a statistical test that determines optimal sample size based
on the probability of correctly rejecting the null hypothesis.
• Page 317 Table 5
Meta-analyses
• File-drawer phenomenon:
– Only the statistically significant results get published!
• A meta-analysis is a statistical summary of all studies on a given
variable.
Generalizing to other populations of research
participants
• Most common samples:
– college students
– volunteers
• Gender considerations
• Locale
• Generalization as a statistical
interaction
BUT:
* “In Defense of College Students and Rats”
Cultural considerations
• It is important to be aware of the ways in which the operational
definitions of the constructs that we study are grounded in a particular
cultural meaning.
– (e.g., cultural definitions in self-concept)
Generalizing to other experimenters:
• It is important to be aware of the ways in which the operational
definitions of the constructs that we study are grounded in a particular
cultural meaning.
– (e.g., cultural definitions in self-concept)
Pretesting
• Used to make sure that groups are equivalent before the experimental
manipulation.
• Useful for assessing mortality effects (i.e., you can see if people who
withdrew were different from people who completed the study).
Mundane and experimental realism
• Mundane realism: Whether the experiment bears similarity to events
that occur in the real world
• Experimental realism: Whether the experiment has an impact on the
participants, involves them, and makes them take the experiment
seriously
– (experimental realism = more important!)
Mutual benefits of lab and field research
• Conducting research in both laboratory and field settings provides the
greatest opportunity for advancing our understanding of behavior.
• Results of lab and field experiments have tended to be extremely
similar.
– Even the effect sizes are (usually) similar!
Replications are important
• Exact Replications: An attempt to replicate precisely the procedures of
a study to see whether the same results are obtained.
• Conceptual Replications: The use of different procedures to replicate a
research finding.
• Literature Reviews and Meta-Analyses help summarize the state of
literature on a given topic.
The importance of psychology
• Improving lives
• Improving our knowledge about the world
* Further info. about psychology & its applications:
• [Link]
• [Link]
• [Link]
• [Link]
• [Link]