Chapter1 Research Methods Notes
Chapter1 Research Methods Notes
Aim The intention of a study — the idea being tested, problem to be solved, or question being answered
IV Independent Variable — factor manipulated to create 2+ conditions, expected to cause change in the DV
Uncontrolled variable Variable with random OR systematic effect on the DV, not deliberately manipulated
Confounding variable Uncontrolled variable acting systematically on ONE level of the IV — confuses/hides/exaggerates the real effect
Independent measures design A different group of participants is used for each level of the IV
Repeated measures design The same participants perform in every level of the IV
Matched pairs design Participants paired on relevant traits; one member of each pair per condition
Random allocation Equal chance of being placed in any condition — spreads individual differences evenly
Demand characteristics Features giving away the study's aim; participants change behaviour to match expectations
Order effects Practice or fatigue effects from doing a task more than once (repeated measures)
Randomisation Random order of conditions allocated to each participant — fixes order effects
Counterbalancing ABBA design — half do condition A then B, half do B then A — fixes order effects
Participant variables Individual differences (age, personality, intelligence) that could distort results
Laboratory experiment IV + DV + strict controls, in an artificial (not the participant's usual) setting
Field experiment IV manipulated in the participant's NORMAL environment for that behaviour
Standardisation Keeping the procedure identical for every participant, raising reliability
Validity Extent to which the researcher is testing what they claim to be testing
Pilot study Small-scale preliminary test of the procedure before the main study
Replication Keeping procedure/materials exactly the same between studies to verify results
Operational definition Clear description of a variable so it can be manipulated, measured and replicated
Generalise To apply findings more widely, e.g. to other settings and populations
Ecological validity Extent to which findings from one situation would generalise to other (real-life) situations
Alternative hypothesis The main testable prediction of a difference (experiment) or relationship (correlation)
Informed consent Knowing enough about a study to decide whether to agree to participate
Right to withdraw A participant can remove themselves, and their data, from a study at any time
Confidentiality Results/personal info kept safe, not released outside the study
Self-report Method obtaining data by asking participants to give info about themselves
Open question Full, descriptive answers in participant's own words — produces qualitative data
Social desirability bias Trying to present oneself in the best light rather than answer honestly
Filler questions Irrelevant items added to a questionnaire to disguise its real aim
Structured interview Fixed, scripted questions in fixed order for every participant
Semi-structured interview Fixed core of open/closed Qs + interviewer can add more if needed
Triangulation Using different techniques on the same case to check consistency (validity)
Controlled observation Watching behaviour where the social/physical environment has been manipulated
Non-participant observer Researcher does not become involved in the situation studied
Inter-observer reliability Consistency between two observers watching the same event
Causal relationship A change in one variable is responsible for (causes) a change in another
Longitudinal study Follows the same group (cohort) of participants over time
Random sample Every population member has an equal chance of being chosen
Standard deviation Average distance of every score from the mean; bigger = more spread out
Bar chart Graph for discrete categories/totals — has gaps between bars
Histogram Graph for continuous data — bars touch (no gaps unless empty category)
Scatter graph Displays correlational data — each dot = one participant's two scores
Ethical issues Problems raising concerns about the welfare of participants or society
Presumptive consent Asking a similar (not actual) group whether a study would be acceptable
Bateson's Cube 3-factor model (benefit, research quality, suffering) justifying animal research
Test-retest Using a measure twice; a high correlation between scores = high reliability
Replicability The extent to which a study's procedure can be kept the same when repeated
Generalisability How widely findings apply, e.g. to other settings and populations
1.1 Experiments
DEFINITION: An experiment is an investigation that allows researchers to look for a cause-and-effect relationship. The researcher
manipulates the IV to produce two or more "levels"/conditions and measures the effect on the DV. If there is a big difference in the DV
between conditions, this suggests the IV caused the difference.
MEMORY TRICK: IV = "I Vary it" (researcher controls it). DV = "Data" you collect (what changes because of the IV).
EXAMPLE: IV = brightness of lighting (bright/dull). DV = how well people pay attention. To be certain the difference is caused by the
IV, the researcher must control other variables that might affect the DV (e.g. whether people have recently eaten, exercised, or sat
through a dull class).
IMPORTANT — don't confuse: a RANDOM uncontrolled variable affects all conditions equally (less serious). A CONFOUNDING variable
affects only ONE condition (serious — hides or exaggerates the IV's real effect).
EXAMPLE: Studying chocolate's effect on attention — one bar vs two bars = two experimental conditions. One bar vs no chocolate
at all = one experimental + one control condition.
MEMORY TRICK: "I R Match" = Independent measures / Repeated measures / Matched pairs
A DIFFERENT group of participants is used for each level of the IV. Data for each level is "independent" — not related to any other data
because it comes from different people.
Random allocation reduces individual-differences bias: give each participant a number, then randomly divide numbers into groups
(drawing from a hat, or a random number generator).
DON'T CONFUSE: Practice effect = performance IMPROVES (familiarity/learning). Fatigue effect = performance DECLINES
(tiredness/boredom). Both are "order effects".
MEMORY TRICK — fixing order effects: "R and C save the day" = Randomisation (random order per participant) OR
Counterbalancing (ABBA — half do A then B, half do B then A).
EXAMPLE: Experiment comparing learning with music (M) vs no music (N). Randomisation = participants randomly allocated to do M-
then-N or N-then-M. Counterbalancing = the group is split — half do M then N, half do N then M.
Participants are arranged into PAIRS, matched on variables relevant to the study (age, gender, intelligence, personality). One member
of each pair performs in each condition. Identical twins make ideal matched pairs — genetically identical, similar experiences.
Independent No order effects. Less demand characteristics (see one level only). Random Participant variables can distort results. More participants
Measures allocation can reduce individual differences. needed → less ethical/effective if hard to find or at risk.
Repeated Participant variables unlikely to distort effect (each participant does all Order effects could distort results. Greater exposure to demand
Measures levels). Counterbalancing reduces order effects. Fewer participants needed. characteristics (task done more than once).
Matched Participants see only one level → less demand characteristics. Participant Similarity between pairs limited by matching criteria chosen.
Pairs variables less likely to distort effect (matched). No order effects. Availability of matching pairs may be limited → small sample.
Types of Experiment
Laboratory experiment: has an IV, DV and strict controls; conducted in a setting that is NOT the usual environment for the
participants.
Field experiment: IV manipulated, DV measured; conducted in the participants' NORMAL environment for the behaviour being
investigated; some control of variables still possible.
EXAMPLE: Lab: testing children's attention in bright/dull lighting via a computerised task in a university room (not their classroom).
Field: testing the same IV/DV by altering the number of lights turned on in the children's normal classroom, measured by scores on a
topic test.
Strengths Good control of variables, raising validity. Causal relationships Participants likely to behave naturally, results representative. If unaware
can be determined. Standardised procedures raise reliability and they're in a study, demand characteristics are lower than in the lab.
allow replication.
Weaknesses Artificial situation could make behaviour unrepresentative, Control of variables is harder → lower reliability, harder replication. Less
lowering ecological validity. Participants could respond to certain the IV caused DV changes. Participants may be unaware they're in a
demand characteristics. study → ethical issues.
Controls Ways to keep confounding variables constant between levels of the IV, raising validity
Standardisation Keeping the procedure identical for each participant, raising reliability
Validity The extent to which the researcher is testing what they claim to be testing
Pilot study A small-scale test of the procedure BEFORE the main study, to identify and fix problems
Replication Keeping procedure/materials exactly the same between studies to verify results
Ecological validity Whether findings generalise to the real world — depends on realism of setting + task (mundane realism)
IMPORTANT: A pilot study should NOT be used to test whether a study is ethical — that is the responsibility of the researcher/ethical
committee.
Hypotheses in Experimental Studies
Hypothesis: a testable statement based on the aims of an investigation. Must be falsifiable (possible to prove wrong) and must have
operationalised variables.
Alternative hypothesis: the main testable prediction of a difference (experiment) or relationship (correlation).
MEMORY TRICK: Non-directional = "No direction stated". Directional = "Direction is stated" (needs previous evidence). Null = "it's just
chance".
Non- Predicts an effect WILL happen, but No previous research to "There is a difference between the effectiveness of mind maps and
directional not the direction suggest direction revision apps in helping students to learn."
(two-tailed)
Directional Predicts the DIRECTION of the effect Previous evidence suggests "Students using revision apps will learn better than students using mind
(one-tailed) (which condition is a direction maps."
"best"/higher/lower)
Null States any difference/correlation is Always written as the "There will be no difference between the effectiveness of mind maps and
hypothesis due to CHANCE "backup" to the alternative revision apps in helping students to learn" / "Any difference is due to
hypothesis chance."
EXAM TIP: Inferential statistics decide whether a result is "significant" (a mathematically significant probability the pattern could NOT
have arisen by chance) — if significant, researchers REJECT the null hypothesis and ACCEPT the alternative hypothesis. If non-
significant, they accept the null hypothesis.
Ethics in Experiments
Lab participants likely know they're taking part → can give informed consent — BUT for validity, may need to hide the aim (deception)
to avoid demand characteristics
Field experiments: participants often unaware they're in a study → cannot give informed consent, cannot exercise their right to withdraw
→ presumptive consent may be needed, and it's vital participants are protected from harm
Deception should be avoided; where necessary, explain the reality afterwards in a debrief
Privacy easier to respect in lab (pre-planned questions/tests); harder in field (risk of invading personal space)
Confidentiality respected by keeping data secure/anonymous; in field experiments where participants don't know they're being
studied, it's vital they can't be individually identified (e.g. by workplace)
1.2 Self-Reports
Self-report: a research method (questionnaire or interview) that obtains data by asking participants to provide information about
themselves directly — different from experiments/observations where the researcher finds the data out FROM the participant.
MEMORY TRICK: Closed = "Countable" → quantitative data. Open = "Own words" → qualitative data.
Questionnaires
Questionnaire: a self-report method using written questions, 'paper and pencil' or online.
Closed Fixed set of possible responses, no opportunity to Quantitative Yes/no; item lists; rating scales (0–5); Likert scales (strongly agree →
expand strongly disagree)
Open Full, descriptive answers in the participant's own Qualitative "Why do you believe it is important to help people who suffer from
words phobias?"
Evaluating Questionnaires
✓ Strengths ✗ Weaknesses
Closed Qs easy to analyse (totals per Open answers need interpretation → risk of low inter-rater reliability if researchers differ
category, averages)
Open Qs give detailed, in-depth (qualitative) Easy to ignore → low return rate → respondents may share characteristics (e.g. unemployed/retired with spare
info time) → poor generalisability
Participants may lie to look more acceptable (social desirability bias) or if they've guessed the aim
Filler questions: irrelevant items added to disguise the real aim of a study (not analysed in the results).
Interviews
Interview: a research method using verbal questions asked directly (face-to-face, telephone, or real-time chat). Interviews often use
MORE open questions than questionnaires.
MEMORY TRICK: "SUS" = Structured (fixed, scripted) / Unstructured (flows from answers) / Semi-structured (fixed core + can add
extra Qs = best compromise).
Format Definition
Structured Questions in a fixed order, may be scripted; consistency required for interviewer's tone/posture too — fully standardised
Unstructured Most questions (after the first) depend on the respondent's answers; a list of topics may be given — very flexible but hard to compare between
participants
Semi- Fixed list of open + closed questions, but interviewer CAN add more if necessary — comparable data AND explores individual issues
structured
Evaluating Interviews
Interviewees may lie — either social desirability bias, OR because they think they know the aim and are trying to help (or disrupt) the
research
Time-consuming → restricts the type of participant who volunteers → narrow representation of feelings/beliefs/experiences
Researchers must avoid being subjective when interpreting responses — aim for objectivity; may ask other experienced (but aim-
unaware) researchers to interpret findings
DON'T CONFUSE: Subjectivity = interpretation biased by personal feelings/beliefs (bad for validity). Objectivity = unbiased,
consistent regardless of who's interpreting (good for validity).
EXAM TIP: Questionnaires/structured interviews tend to be MORE reliable (consistent administration, numerical results, no
interpretation needed). Open questions give more VALID, in-depth info but are harder to interpret consistently.
Case study: a detailed investigation of a SINGLE instance — usually one person, but could be a family or institution. Uses a VARIETY of
techniques (interviews, observations, tests, questionnaires) AND different sources (the participant themselves, relatives, colleagues,
existing medical/school records).
Particularly useful for: rare cases needing detailed description; tracking developmental change (progress of a child, or
improvement/decline of a disorder)
Often linked to therapy, BUT the therapeutic purpose is NOT the main research aim
MEMORY TRICK: Triangulation = using different techniques on the same case — if results agree, this suggests the findings are
VALID.
Highly valid — individual explored in great depth, in genuine Close researcher-participant relationship risks subjectivity → lowers validity
context (work/family)
Validity improved further by triangulation Detail can threaten confidentiality/privacy (ethical issue) — hard to disguise identity even
without real names
EXAMPLE: A sleep clinic case study on patient "PMA" with severe nightmares — used EEG (brain wave recording), a dream diary, and
family interviews (triangulation across multiple techniques).
1.4 Observations
MEMORY TRICK: 4 decisions for every observation — "S-F-R-O": Setting, Focus, Role, Obviousness.
Setting Naturalistic: watching behaviour in the normal environment, no researcher interference (social or physical). Controlled: environment (social or
physical) manipulated by researchers — can happen in normal OR artificial (lab) settings.
Focus Unstructured: records the WHOLE range of possible behaviours (often a pilot stage to refine categories). Structured: records only a LIMITED,
(range pre-defined range using behavioural categories (must be operationalised — observable actions, not inferred states).
recorded)
Role of Participant observer: part of the social setting (may talk/play with people being observed). Non-participant observer: does not become
observer involved (e.g. one-way glass, sits apart).
Obviousness Overt: role obvious to participants (e.g. holding a clipboard). Covert: role hidden/disguised (e.g. disguised as a group member, or physically
hidden via CCTV).
Evaluating Observations
Naturalistic: ✓ behaviours are "true to life"; ✗ no guarantee the target behaviour will actually occur
Unstructured: ✓ important unexpected behaviours are captured; ✗ very difficult to accurately record ALL activities, many may be
irrelevant
Structured: ✓ more reliable data (focused on a small number of categories)
Overt observer: more ethical but may alter participants' behaviour → lowers validity
Covert observer: more valid (no demand characteristics/social desirability) BUT practical issues (must stay hidden, harder data
collection) AND ethical issues (no informed consent, distress if discovered)
IMPORTANT: Covert + Participant observation = BEST validity (most natural behaviour) but WORST ethics (no informed consent,
invades privacy, risk of distress if the participant's role is discovered).
Inter-observer reliability = consistency between two researchers watching the SAME event — whether they produce the same records.
EXAM TIP: Observation can be used as a research method IN ITSELF, OR as a TECHNIQUE within another method (e.g. measuring the
DV in an experiment, a co-variable in a correlation, or gathering data within a case study).
1.5 Correlations
Correlation: a research method that looks for a relationship between two MEASURED co-variables. Used when it's not practical or
ethical to manipulate variables (i.e. can't do an experiment). Both variables must exist over a range and be measurable numerically
(durations, tallies, ratings, test scores).
THE #1 EXAM TRAP: Correlation ≠ Causation. NEVER say one co-variable "causes"/"makes"/"leads to" the other. A THIRD variable
may be responsible for both. Only ever refer to "co-variables" — never "IV/DV" — in a correlation.
CLASSIC EXAMPLE: A bizarre POSITIVE correlation exists between ice cream consumption and murder rates — this does NOT mean
ice cream causes murder. Likely both are linked to a third variable: hot weather.
Classroom example: attention in class and test scores may correlate, but this doesn't prove paying attention CAUSES good scores — a
third variable (how hard-working the student generally is) may cause both.
MEMORY TRICK: Positive correlation = "both go up the stairs together". Negative correlation = "a see-saw — one up, one down".
Positive An increase in one variable accompanies an increase in Exposure to aggressive models & violent behaviour — greater exposure =
correlation the other higher violence
Negative An increase in one variable accompanies a decrease in Years in education & obedience — fewer years of education = more obedient
correlation the other
Good starting point for research — shows if a relationship is worth investigating Only valid if BOTH co-variables are measured in clearly-defined, effective
further (e.g. with an experiment) ways
Useful when not practically/ethically possible to conduct an experiment Reliability depends on measures being consistent — lower for self-
reports/observations than scientific scales
Strength shown by an 'r' value: close to +1 = strong positive; close to −1 = strong negative; close to 0 = no/weak relationship. Displayed
on a scatter graph.
Longitudinal study: follows a SINGLE group of participants (a cohort) over time (weeks to decades), studying variables at intervals to
explore development/change due to experiences (interventions, drugs, therapies).
Cross-sectional study: compares DIFFERENT groups of people at differing ages/stages at ONE point in time.
MEMORY TRICK: Longitudinal = "LONG" time, SAME group (one cohort). Cross-sectional = "CROSS-comparing" DIFFERENT age groups
at ONE time.
PROBLEM WITH CROSS-SECTIONAL: hard to separate changes over time from differences due to individuals growing up at different
times (different upbringing/societal expectations), rather than age itself. Longitudinal studies avoid this by using ONE cohort tested
repeatedly.
Different designs of longitudinal study: tracking ONE variable over time; recording TWO+ variables to look for correlations; or an
EXPERIMENTAL design ("quasi-experimental") — baseline (pre-intervention) measured, then post-intervention measure(s)/follow-up
Evaluating Longitudinal Studies
✓ Strengths ✗ Weaknesses
Confident changes are due to TIME passing, not cohort differences → most valid Sample attrition — loss of participants over time (withdrawal, boredom,
test of developmental change moving, poor health, homelessness, death)
Like repeated measures, retests same individuals → no participant-variable Smaller sample = less representative; survivors may be atypical/"special"
confound (motivated to compete/succeed)
Overcomes confounding situational variables (e.g. different educational Reliability issue — measures/researchers may change over a long period
methods) that would affect a cross-sectional study
CLASSIC EXAMPLE: Terman's "Life Cycle Study of Children with High Ability" (began 1922) — "Terman's Termites". Debunked the
myth that high IQ automatically leads to success (though some, e.g. psychologist Lee Cronbach, did become famous). Initial sample =
1528 children. Sample attrition: comparing 1996 & 1999 data, only 119 of 162 returned questionnaires provided data on BOTH
occasions (Holahan & Velasquez, 2011).
MEMORY TRICK: Participant variable = about the PERSON (age, personality, intelligence). Situational variable = about the
SITE/environment (noise, light, weather).
EXAMPLE: Study on age and false-memory susceptibility: IV = age group, operationalised as 'young' (under 20), 'middle-aged' (40–50),
'old' (over 70). DV = number of details "remembered" about the false memory, or how convinced participants were it was true.
Confounding variables act SELECTIVELY on one level of the IV, so can: (1) work AGAINST the IV's effect (masking a real effect), or (2)
increase the APPARENT effect of the IV (suggesting an effect that doesn't exist)
A pilot study helps identify problematic uncontrolled variables BEFORE the main study starts
EXAMPLE (Dr Huang's chair study): Comparing a class with soft chairs in one room vs hard chairs in another. If the soft-chair room
happens to have better lighting, this is a situational variable (confounding, related to the environment). If soft-chair students happen
to all do arts subjects and hard-chair students all do maths/science, this is a participant variable (confounding, related to individual
ability).
Standardisation
Standardised instructions: the same written/verbal info given to every participant, ensuring their experience (regardless of IV level)
is as similar as possible.
Every participant must be treated the SAME way — includes standardised instructions AND a standardised procedure (consistent
equipment/tests, measuring the same variable the same way every time)
Standardisation is easier in lab experiments (consistent equipment, e.g. stopwatches) than other studies; some measures (e.g. brain
scans) need interpretation, which must also be standardised
1.8 Sampling of Participants
Population: the group, sharing one or more characteristics, from which a sample is drawn.
Sample: the group of people selected to represent the population in a study.
Sampling technique: the method used to obtain participants from the population.
MEMORY TRICK: "O-V-R" = Opportunity (available) / Volunteer (they come to YOU via advert) / Random (equal chance, numbers from
a hat).
Samples should ideally be REPRESENTATIVE so findings can be generalised. Details commonly reported: age, ethnicity, gender, socio-
economic status, education, employment, geographical location, occupation
Sample SIZE matters — small samples are less reliable/representative (less likely to contain the full range of population variation)
Opportunity Participants chosen because they are available Quicker/easier — larger sample readily Non-representative — variety of people
(e.g. uni students present at the university) obtained available is likely limited/similar → biased
sample (e.g. young, above-average education)
Volunteer Participants invited via advert/announcement; Relatively easy; participants likely Non-representative — people who respond
(self- those who reply become the sample committed (willing to return for repeat may share characteristics (free time, higher
selected) testing); useful for finding UNUSUAL education/motivation)
participants
Random Every person in the population has an EQUAL Likely to be representative — all types of In reality not everyone equally accessible
CHANCE of being chosen (numbered list + people equally likely to be chosen (incomplete list, or mainly one type selected)
random number generator, or names in a hat) — esp. important if sample is small
EXAM TIP — quotable example: Milgram used a VOLUNTEER sample — a newspaper advert offering $4 for "a study of memory" at
Yale University. Baron-Cohen et al.'s study used volunteer sampling to find people with autism spectrum disorder (needed for their
specific research question).
EXAM TIP: Individual differences psychology and developmental psychology are specifically devoted to studying differences between
people — so recognising limitations in a sampling technique used for these topics is especially important.
Definition Numerical results about the amount/quantity of a measure Descriptive, in-depth results indicating the quality of a characteristic (e.g. open-
(e.g. pulse rate, IQ score) question answers)
Strengths Typically objective; scales/questions often very reliable; easy Often valid — participants express themselves exactly, not limited to fixed
to analyse (central tendency & spread) and compare choices; unusual responses aren't lost to averaging
Weaknesses Data collection often limits responses → less valid if participant Often relatively subjective; findings may be invalid if interpretation is biased by
wants an unavailable answer the researcher; may not generalise if from only a few individuals
Mode The most frequent score(s) in a data set ONLY measure usable on data in separate/discrete/named categories. Can have 2+ modes if
scores are equally common. Doesn't consider score VALUES, so less informative than
median/mean.
Median The middle score of a RANKED (smallest→largest) Cannot be used on discrete/named categories — only numerical/linear scale data. Unaffected
data set. If two middle numbers, add together & by outliers (benefit) but doesn't take their VALUE into account (so less representative than
divide by 2. the mean).
Mean Add up the values of ALL scores, divide by the total Most informative — considers the VALUE of every score. Can be swayed by a small number of
number of scores (including zeros) extreme/outlying scores, but by taking them into account is MORE representative than
median/mode.
Range (Largest value − smallest value) + 1 +1 added because psychology scales measure GAPS between points, not the points themselves. Problem:
doesn't accurately reflect outliers — one extreme score can drastically change the range while barely
affecting the mean.
Standard A calculation of the average difference Bigger SD = more spread out; smaller SD = more clustered. Advantage over range: takes EVERY score
deviation (deviation) between EACH score and into account, so not distorted by outliers alone.
the mean
Graphs
Bar chart Used for data in SEPARATE/discrete categories; totals or GAPS between bars (categories not linearly related). x-axis = IV levels/categories,
averages plotted y-axis = DV/total.
Histogram Used for CONTINUOUS data — e.g. distribution of scores No gaps between bars (unless a category is empty, leaving a visible gap). x-axis
= DV score(s), y-axis = frequency.
Scatter Displays correlational data — a dot marks each participant's A "line of best fit" may be drawn. Strength described by 'r' value from +1 to −1.
graph scores on BOTH co-variables
IMPORTANT: You CANNOT draw a causal conclusion from a scatter graph — it only shows a relationship exists, not which (if either) co-
variable causes the other.
1.10 Ethical Considerations
Ethical issues: problems in research raising concerns about welfare of participants (or wider negative impact on society).
Ethical guidelines: pieces of advice guiding psychologists to consider the welfare of participants/society. Based on the British
Psychological Society (BPS) Code of Ethics and Conduct (2018) — key principle: MINIMISING HARM and MAXIMISING BENEFIT.
MEMORY TRICK: "Concerned People Won't Deceive, Care Properly, Debrief" = Consent, Protection from harm, Withdraw (right to),
Deception (avoid), Confidentiality, Privacy, Debrief.
Informed Knowing enough about a study to freely decide whether to participate. Must be freely given by a COMPETENT individual. Difficult for children,
consent people with mental health problems/learning difficulties, low literacy, or non-native speakers. Children under 16 need parent/guardian consent
AND their own (in a child-friendly way).
Protection No greater physical/psychological risk than participants would expect in everyday life. Study may cause psychological harm (embarrassment,
from harm self-doubt, stress) or physical harm (risky behaviours). Risk should be minimised via screening, experienced researchers, stopping if
unexpected risks arise.
Right to Participants can leave a study, and remove their data, at ANY time. Made clear at the START. Incentives can be offered but not taken away if
withdraw they leave. Researchers must not use authority to pressure continuation.
Lack of Participants should not be deliberately misinformed. If essential (e.g. to avoid demand characteristics), tell the real aim ASAP and allow
deception removal of results.
Confidentiality Data stored separately from names/personal info; never published unless individuals specifically agree. A NUMBER can identify participants
instead of a name (e.g. to pair scores in repeated measures).
Privacy Participants' emotional/physical space should not be invaded. In questionnaires/interviews, participants can ignore questions they don't want
to answer. In observations, people should only be watched where they'd expect to be on public display.
Debriefing Full explanation of aims/consequences given at the END, so participants leave in at least as positive a state as they arrived. NOT a substitute
for designing an ethical study in the first place.
IMPORTANT: Presumptive consent = when actual informed consent is impossible (e.g. naturalistic observations, field experiments),
a SIMILAR group (not the real participants) is asked if they'd find the study acceptable.
EXCEPTION to confidentiality: personally identifiable info CAN be shared if the participant gives informed consent for this, OR in
exceptional circumstances where safety/interests of the individual or others are at risk.
MEMORY TRICK: "Replace Sick Numbers, Protect Housing & Rewards" = Replacement, Species choice, Number (minimum),
Procedures, Pain/suffering/distress, Housing, Reward/deprivation/aversive stimuli.
Replacement Consider replacing animal experiments with alternatives — videos of previous studies, computer simulations
Species Choose the species LEAST likely to suffer pain/distress; consider whether bred in captivity, previous experimentation experience, and
sentience (ability to think/feel)
Number of Only the MINIMUM number needed for valid/reliable results; minimised via pilot studies, reliable DV measures, good design, appropriate
animals data analysis
Procedures Controlled by legal requirements/organisation guidelines; animal's experience should be as normal/positive as possible within research
constraints
Pain, suffering, Should be avoided where possible; use designs that IMPROVE rather than worsen the animal's experience (e.g. studying early enrichment
distress vs normal, rather than early deprivation); costs justified by scientific benefit
Housing Depends on the species' social behaviour (isolation worse for social animals, overcrowding causes distress/aggression); enough space,
food, water; artificial environment only needs to recreate aspects important to welfare/survival (not what looks nice to humans)
Reward, Consider each species' natural feeding/drinking needs when using deprivation; prefer using PREFERRED food as a reward instead of
deprivation, deprivation; avoid aversive (unpleasant) stimuli where possible
aversive stimuli
EXAM TIP: Bateson's Cube (1986) = 3-dimension model deciding if animal research is justified: certainty of (medical) BENEFIT
(high=good), quality of RESEARCH (high=good), animal SUFFERING (low=good). Research only justified if certainty of benefit is HIGH
and suffering is LOW.
EXAM TIP — justifying an ethical decision, explain 3 points: (1) choices about animals & their care (species, housing,
procedures); (2) WHY the study is being done (benefits to humans/animals); (3) STRENGTHS of the design (controls, objectivity, validity,
reliability).
MEMORY TRICK — THE BIG 4, always in this order: Reliable? → Valid? → Generalisable? → Replicable?
Reliability The extent to which a procedure/measure is consistent — produces the same results with the same Would I get the SAME result if I
people on each occasion repeated this?
Validity The extent to which the researcher is testing what they claim to be testing. Affected by reliability, Am I actually measuring what I
objectivity/subjectivity, face validity, demand characteristics, ecological validity. CLAIM to measure?
Generalisability How widely findings apply — depends on ecological validity AND the sample's representativeness/size Do the findings apply to OTHER
people/places?
Replicability The extent to which a study's procedure can be kept the same when repeated (by same or different Could someone else repeat this
researchers) to verify results EXACTLY using my write-up?
DON'T CONFUSE: Inter-rater reliability = 2+ researchers interpreting QUESTIONNAIRE/INTERVIEW answers consistently. Inter-
observer reliability = 2+ researchers watching the SAME behaviour/event and recording it consistently.
EXAM TIP — Test-retest: checks RELIABILITY. Use the SAME test TWICE, in the same situation, on the same people. A HIGH
CORRELATION between the two sets of scores = high reliability.
EXAM TIP — Face validity vs Ecological validity: Face validity = does the TEST/TASK itself seem to test what it claims to? (e.g. a
helping test using spiders would lack face validity for people who are scared of spiders). Ecological validity = do findings from this
SITUATION apply to real life/other situations? Improved by mundane realism (a realistic task).
Design Independent measures = fewer order-effect problems; repeated measures = fewer individual-difference problems
Sample Opportunity = larger sample; volunteer = specific/rare participants; random = better generalisability
EXAM TIP: Always tie your suggested improvement to the SPECIFIC scenario given in the question — generic textbook answers score
fewer marks than answers that clearly apply to the study described.
QUICK REVISION CHEAT SHEET
Experiment basics IV manipulated → DV measured → looking for cause & effect. Control confounding variables to raise validity.
3 Experimental designs "I R Match" — Independent (diff people, no order effects, participant variable risk) / Repeated (same people, order effect
risk, no participant variable risk) / Matched pairs (paired, best of both, hard to match)
Order effects fix Randomisation OR Counterbalancing (ABBA — half A-then-B, half B-then-A)
Lab vs Field Lab = more control, causal certainty, but low ecological validity | Field = natural behaviour, but harder to control &
consent issues
Hypothesis types Non-directional = effect, no direction | Directional = states direction (needs evidence) | Null = due to chance
Closed vs Open Qs Closed = quantitative (countable, fixed options) | Open = qualitative (own words, needs interpreting)
3 Interview types "SUS" — Structured (fixed/scripted) / Unstructured (flows from answers) / Semi-structured (fixed + extra Qs, best
compromise)
Case study 1 instance, many methods+sources, triangulation = validity check, but n=1 = very poor generalisability, risk of
researcher subjectivity
Correlation golden rule NEVER claim causation — only a relationship; classic example = ice cream & murder rate, 3rd variable = hot weather
Positive vs Negative correlation Positive = both increase together | Negative = one up, one down
Longitudinal vs Cross-sectional Longitudinal = same cohort over time (risk: sample attrition) | Cross-sectional = different groups, one time point (risk:
cohort/generation differences)
Terman's Termites 1922 gifted-children longitudinal study; debunked "high IQ = automatic success"; massive sample attrition over decades
Participant vs Situational variable Participant = about the PERSON (age, personality) | Situational = about the ENVIRONMENT (noise, light)
3 Sampling techniques "O-V-R" — Opportunity (available, biased/limited) / Volunteer (advert response, biased/motivated) / Random (equal
chance, best generalisability but list may be incomplete)
Milgram's sample Volunteer sample — $4 newspaper ad, Yale University, "a study of memory"
3 M's (central tendency) Mode = most frequent (only one usable on named categories) | Median = middle score, ignores outlier values | Mean =
average of all scores, most informative but swayed by outliers
2 Spread measures Range = (highest−lowest)+1, simple but hides outlier detail | Standard deviation = avg distance from mean, uses every
score
3 Graphs Bar chart = gaps, discrete categories/totals | Histogram = no gaps, continuous data | Scatter graph = correlation dots +
best-fit line, r value +1 to −1
7 Human ethics guidelines "Concerned People Won't Deceive, Care Properly, Debrief" — Consent, Protection from harm, Withdraw (right to),
Deception (avoid), Confidentiality, Privacy, Debrief
Presumptive consent Ask a SIMILAR group (not real participants) — used when real consent is impossible (naturalistic obs, field experiments)
7 Animal ethics guidelines "Replace Sick Numbers, Protect Housing & Rewards" — Replacement, Species, Number (min), Procedures,
Pain/suffering/distress, Housing, Reward/deprivation/aversive stimuli
Bateson's Cube Benefit certainty (high) + Research quality (high) + Animal suffering (low) = ethically justified
The Big 4 (evaluation) Reliable? → Valid? → Generalisable? → Replicable? — always in this order
Test–retest Same test, twice, same people, same conditions — high correlation between scores = high reliability
Face validity vs Ecological validity Face = does the TASK look like it tests the right thing? | Ecological = do findings apply to REAL LIFE?
Improving methodology Method / Design / Sample / Tool / Procedure — always tie the fix to the SPECIFIC scenario given