Chapter 1 Research Methods Notes
Chapter 1 Research Methods Notes
Written so that someone learning this for the very first time can understand it —
every idea explained in plain English before the exam-ready detail.
Contents
1. 1.1 Experiments
4. 1.4 Observations
5. 1.5 Correlations
IN PLAIN ENGLISH
An experiment is the most "scientific" method psychologists use. The researcher deliberately changes ONE thing (to see
what effect it has) and measures what happens as a result. Think of it like testing whether fertiliser helps plants grow: you
give one set of plants fertiliser and not the other, then measure how tall they grow. The thing you change (fertiliser or not) is
the independent variable, and the thing you measure (plant height) is the dependent variable.
Before running any study, a psychologist needs to know exactly what they are trying to find out.
Aim — the purpose of the study; what the researcher is trying to find out or the question they want answered. Example: "to
investigate whether mind maps or revision apps are more effective at helping students to learn."
Once there's an aim, the researcher writes a hypothesis — a precise, testable prediction.
Hypothesis — a testable statement based on the aim of the study. It must be falsifiable, meaning it must be possible to
prove it wrong. (Sigmund Freud's ideas about unconscious motives were criticised because they could never be proven
wrong — whatever happened, he could claim it fit his theory. That's a classic example of a non-falsifiable idea.)
Alternative hypothesis — predicts there WILL be a difference between conditions (or a relationship, in a correlation).
Non-directional (two-tailed) hypothesis — predicts a difference/effect, but does NOT say which direction it will go.
Used when there's no previous research to base a direction on. Example: "There is a difference between the effectiveness
of mind maps and revision apps."
Directional (one-tailed) hypothesis — predicts the direction of the effect. Used when previous evidence points that
way. Example: "Students using revision apps will learn better than students using mind maps."
Null hypothesis — states that any difference found is simply due to chance, not a real effect. Example: "Any difference
in effectiveness between mind maps and revision apps is due to chance."
Hypotheses must be operationalised — meaning the vague ideas are turned into something measurable. Saying "revision apps
help students learn better" isn't good enough, because we don't know how "better" will be measured. A properly operationalised
version would be: "Students using the Gojimo revision app will gain higher test marks than students using mind maps."
Independent Variable (IV) — the thing the researcher deliberately changes or manipulates to create different conditions.
Dependent Variable (DV) — the thing the researcher measures, which is expected to change because of the IV.
Control condition — a baseline group/condition that the experimental condition is compared against.
Types of Experiment
Not all experiments happen in a lab! The two main types are:
Laboratory Conducted in an artificial, controlled setting (e.g. a university lab). Participants usually know they're being
experiment tested.
Field experiment Conducted in participants' everyday, natural environment (e.g. their school or workplace), but the researcher
still manipulates the IV.
Strengths Good control of variables → raises validity. Causal relationships can Natural setting → participants behave naturally →
be shown (only the IV should affect the DV). Standardised more representative results (higher ecological
procedures → raises reliability and allow replication. validity). If unaware they're in a study, fewer
demand characteristics.
Weaknesses Artificial situation → participants may behave unnaturally → lower Harder to control variables → lower reliability,
ecological validity. Participants may notice they're being tested → harder to replicate. Less certain the IV really
demand characteristics. Participants aware they're in a study caused the change in DV. If participants unaware,
raises ethical issues about consent. they can't give informed consent.
WORKED EXAMPLE
A team wants to know if watching TV affects how kind (pro-social) children are. In the lab experiment, each child watches
either a "helpful" or "neutral" cartoon in a research room, then is observed with a crying doll. In the field experiment,
parents show their own child the cartoon at home, and the child is later videoed playing with the doll at home. The field
version has higher ecological validity (child is in their normal environment) but is harder to control (distractions at home).
Experimental Designs
IN PLAIN ENGLISH
"Design" here just means: how do we decide which participants do which condition? Do the same people do both conditions,
or different people do each one?
Independent Different participants are used in each condition No order effects (each person Participant variables —
measures (level of the IV). only does it once). Less differences between the groups of
demand characteristics (they people (age, ability, etc.) might
only see one condition). cause the result, not the IV.
Repeated The SAME participants take part in every Individual differences can't Order effects (practice/fatigue).
measures condition (they "repeat" their performance). bias results, because each Participants see both conditions,
person acts as their own so more likely to guess the aim
comparison. Needs fewer (demand characteristics).
participants.
Matched Different participants are used, but each person Reduces participant variables Time-consuming/hard to find good
pairs in one condition is matched with someone similar without causing order effects. matches. Matching may not be
(e.g. same age/ability) in the other condition. perfect — matched people can
Identical twins are ideal matched pairs. still differ in unmeasured ways.
Because participants do every condition in a repeated measures design, we get order effects:
Practice effect — participants get BETTER simply from doing the task before (familiarity/learning).
Fatigue effect — participants get WORSE from doing the task before (tiredness or boredom).
Two ways to fix this:
Randomisation — randomly decide which order each participant does the conditions in.
Counterbalancing — split participants into two halves: half do condition A then B; half do B then A (called an "ABBA"
design). This cancels out order effects more reliably than simple randomisation.
For independent measures, the equivalent fix for participant differences is random allocation — randomly assigning
participants to conditions (e.g. drawing numbers from a hat) so any individual differences are spread evenly across both
groups.
Participant variable — a type of confounding variable caused by differences BETWEEN people (age, gender, personality,
intelligence).
Situational variable — a type of confounding variable caused by the ENVIRONMENT (e.g. one room being noisier or
brighter than another).
Demand characteristics — clues in the study that give away its aim, causing participants to (often unconsciously) change
their behaviour — this reduces validity.
WORKED EXAMPLE
Dr Huang compares students on soft chairs vs hard chairs to see which group works harder. If, by accident, all the "soft
chair" students happen to also be arts students (who might naturally work differently to science students), then "subject
studied" becomes a confounding participant variable — we can no longer be sure it was the chairs that caused any difference
in work rate.
Ethics in Experiments
We'll cover ethics fully in Section 1.10, but here's how it applies specifically to experiments:
In lab experiments, it's usually easy to get informed consent because participants know they're being tested — but
sometimes deception is needed so participants don't guess the aim (which would create demand characteristics).
In field experiments, it's often hard or impossible to get consent, because participants may not even know they're in a study.
This raises issues around their right to withdraw too — you can't withdraw from something you don't know you're part of!
Privacy is easier to protect in a lab (the tasks are pre-planned); in a field setting there's more risk of accidentally invading
someone's personal space.
Confidentiality (keeping data secure and anonymous) matters in every experiment — but in a field study it's especially
important participants can't be identified by things like their workplace.
IN PLAIN ENGLISH
A self-report is simply when you ask someone directly to tell you about themselves — their opinions, feelings, or behaviour —
rather than watching them or testing them. There are two main ways to do this: questionnaires (written questions) and
interviews (spoken questions).
Self-report — a research method (questionnaire or interview) that collects data by asking participants to give information
about themselves.
Questionnaires
Closed questions — offer a fixed set of possible answers (e.g. yes/no, a rating scale, or a Likert scale like "strongly agree"
to "strongly disagree"). These produce quantitative (numerical) data.
Open questions — ask participants to answer in their own words, with no fixed choices (e.g. "Why do you think helping
behaviour is important?"). These produce qualitative (descriptive) data.
"What is your gender: boy or girl?" "Why do you believe it's important to help people with phobias?"
"How do you travel to school? walk/bicycle/bus/train/car" "Describe your views on social media and helping behaviour."
"Rate how much you like psychology from 0–4"
Evaluating questionnaires
Closed questions: easy to analyse (just totals/averages), but limited in depth — a participant's true feeling might not fit any
of the offered options, which can lower validity.
Open questions: rich, detailed, valid data — BUT harder to score consistently. Different researchers might interpret the same
answer differently. This is a lack of inter-rater reliability.
Low response rate: people can easily ignore a questionnaire, so the people who DO reply might all share certain traits (e.g.
more free time), making the sample less representative — poor generalisability.
Social desirability bias: people often answer in a way that makes them look good, rather than being fully honest.
Filler questions — extra irrelevant questions added to disguise the study's real purpose, so participants are less likely to
guess the aim and change their answers.
Inter-rater reliability — how consistently two or more researchers interpret and score the same qualitative answers.
Social desirability bias — answering in the way that seems most acceptable to others, rather than answering completely
honestly.
Generalisability — how widely a study's findings apply to other people, places or times.
Filler questions — irrelevant items placed among the real questions to hide the true aim of a study.
Interviews
Interview — a self-report method using spoken questions, usually face-to-face or by telephone (though real-time chat is also
possible).
Structured interview — same fixed questions, in the same fixed order, for every participant. May even standardise the
interviewer's tone/posture. Easy to compare between participants.
Unstructured interview — questions depend on what the participant says (very flexible), so questions may differ
completely between participants. Hard to compare answers.
Semi-structured interview — a mix: some fixed questions (so answers CAN be compared) plus the freedom to ask extra
follow-up questions specific to that person.
Evaluating interviews
Participants may lie — either from social desirability bias, or because they've guessed the aim and want to "help" (or
deliberately disrupt) the study.
Time-consuming, which may put off certain types of people from volunteering — reducing how representative the sample is.
Interpreting answers risks subjectivity (the researcher's personal views affecting how they interpret the data) — the goal is
objectivity (an unbiased, external viewpoint), which can be improved by having other, unaware researchers help interpret
the data.
Closed questions in an interview → quantitative data → easier to analyse, generally more reliable.
Open questions in an interview → qualitative data → more in-depth and potentially more valid (participant isn't limited to fixed
choices), but interpretation may be less reliable.
Subjectivity — when a researcher's personal feelings or beliefs bias how they interpret data.
Objectivity — interpreting data in a way that isn't affected by personal feelings, beliefs or experience — consistent no
matter who interprets it.
IN PLAIN ENGLISH
Instead of studying lots of people briefly, a case study is when researchers study ONE person (or sometimes one family, or
one organisation) in huge detail, often using several different methods together. It's a bit like a detective building a complete
picture of one individual rather than taking a quick snapshot of many people.
Case study — a research method studying a single instance (usually one person, but possibly a family or institution) in great
detail.
Case studies typically combine several techniques — interviews, observations, tests, and questionnaires — and may gather
information from multiple sources: the person themselves, their relatives, colleagues, and existing records (e.g. medical or school
records).
Studying rare cases where a detailed description is valuable (e.g. an unusual brain injury).
Tracking developmental changes over time, such as a child's progress or a patient's recovery.
Triangulation — using several different techniques (e.g. observation + interview + questionnaire) to study the same thing.
If they all produce similar results, this suggests the findings are valid.
Strengths Weaknesses
Very high validity — the person is studied in real Subjectivity risk: researchers often build a close relationship with the
depth and in a genuine, real-life context. participant, which can bias their interpretation, lowering validity.
Triangulation (using multiple methods) can further
support validity.
Ethical risk: such personal, detailed questions can feel intrusive; participants
may feel unable to refuse to answer. It can also be hard to keep the person's
identity confidential — even using initials may not be enough if they are well
known.
Reliability risk: usually only one (or very few) researchers and one participant
are involved, so it's hard to be sure the interpretation is objective — a different
researcher might interpret it differently.
Low generalisability: findings are specific to that one person, so they may not
apply to anyone else at all.
IN PLAIN ENGLISH
An observation is exactly what it sounds like: watching people (or animals) and recording what they do, rather than asking
them questions or testing them. There are several choices to make about HOW you observe — where, how much detail,
whether you join in, and whether people know they're being watched. Each choice changes the study's strengths and
weaknesses.
Naturalistic observation — watching participants in their normal, everyday environment, with no interference from the
researcher (social or physical).
Controlled observation — the social or physical environment has been deliberately manipulated by the researcher (e.g.
changing group size, adding objects). Can happen in a natural setting OR an artificial one like a lab.
Unstructured observation — the observer records the WHOLE range of behaviours they see. Usually just used at the start
(a "pilot" stage) to help decide what specific behaviours matter.
Structured observation — the observer records only a limited, pre-decided set of behaviours, called behavioural
categories.
Behavioural categories — the specific, clearly defined ("operationalised") actions being recorded. They must break up a
continuous stream of behaviour into distinct, observable events (not "inferred" mental states you can't actually see).
Participant observer — the researcher becomes part of the social group being studied (e.g. joining in conversation or
play).
Non-participant observer — the researcher stays separate/apart from the group being studied (e.g. watching through one-
way glass).
Overt observer — participants know the researcher is observing them (role is obvious).
Covert observer — participants do NOT know the researcher is observing them (hidden or disguised, e.g. watching via
CCTV, or disguised as a group member).
Inter-observer reliability — how consistently two or more observers record the same event.
Evaluating observations
Naturalistic Behaviour is true-to-life (high ecological validity) No guarantee the target behaviour will even happen
Controlled Ensures the behaviour of interest actually occurs Less natural — may reduce ecological validity
Unstructured Captures unexpected/important behaviours you didn't Hard to record everything accurately → can lower reliability
predict
Structured More reliable — observer only focuses on a small set of Might miss behaviours that aren't on the list
categories
Covert Higher validity — no demand characteristics, since Ethical issue (no informed consent); practically harder to
participants don't know they're watched arrange (observer must stay hidden)
Overt Easier and more ethical to arrange Participants may change behaviour because they know
they're being watched — lowers validity
REMEMBER
Observation isn't just a stand-alone method — it can also be used as a technique inside other methods. For example, in an
experiment the DV might be measured by observing behaviour; in a case study, observation might be one of several
techniques used (alongside interviews etc).
IN PLAIN ENGLISH
A correlation looks at whether two things that CAN'T (or shouldn't) be experimentally changed tend to rise and fall together.
Unlike an experiment, nobody manipulates anything — you simply measure both variables as they naturally occur and see if
there's a pattern. Crucially, a correlation can NEVER prove that one thing causes the other — only an experiment can do that.
Correlation — a research method that looks for a relationship between two measured variables (co-variables). A change in
one is related to a change in the other, but this is NOT necessarily a causal relationship.
Causal relationship — when a change in one variable is directly RESPONSIBLE for (causes) a change in another — this can
only be established with an experiment, never with a correlation alone.
Correlations are useful when it isn't practical or ethical to manipulate a variable directly — e.g. we can't ethically force children to
watch years of violent TV, but we CAN measure how much violent TV they already watch and correlate it with their aggression
levels.
Positive correlation — as one variable increases, the other ALSO increases (they rise together).
Negative correlation — as one variable increases, the other DECREASES.
No correlation — no consistent pattern between the variables at all.
Evaluating correlations
Validity depends on both co-variables being clearly and effectively measured.
Reliability depends on how consistent the measurements are — scientific scales (e.g. time in seconds) are very reliable; self-
report or observation-based measures are often less reliable.
The single biggest thing to remember: you cannot conclude cause and effect from a correlation — a hidden "third
variable" could be responsible for both changes.
Correlations are great for a first look at whether something is worth investigating further with an experiment.
IN PLAIN ENGLISH
A longitudinal study follows the SAME group of people over a long period of time — months, years, or even decades —
retesting them again and again to see how they change. It's the opposite of quickly comparing different age groups all at
once (which is called a "cross-sectional" study).
Longitudinal design — the same participants are tested on 2 or more occasions over a long period of time (e.g. before and
after a six-month intervention, or repeatedly across many years).
Cohort — a group of participants all selected at the same age or stage of life, who are then tracked over time.
Cross-sectional study — instead of following one group over time, this compares DIFFERENT groups of people of different
ages/stages, all measured at the SAME point in time.
REAL EXAMPLE
The Terman "Life Cycle Study of Children with High Ability" began in 1922, following highly intelligent children for decades. It
found that high IQ doesn't guarantee success in life — but did produce famous graduates such as psychologist Lee Cronbach.
KEY STRENGTH
Because the SAME people are re-tested, researchers can be confident that any change is due to the passage of
time/development — not due to differences between separate groups of people (as could happen in a cross-sectional study).
It avoids confounding situational variables (e.g. one generation having different schooling to another).
Sample attrition — the loss of participants from a study over time (they may withdraw consent, get bored of repeated
testing, move away and become uncontactable, become unwell, or even die).
Sample attrition means the sample shrinks over time and becomes biased towards stable, healthy, cooperative people —
reducing generalisability.
Being part of a famous long-term study can make people feel "special," which may itself change their behaviour (a validity
problem) — e.g. Terman's participants were nicknamed "Termites" and knew they were considered gifted.
Reliability risk: measuring tools may need updating over such a long period, and the researchers running the study may
change over time.
Ethical issues: consent must be repeatedly re-confirmed at every time point (harder with children); keeping large banks of
contact details over years raises confidentiality concerns; harder to judge if a child wants to withdraw.
IN PLAIN ENGLISH
This section is really about being PRECISE. Psychologists can't just say "we tested if comfy chairs help people work" — they
need to define exactly what "comfy" means and exactly what "work better" means, and they need to make sure nothing
ELSE is sneaking in to affect the results. This section covers how to define variables clearly and how to stop unwanted
variables from messing up a study.
Operationalisation
Operationalisation — turning a vague idea into something clearly defined and measurable. E.g. instead of "young people,"
say "under 20 years old."
WORKED EXAMPLE
"Hard chairs" vs "soft chairs" isn't precise enough. Better: "chairs with wooden/plastic seats" vs "chairs with padded seats."
Likewise, "students working better" isn't precise. Better: "the number of pieces of homework handed in on time" or "time
spent doing extra work."
Controlling Variables
IN PLAIN ENGLISH
If we don't control other variables, we can never be sure that our IV was really what caused the change in the DV —
something else might have snuck in and caused it instead.
Confounding variable — an uncontrolled variable that acts systematically on just ONE level of the IV, so it can hide OR
exaggerate the true effect of the IV, making the results hard to interpret.
Participant variable — a confounding variable caused by differences between the people themselves (e.g. their natural
ability, personality, mood).
Situational variable — a confounding variable caused by an aspect of the environment/setting (e.g. lighting, noise,
temperature, weather).
Pilot study — a small trial run of the study's procedures BEFORE the real study begins. Its purpose is to spot problems (like
uncontrolled variables) early, so they can be fixed.
Standardisation
Standardisation — making sure every participant has exactly the same experience (same instructions, same procedure,
same equipment), no matter which condition they are in.
Standardised instructions — the exact same written or spoken information given to every participant, so their experience
is as similar as possible.
IN PLAIN ENGLISH
You can't test every single person in the world, so researchers pick a smaller group to represent everyone else. The BIG
question is: how do you pick that smaller group so it's actually a fair representation of everybody? Different methods of
picking people (sampling techniques) have different strengths and weaknesses.
Population — the whole group of people (or animals) who share a certain characteristic (e.g. all people in a country, all
football fans).
Sample — the smaller group of people actually selected to take part in the study, ideally representative of the population.
Sampling technique — the method used to choose the sample from the population.
Opportunity Choosing whoever is conveniently Quick and easy — bigger samples can be Likely unrepresentative —
sampling available (e.g. asking students who gathered fast available people tend to be alike
happen to be at the university) (e.g. all young, similarly
educated)
Volunteer Putting out an advert/announcement; Easy to arrange; volunteers tend to be Volunteers often share traits
(self- people who respond become the sample committed (e.g. willing to come back for (e.g. more free time, more
selected) retesting); great for finding rare/unusual curious personality) —
sampling participants unrepresentative
Random Every person in the population has an Most likely to be genuinely Time-consuming to arrange
sampling EQUAL chance of being picked (e.g. representative properly; if the original list of
numbers drawn from a hat, or a random people is incomplete, it isn't
number generator) truly random after all
REMEMBER
The size of the sample matters too — smaller samples are less likely to capture the full range of variation that exists in the
population, so they tend to be less reliable and less representative.
IN PLAIN ENGLISH
Once a study is done, researchers end up with a big pile of numbers or descriptions — this is called "raw data." On its own,
raw data is hard to make sense of. This section is about the tools used to simplify, summarise, and display that data so we
can actually understand what it's showing us.
Types of Data
Quantitative data — numerical data about the AMOUNT of something (e.g. pulse rate, a test score).
Qualitative data — descriptive, in-depth data about the QUALITY of something (e.g. detailed answers to open questions).
Strengths: Usually objective; scales/questions tend to be highly Strengths: Often more valid — participants can express
reliable; can use averages and spread to compare easily; unusual (but themselves fully instead of being squeezed into fixed choices.
important) responses aren't hidden by averaging out.
Weaknesses: Data collection method may limit what participants can Weaknesses: More subjective — recording or interpretation
express, making it less valid if their true view doesn't fit the options may be biased by the researcher's own opinions; results from
given. a few individuals may not generalise widely.
The first step is usually to build a clear summary table — with a title, and clearly labelled rows and columns (units like seconds
or cm should be stated once in the heading, not repeated in every cell).
IN PLAIN ENGLISH
An "average" is one single number that tries to represent the WHOLE data set. There are three different kinds of average,
and each one is useful in different situations.
Mode The most frequently occurring score (there can be Any type of data, Only average usable with categories, but
more than one mode if scores are tied) including categories ignores the actual VALUES of scores, so it's
(e.g. favourite subject) the least informative
Median The middle value once all scores are lined up Numerical/linear scale Not distorted by extreme outlier scores,
smallest → largest (average the middle two if data only (not but ignores the actual values of most
there's an even number of scores) categories) scores
Mean Add up ALL the scores, then divide by how many Numerical/linear scale Most informative (uses every score's actual
scores there are (the everyday "average") data only value), but CAN be distorted by one or two
extreme outliers
IN PLAIN ENGLISH
An average tells you the "typical" score — but it doesn't tell you whether everyone scored close to that typical value, or
whether scores were all over the place. That's what a "measure of spread" tells you.
Range — (biggest score − smallest score) + 1. We add 1 because psychological scales measure the gaps BETWEEN points,
not the points themselves.
Standard deviation (SD) — a calculation of the average distance between each individual score and the mean. A BIGGER
standard deviation means scores are more spread out/varied; a SMALLER one means scores are clustered close to the mean.
WHY STANDARD DEVIATION IS BETTER THAN RANGE
The range only looks at the two most extreme scores, so ONE unusual outlier can distort it completely. Standard deviation
takes EVERY score into account, so it gives a much fairer, more complete picture of how spread out the data really is.
Bar chart Data in separate, distinct categories (e.g. totals, Bars have GAPS between them — because the categories aren't
means or modes for different groups) part of one continuous scale
Histogram Continuous data (e.g. a whole distribution of test Bars TOUCH each other — because the x-axis is one continuous
scores) scale (DV on x-axis, frequency on y-axis)
Scatter Correlational data (two co-variables) Each dot = one participant's score on BOTH variables. A "line of
graph best fit" may be added.
r value — a number from −1 to +1 describing the strength of a correlation. Closer to +1 or −1 = a stronger correlation
(points close to the line of best fit). Closer to 0 = a weaker or non-existent correlation (points scattered randomly).
ALWAYS REMEMBER
A scatter graph can show that a relationship EXISTS — but it can never tell you WHICH variable (if either) is causing the
change in the other.
IN PLAIN ENGLISH
Ethics is all about making sure research doesn't harm the people (or animals) taking part, and that it's conducted fairly and
honestly. Psychology has clear guidelines to protect participants' wellbeing AND to protect the reputation and
trustworthiness of psychology as a whole.
Ethical issues — problems in research that raise concerns about participant welfare (or a wider negative impact on society).
Ethical guidelines — official advice (such as the British Psychological Society's Code of Ethics and Conduct, 2018) that
helps psychologists protect participants and the public standing of psychology. University research usually also needs
approval from an ethics committee before it can begin.
Informed consent — participants are given enough information about the study to decide, freely and while fully
understanding, whether they want to take part.
Presumptive consent — used when it isn't possible to get informed consent from the real participants (e.g. in natural
observations or field experiments). A similar group of people is told about the proposed study in advance and asked whether
THEY would be happy to take part — if they say yes, this "presumes" the real participants would agree too.
Extra care is needed with groups who may find it harder to give fully informed consent: children, people with mental health
conditions or learning difficulties, non-native speakers, or people who might feel pressured (like prisoners).
Right to Withdraw
Right to withdraw — participants must be able to leave a study (and have their data removed) at ANY time. Any
rewards/incentives offered can't be taken away if they choose to leave, and researchers must never use their authority to
pressure someone to stay.
Protection from harm — participants should never face greater physical or psychological risk than they would in ordinary
daily life. Researchers should screen participants for risk, use experienced staff, and stop the study immediately if
unexpected risks appear.
Deception
Deception — deliberately misinforming (lying to) participants about the study. Should be avoided wherever possible; if truly
necessary (e.g. to stop participants guessing the aim), the study should be planned to minimise distress, and participants
must be fully debriefed afterwards.
Confidentiality
Confidentiality — participants' data and personal information must be stored safely and never released to anyone outside
the study. Names are usually replaced with numbers or initials. Institutions like schools or hospitals should also be kept
anonymous where possible.
Privacy
Privacy — a participant's physical space and emotions should never be invaded. They should be free to decline answering
any question, and should only be observed in places where they'd expect to be seen by others anyway.
Debriefing
Debriefing — a full explanation given to participants AFTER the study, covering its real aims and any consequences, so they
leave in at least as positive a state as when they arrived. Debriefing is NOT a substitute for designing an ethical study in the
first place — it can't undo real harm.
Animals are used in psychology because they can be convenient models for studying processes (like learning), allow procedures
that couldn't ethically be done on humans (like brain surgery), or are simply interesting to study in their own right (e.g.
communication in whales).
Bateson's cube (1986) — a model for deciding whether animal research is justified, based on three factors: how CERTAIN
the benefit is, how HIGH the QUALITY of the research is, and how much SUFFERING the animals experience. Research is
considered justified when certainty and quality are high, and suffering is low.
Human guidelines: informed consent, right to withdraw, protection from harm, avoiding deception, confidentiality, privacy,
debriefing.
Animal guidelines centre on Bateson's cube: weighing certainty of benefit and quality of research against the level of
suffering caused.
Debriefing helps reduce harm afterwards, but doesn't replace designing an ethical study from the start.
1.11Evaluating Research: Methodological Issues
IN PLAIN ENGLISH
This section pulls together ideas you've already met throughout the chapter — reliability, validity, generalisability,
replicability — and asks you to use them to judge whether a piece of research is actually "good science."
Reliability
Test-retest reliability — giving the same test to the same participants on two separate occasions; if reliability is high, the
two sets of scores should correlate strongly.
Inter-rater reliability — how consistently different researchers interpret the same qualitative data (e.g. answers to open
questions).
Inter-observer reliability — how consistently different observers record the same event in an observation.
Reliability can be improved through standardisation (same instructions/procedures/materials for everyone), clear operational
definitions, and researchers discussing/training together to interpret data consistently.
Validity
Validity — whether a test, task or study actually measures what it CLAIMS to measure.
Face validity — whether a test simply LOOKS, on the surface, like it measures what it's supposed to.
Ecological validity — whether findings from one situation (e.g. an artificial lab) would also apply to other, more realistic
real-world situations.
Mundane realism — how true-to-everyday-life a task feels. Higher mundane realism generally leads to higher ecological
validity.
Replicability — whether a study's exact procedure can be repeated by the same or different researchers, to check if the
same results occur again. This requires very detailed, clear reporting of the original method.
Generalisability — how widely a study's findings apply beyond the specific sample studied — to other people, places, or
times. Depends heavily on how representative and large the sample was, and on ecological validity.
Whenever you're asked to evaluate ANY study (one you know, or a brand new example), ask:
1. Is it valid? Does it test what it claims to? Think about ecological validity, demand characteristics, and subjectivity.
2. Is it reliable? Are the measures/tools consistent? Could interpretation of the data be subjective?
3. Is it generalisable? Would the findings apply to other people/places/times? Depends on the sample and ecological
validity.
4. Could it be replicated? Is there enough procedural detail reported to repeat it exactly?
Ways to improve a study's methodology
Method — e.g. switch from a field to a laboratory experiment (or vice versa), or from a questionnaire to an interview.
Design — independent measures avoids order effects; repeated measures avoids individual-difference problems.
Sample — opportunity sampling for a bigger sample; volunteer sampling to find rare participants; random sampling for better
generalisability.
Tool — check and improve inter-rater or test-retest reliability.
Procedure — reduce demand characteristics, and increase the realism (mundane realism) of the task.