0% found this document useful (0 votes)
21 views67 pages

Understanding Research Design in Psychology

The document discusses various aspects of psychological research, including the importance of understanding demand characteristics, research designs (P*E and ATI), and the goals of research such as description, explanation, prediction, and application. It also covers research ethics, including the APA Ethical Principles, informed consent, and the treatment of human and animal subjects. Additionally, it outlines the process of conducting research, from literature review to data analysis, emphasizing the significance of reliable and valid measurements.

Uploaded by

a17538856z
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
21 views67 pages

Understanding Research Design in Psychology

The document discusses various aspects of psychological research, including the importance of understanding demand characteristics, research designs (P*E and ATI), and the goals of research such as description, explanation, prediction, and application. It also covers research ethics, including the APA Ethical Principles, informed consent, and the treatment of human and animal subjects. Additionally, it outlines the process of conducting research, from literature review to data analysis, emphasizing the significance of reliable and valid measurements.

Uploaded by

a17538856z
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Demand characteristics: if subjects know what is expected of them, they might be “good

subjects” and not behave naturally.

Distinguish between P*E and ATI design:


A P × E design has at least one subject factor (P) and one manipulated (E) factor; an ATI design is a
type of P × E design in which the “P” factor refers to some kind of ability or aptitude; these ATI
designs are frequently seen in educational research.

Some have argued that results from single‐subject designs do not generalize beyond the specific
situation of the study. Defenders respond that there have been studies that directly test for
generalization (e.g., the study on stuttering).

Lecture 1: Scientific Thinking p1-7


P1
Why take this course?
-Foundation for understanding other psychology courses:
Process course (how to conduct experiment, process/manner of gain knowledge), content course
Example: 2 studies with Single factor (with two levels) between-subjects design
The influence of tactile experiences of physical warmth on judgement of other people
The impact of presentation format on visual working memory (simultaneous, sequential)
Multiple factor design/factorial design, condition=level, same group: with-in subject design
-Essential for graduate school
P2
-Make you a more informed and critical thinker: judge quality of evidence to support a claim
Fair, unbiased when examine conflicting claims, draw reasonable conclusions based on evidence.
Confirmation bias: A tendency to seek and pay special attention to information that supports
one’s beliefs, while ignoring information that contradicts a belief.
E.g. take amulet with you...

The Goals of Research in Psychology


-Describe: identify regularly occurring sequences of psychological events, e.g. span of STM
most survey/questionnaire and observational research-->descriptive
without it, predictions cannot be made and explanations are meaningless.
-Explain: what caused it to happen
(in terms of their relationship to other factors, casual are ideal e.g. memory expert? chunking)
P3
-Predict: psychological events follow certain laws, regular, predictable
If that series of events happened often enough.. correlation/regression
E.g. if chunk-->easier to remember
-Apply: real-world applications of psychological events...translational research

Ways of knowing
-Authority: stability, consistency, beneficial if knowledge is brand new e.g. coronavirus
-Reason/logical argument: initial assumptions may be incorrect
P4
-Experience: observation, experience, limited experience, bias interpretation based on social
cognition biases:
confirmation bias, availability heuristic (mental shortcut, ease examples/instance come to mind,
unusual, memorable, overestimate, e.g. plane, MCQ)
Availability heuristic --> confirmation bias...
P5
-Scientific method: determinism (have causes, regularities), discoverability (agreed-upon
scientific tools discover causes, a degree of confidence)
Systematic observation (less bias in everyday observation),
produce public knowledge (publicly verified, repeat),
Data-based conclusion,
Answerable questions (verified by observation/experiment, not theory, precise, prediction),
Tentative conclusion (not absolute, change/refine if additional evidence, self-correcting)
Theories can be falsified (results support or reject hypothesis...) e.g. A/B improve memory
P6-7
Pseudoscience (e.g. graphology): associates with real science (appear legitimate/resonable)
1) Rely on anecdotal evidence: selective, examples don’t fit ignored (data are biased)
2) Sidesteps falsification: Avoids falsification by explaining away anomalies, any contradictory
outcome can be explained. (lack specificity, avoid true test of theory)
3) Oversimplification of complex process

Tutorial 1 p8
Step 1: Specify variables: IV,DV (Experiment),Predictor & Outcome variable (Correlational Design)
Confounding variable?
Step 2: Select appropriate approach and design
Experiment vs correlational study, cross-sectional vs longitudinal

Step 3: Literature review: Web of science, sci-hub, google scholar, library search
Step 4: Prepare stimuli and/or measurement
Step 5: Ethics approval: least harm, reviewed by the ethics committee in advance.
Step 6: Recruit participant & Data collection: Criteria, Sample size: Rule of thumb, N=30,
Power analysis: G*power, Data collection: Consent -> Experiment/Survey -> Debriefing
Step 7: Data cleaning & Data analysis: Remove duplicate or irrelevant observations, Fix structural
errors (e.g., label issue), Filter outliers, Handle missing data
Correlation analysis & Regression

Lecture 2: Research Ethics p9-17


P9
Ethics? Standards governing the conduct of a person or the members of a profession.
Research ethics: what scientists should do/not
Historical cases of ethically questionable research:
• Watson & Rayner (1920) – scaring Little Albert, long-term, negative consequence, stress..
• McGraw (1941) – effects of repeated pinpricks: nervous system maturation, distress
• Dennis (1941) – raising children in isolation: min sensory stimulation, deprived
P10
APA Ethical Principles of Psychologists and Code of Conduct
First code, 1953, critical incidents technique: survey, examples of unethical incidents..
Most recent (2017): The Five General Principles
P10-11
The Five General Principles:
1. Beneficence and Nonmaleficence: greatest good, little harm, weigh costs vs benefits (IRB)
2. Fidelity and Responsibility: professional, serious, conscientious, aware responsibility society
3. Integrity: honest in all aspects, fake/false data, no misleading research (draw tentative
conclusion, overgeneralize finding), honest participants (deception if necessary), criterion for
excluding, cannot delete data after collecting
4. Justice: treat everyone fairly, expertise, reduce bias
5. Respect for Peoples’ Rights and Dignity: privacy, Confidentiality
P12:
-The Institutional Review Board (IRB): at least 5 ppl, faculty members from several departments,
at least one member outside community and a minimum of one non-scientist
key factor of decision is degree of risks to participants
Rationale, description of procedure, risk, alleviate, justify, informed consent..
Exempt: no risk,observation in natural setting, reaction time, survey.. no need ethic reports
Expedited: minimal risk, cognitive tests, quicker review, stress no more than everyday
Full review: several months
P13
Ethical guidelines for research with humans:
-Informed consent: procedure, goal, risks, quit anytime, confidentiality, anonymity, contact
information of researcher and IRB, opportunity of obtain final results e.g. Nuremberg code
Example: Willowbrook hepatitis study (parents forced to give consent)
Consent with special population: infants, children (parents, children assent, end if child undue
stress), prisoners (avoid compulsion, not force by parole board 假释委员会)
P14..
-Deception: desire participants act naturally, after: fully inform + why required
1) cover story (e.g. Milgram, punishment learning vs obedience, 65%, learner, teacher..., IRB..)
2) omitting some information: memory study
P15
-Debriefing:
Dehoaxing (true purpose, hypothesis),
Desensitizing (reduce stress, negative feelings)
Participants crosstalk (partial debriefing-->full when study complete)
P16
-Research ethics and internet: qualtrics; gorilla, pavlovia; Pacific, Sona System
Problem: 1) consent: not read, no answer question 2) debrief: desensitize, read

Ethical Guidelines for Research with Animals:


-Justifying the study: sufficient potential significance outweigh harm or distress
Justification increase as discomfort increase, aversive-->appetitive, natural habit, pain on them...
-Caring for the animals: expert in species use, carefully train people, aware of federal regulations,
legal suppliers, trapped humanely, alternative to destroying, if killing-->humane manner
-Minimize using animals for educational purposes: computer simulation..
P17
Scientific fraud: 学术造假
1) Plagiarism: ideas of someone else
2) Falsifying data: fabricate 捏造 data set, manufacture, change, favorable outcome
Discover due to failure to replicate data
Never discard data unless clear procedures specified before experiment, keep raw data

Tutorial 2 Literature Search & APA Citation Style p18-21


How to search literature? Topic--> research question, key words, other expression
Search rules: “”exact word
NOT --> AND --> OR
Simple search vs advanced search
Reference: Mendeley
APA citation:
Individual author: surname and initials up to 20 authors

Lecture 3: Developing Ideas for Research in Psychology p22-34


P22
Varieties of Psychological Research
- The Goals:
--Basic research (describe, predict, explain, foundation of applied research):
understand fundamental psychological phenomena
[Link] attention in read,dichotic listening(shadowing,difficult recall, unless meaningful)
P23
--Applied research: solve real-world problems
e.g. cell phone using while driving
Study1 (1, drive itself 2,shadowing 3,word generation task: not much affected by
shadowing task, it decreased sharply by more attention-demanding word‐generation task)
Study2 (hands-free/handheld, texting, in-vehicle speech-to-text)
Merging of basic & applied research --> translational research
P24
- The Setting: Laboratory versus Field Research
Laboratory research: low Mundane realism (study mirrors real-life experiences)
Experimental realism: has an impact on subjects, forces them to take the matter seriously, and
involves them in the procedures -->valid conclusion (whether in the laboratory or in the field)
Field Research: closely matches the situations everyday life, high mundane realism
Example: Effects of violent media on helping others (lab + field)
P25
Problem of field research: people chose to see violent/nonviolent films different types of people.
To control: half the trials were ran before the movie started
Manipulation check: sure intended manipulations in a study have desired effect
participants rate the level of violence in questionnaire
P26
Pilot study: Test aspects of the procedure to be sure the methodology is sound: reality of fight

- The Data: Quantitative versus Qualitative Research


Quantitative research: Data is in the form of numbers, Most research
Qualitative Research: analytical narrative, interview, case study, observational study, quote
Example: gender differences in control of TV remote would affect relationships among couples
men usually had control over what was being watched, a source of stress, misunderstanding.
Much research includes elements of both
P27
Empirical questions:
1) answered with data (quantitative/qualitative)
2) precisely defined terms (operational definitions: how concept studied in experiment, precisely
specified/ described operations, allow repeated)
Converging operations: confidence in understanding of some behaviour increased when you get
a whole bunch of studies, all using different operational definitions and different experimental
procedures, that all still converge on the same conclusion. E.g. color/orientation?
P28
Where do Research Ideas Come From?
1. Develop research from Observations (daily life, e.g. bystander effect, Kitty Genovese) &
serendipity (discover when looking sth else, accidental)
2. Develop research from Psychological Theories:
-Nature of theory: logically consistent statement about some phenomenon
(a) summarizes existing empirical knowledge (b) organize variables, precise statements of
relationships among variables (c) proposes an explanation (d) basis for making predictions
P29
Theory: Working truth, subject to revision pending the outcome
fact: outcome, theory: explain facts
Examples of theory: 1) working memory 2)cognitive dissonance: opposing cognition-->discomfort
P30
-Relationship between theory and research: reciprocal relationship
Induction (specific-->general; observation/results-->theory),
Inductive support for theory increases when individual studies keep producing the results as
predicted from the theory.
Deduction (general-->specific; theory-->study to test hypothesis (predict certain outcome...))
Reciprocal relationship between theory development and data collection:
Deduct (develop hypothesis)-->inductive support (outcome support hypothesis)
Fail a study--> not theory false; Repeated fail --> discard/alter theory; supports but cannot prove
P31
-Attribute of good theory:
Productive: good theories produce much research and advance our knowledge
Falsification: can be shown wrong (precise enough to be), maybe true if resistant to falsification..
Parsimony: conscise, simple explanation e.g. evolution theory, min constructs, assumptions
P32
3. Develop research from Existing Research: “programs of research”
-“what’s next?”, interrelated study, unanswered questions, new direction
Example: 1) retrieval practice --> long-term memory, SS ST-->SSSS SSST STTT
2) self-reference effect: structural, phonemic, semantic, self-reference --> close others,
other type of infor (color, location, images), other age group
P33
-Replication: Direct replication (exact procedure), Conceptual replication (partial, new features)
P34
Reviewing literature: computerized database searches
Best strategy: trial and error

Tutorial 03 Empirical question & Operational definition p35-37


Empirical questions: 1. precisely defined terms 2. answerable with data (qualitative, quantitative)
Operational Definition: A definition of a concept or variable in terms of precisely described
operations, measures, or procedures.
One concept could have multiple operational definitions.
Examples for operational definition:
1. Stroop effect, Delay = RT(incongruent) – RT(congruent)
2. memory: recall, recognition: no. of recalled items, accuracy, sensitivity index:
d’ = Z(hit) – Z(false alarm)
Positive effect for old ppl
Reliability: extent the outcomes are consistent when experiment repeated more than once
Low reliability: mistakes in designing experiment
Low validity: instrument used measure what u want to measure

Lecture 4: Sampling, Measurement, and Hypothesis Testing p38-55


P38
Population vs sample (typical, representative)
Sampling methods: probability vs nonprobability (whether random or not)
• Probability sampling: definable probability, fixed
if reflect attribute of population-->representative, if not-->bias
-Simple random sampling: equal chance, problem:1)systematic features, 2)large population
P39
-Stratified sampling: Proportions of important subgroups in population represented in sample
Different characteristics, Problem: large population, cannot have complete list of names
-Cluster sampling: Randomly select a cluster of individuals all having some feature in common
More convenient, If a cluster is too large, sample a smaller cluster within the larger one.
P40
• Non-probability sampling: non-random, easier, may not representative sample
- Convenience sampling: meet general requirements, recruited in variety of non-random ways.
Available, convenient, subject pool, Purposive sampling: specific type, non-psychology maj
- Quota sampling: representing subgroups proportionally, but non-random (recruit until filled)
- Snowball sampling: network of friends, share link & social media
P41
Measurement:
-What to Measure - Varieties of Behavior: construct-->operational definition-->measures
Construct: not directly observable, need inferred from measurement
--Example1: Do preverbal infants understand the concept of gravity--> “preferential looking”
5 month old-->2 differences (no. Stimulus dimensions), 7 months old-->1 difference
P42
--Example 2: Can you demonstrate that people use visual images? Mental rotation task
Construct: visual imagery, measures: reaction time, larger degree-->longer time
-Evaluating Measures:
•Reliability: repeatability and consistency of measures, low measurement error (close to true)
Meaningful comparison, confidence in reliability of measure develops overtime, correlational
P43
•Validity: measure what is designed to measure, study/Hypothesis properly conducted/ tested?
-Content validity: actual content of items measure construct, expert to evaluate
vs face validity: seems valid to those taking it
-Criterion validity: a measure is related to some criterion established prior, correlation
Predictive validity: forecast future behavior,
Concurrent validity: meaningfully related to other measure of behavior
P44
-Construct validity: a particular measurement truly measures the construct as a whole
Convergent validity: correlate with measures of theoretically related constructs
Divergent validity: not correlate with measures of theoretically unrelated construct
--Example: A test to measure self‐efficacy, locus of control (related), personality type
Measures can be reliable but not valid; valid measures must be reliable.??
P45
-Scales of Measurement:
Nominal scale: classify into group
Ordinal scale: relative standing, likert scale, cannot tell difference between each level
P46
Interval Scale (Order + equal intervals): IQ score, temperature, interval equal between...
0: another point on scale
Ratio Scale (Order + equal intervals + true zero point):
true 0: complete absence of attribute measured, e.g. physical measures.

Whenever a behavior is measured, numbers are assigned to it in some fashion


P47
Statistical Analysis:
-Descriptive Statistics: organize, summarize data
--Measure of central tendency: mean, median (half scores higher & half lower, outliers), mode

P48
Problem with central tendency: no dispersion of scores
--Variability:
Range: Tells little about the scores in the center of the distribution, outlier
IQR: range of scores between the bottom 25 percent of scores (25th percentile) and the top 25
percent of scores (75th percentile), show in boxplot, little influenced by outliers
P49
Boxplot
Measures of average deviation:
Variance: how distributed the scores are, relative to the mean
((X1-M)2 + ..+ (X20 - M)2 )/ n-1
Standard deviation: Square root of variance, only if normal distributed?
Variance is reported when the data represent the entire population of scores?
standard deviation is reported when the data represent a sample of scores from the population?
P50
Frequency Distribution: frequency table, histogram (in a defined range)
Normal distribution: hypothetical if all in population

P51
-Inferential Statistics: make inferences about the population based on the sample data,
determine if statistically meaningful
2 possibility: 1) from 2 populations, 2)from one, difference due to random error
P52
--Null Hypothesis Significance Testing (NHST): Null hypothesis(H0), Alternative hypothesis (H1)
Hypothesis testing:
p-value: probability of the 2 samples coming from one population, by chance
Most of data (95%) located in 2D, probability is too low if they come from one population
P53
Possible errors:
Type I (α) error: reject null hypothesis when it’s actually true; no difference, but detect
False positive reality: nothing, model: something
Type II (β) error: fail to reject null hypothesis when it is actually false; difference, fail to detect
False negative reality: something, model: nothing
P54
--Beyond Hypothesis Testing
-Effect size: size of difference between groups, magnitude + variability
cohen’s d, η2 (ANOVA)

Common metric for evaluating experiments, meta-analysis-->converging operations...


-Confidence intervals:
a range of values expected to include a population value with a certain degree of confidence.
Non-overlapping confidence intervals indicate a meaningful difference (how large the difference)
95% confident the interval capture population mean, also how large difference between groups
P55
-Power: 1- β
When a true effect exists, probability to find it, if power increases, type II error decreases
G*power calculate power, choose best sample size
Affect by alpha level, effect size, size of sample

Tutorial 04 Descriptive Statistics p56-58


P56
Definition (meaning of word, abstract)vs operational definition (concrete, replicable procedure,
measurable)
Why we need descriptive information? To apply conclusion in other conditions...See other
possibility, (cannot control), Criterion for sample...

Frequency distribution usually reported for nominal/ordinal


P57
Median can be reported for ordinal/interval/ratio scale.
Mode is often reported for categorical data (nominal/ordinal scale)
?
Histogram (whole pic, mode/median, but no mean..) vs blox plot (central tendency, variability)
P58
P: whether by chance or not, no magnitude and direction of difference
Confidence interval: magnitude, direction of difference
Treatment better/worse

Lecture 5: Introduction to Experimental Research p.59-70


P59
Experiment: effect of X on Y
Independent variable, manipulated, min 2 levels: x, not x/2groups from 2 categories
Control group/comparison group(treatment is withheld, baseline) vs experimental group
3 examples of control group:1) TV violence on children’s aggressive behaviors(violent TV/non-
violent) 2) Listen to Mozart improve memory? (Mozart/non-lyric music) 3) training program--
>sense of direction (receive/no)
P60
-Not all research need control group
Subject variable: groups of people differ from each other other than those directly manipulated
Select people depend on attribute they possess, already existing characteristics
Manipulated iv: true experiment-->casual relationship, iv precede dv, most reasonable explain
iv is subject variables: quasi experiment-->cannot control certain extraneous factors
(iv precede, but cannot eliminate alternative explanations)
P61
Experimental design vs correlational research (x & y related, don’t know reason)
IV can have >2 levels (control 可以出现在其中)
E.g. authoritarian, authoritative, permissive, uninvolved/ no. of bystanders
P62
Extraneous variables: may influence if not controlled, can systematically influence..
Confounding Variable: uncontrolled extraneous variables, co-varies with IV
E.g. coffee & decaffeinated coffee,
spread study vs learn at once: total study hours, retention interval
The distribution of practice is confounded with total study hours
memory, visual imagery: type of words, presentation type
P63
Measuring DV (use operational defin):
ceiling effect(easy, high score, no difference), floor effect-->pilot testing(moderate difficulty)
P64
Validity of experimental research:
provides understanding of behavior it is supposed to provide
• Statistical conclusion validity: Proper statistical analyses and conclusions based on analysis,
Reduce if wrong analysis, violate assumptions: continuous data, homogeneity of variance,
normal distribution (t-test)
If measures not reliable, error variability-->type II error...
• Construct validity: Well-chosen and well-defined IVs and DVs (accuracy of operational defin)
Example: violence, tom and jerry, aggressive behavior, “no”
P65-66
• External Validity: Generalize to other context:
-Other population (undergraduate, gender, culture):
Gender: Kohlberg’s research on children’s moral development (6 stages),
Overlook difference in thinking: boys (individual rights), girls (relationships)
Culture: individualistic vs collectivist, the rod and frame test (RFT), field (in)dependence,
European Americans, east Asians, females make more errors (framework influence decision)
patterns: Males, individual rights, Females, preservation of individual relationships.
-Environment(lab far from real life, Ecological validity:relevance for everyday cognitive activities),
-Longevity of results (social factors, historical context...), basic research less likely influenced...
Basic research: less likely to influenced by population, environment, time....
P67
External validity not determined by individual research project; accumulates overtime as research
is replicated in various contexts.
External validity is not a major concern, internal validity more important
P68
• Internal validity: cause-and-effect, no confound, methodology sound
-Pretest & post tests: whether change due to experience..
Example: anxiety management program
History: event outside study, grade-->pass/fail
Maturation: developmental changes, accustomed to college life
Regression to the mean: extreme score-->closer to mean, most scores around mean
Decrease of score due to regression to mean, not treatment
Testing (practice effect): same test, familiarity, aware of study purpose, pretest sensitization...
[Link] test, mere effect of pretest on post test
Instrumental: measurement instrument changes
Solutions: control group (for pretest & posttest), change due to treatment if no significant
difference for control
P69
-Participant problems:
subject selection: difference from participants, not treatment(free sign up)-->random assignment
Selection effects can interact with other threats...(history, maturation rate)
P70
Subject Attrition: leave study without complete
a problem when one condition in the study has a higher attrition rate than another
Group finishing study different type of people, could influence external validity.....
The change of DV comes from IV or difference between people (attrition, selection problem)...
To prevent: compensation for every session, follow-ups brief & convenient, routine reminders,
detailed contact information
If “attriters” & “continuers” indistinguishable at the start...?
Internal validity vs external validity

Tutorial 05 Scales of Measurement & Identify Variables


P71
Scales of Measurement
• Difference: different numbers mean different categories or values.
• Order: ranking.
• Similar Interval: the interval between any two adjacent numbers is equal.
• Meaningful Zero Point: 0 represents the absence of the quantity being measured.

P72-73
Likert Scale: interval scale, Usually 5 to 9 points
Odd number → Neutral midpoint (neither agree or disagree)
Each point corresponds to a semantic label, e.g., 1 “strongly disagree” to 5 “strongly agree”
2 types of IV: manipulated IV & subject IV
Spot the confounding variable: A social psychologist is interested in helping behavior

Lecture 6: Control Problems in Experimental Research p.74-90


P74
Between-Subjects Designs:
Necessary when: Subject variable is the IV, naïve participants(experience affect them, cannot
start fresh..) e.g. Bystander effect
Barbara Helm study: IV: type of crime, attractiveness, DV: yrs of sentence
P75
Problems: large number of participants, individual differences
Solutions: creating equivalent groups 1)random assignment, 2)matching
P76
1) Random assignment: not random sampling/selection:
Population --(random selection)--> sample ---(random assignment)-->control/experimental
Equal chance of assigning to any group, spread potential confounds equally
Example: presentation rate (2 or 4s/word) on memory, anxious
All anxious go to 4, no difference (Type II); anxious go to 2, worse memory (Type 1)
P77
If large number, greater chance equivalent group, if small number, may fail
Cannot guarantee equal number of participants in each group-->Block random assignment
Each condition has a randomly assigned participant before any condition is repeated.
P78
2) Matching: group together on matching variable, then distribute randomly to groups
when sample size is small and random assignment might not yield equivalent groups.
Matching variable correlates with DV. Hard if more than 1 matching variable
prefer use large sample size, assume random assignment distribute confounding factors evenly
P79
How to use a match procedure: get a score-->arrange the score in ascending order--> pair scores:
each consist adjacent scores-->randomly assign subject to different groups in each pair
2 conditions for matching: 1. matching variable effect DV (if correlation high-->sensitive to
between-group differences, 2. reasonable way of measuring matching variable (may bring
participants to lab on 2 separate occasions, bias)
P80
With-in subject designs (repeated-measures design)
Sometimes only reasonable choice: just a brief time to test but demand extensive preparation
orientations on Müller‐Lyer illusion: physiological psychology and sensation and perception..
Need fewer participants; if scare participants, small population, need special expertise..
Eliminate possibility: due to individual differences
P81
Example: 2 golf balls
Main problem: order effects
1) Progressive effects: practice effect & fatigue/boredom
2) Carry-over effects: Some sequences produce effects different from other sequences.
Experiencing 1st condition before 2nd affect person much differently than 2nd before 1st
If carry-over-->switch to between-subject design
P82
Example: effects of noise on a problem-solving task
UPN----->PN (worse performance); PN-------> UPN (better performance)
the order of conditions, independently of practice/fatigue effects,may influence the outcome.
P83
Counterbalancing: work better for progressive effects, use more than 1 sequence
If test once per condition:
• Complete counterbalancing: for a few conditions, x!, every possible order at least once
Problem: no of levels increase, orders increase dramatically
• Partial counterbalancing/incomplete counterbalancing: subset of total orders
Random sample of all possible combinations or randomize orders for each subjects (more simple)
Common if participants<no. of conditions...
-Latin square: each letter appears only once in each row and once in each column
(a) Every condition of the study occurs equally often in every sequential position
(b) Every condition precedes and follows every other condition exactly once
No. of rows = no. of conditions
Generate latin square-->randomly assign participants, participants no = or multiply no. of rows
P84
Testing more than once per condition:
• Reverse counterbalancing: presents in one order then presents them again in reverse order

e.g. stroop test


• Blocked randomization: every condition must occur once before any condition can be repeated
Within each block, order is randomized, participants can’t predict next, more frequently used..
same procedure in assign subjects randomly to groups in between‐subjects design..
An example from book: Competitors, red/blue, error bars-->standard deviation
P85
Methodological Control in Developmental Research:
-Cross-sectional design: between-subjects, time-saving, cohort effects
Cohort effects: different cohorts (generations) systematically different (environment, life story)
Example: intelligence, age
-Longitudinal design: single group study overtime, with-in, time-consuming, attrition
Attrition--> Group completing it may different from group starting it
P86
-Cohort sequential design: combine cross-sectional and longitudinal, save time??
A group of participants is selected and retested every few years, additional cohorts are selected
every few years and also retested over time.

Comparing the data in the rows gives you longitudinal designs, while comparing data in columns
(especially 2020, 2025, and 2030) gives you cross-sectional comparisons. Comparing the rows
enables a comparison of overall differences among cohorts.
Compare ppl 55 in 2010, 55 in 2015(same age, different cohort, if no difference, no cohort effect)
The attrition in one cohort can be compensated for by the data from other cohorts
P87
Controlling for the Effects of Bias
influenced by human bias, preconceived expectation about what is to happen in experiment.
two forms of bias often interact
Experimenter bias: experimenter may behave differently (even unconsciously) when they see
desired and undesired responses (e.g., smile vs. frown)-->participants know how to behave
Experimenter expectancy effects, also happens in animal research..
Other than expectation: race, gender, demeanor...
P87
Solutions for control experimenter bias: reduce direct involvement, standardize the process.
• Automatize the study with the aid of computer
• Standardize procedure (protocols: highly detailed descriptions of sequence of steps)
• Double blind: experimenter and participants do not know what to expect (which condition)
Vs single blind: participants unaware (aim), but experimenter know condition
P88
Participant bias: depends on what they’re expecting & their role, usually interact
• Hawthorne effect: change behavior because aware they’re being observed.
Hawthorne: how improve efficiency? Being observed; In spirit of helping study-->good subject...
• Evaluation apprehension: behave in ideal ways, not evaluated negatively, good person/ subject
Sometimes Conflict if want to help researcher, being evaluated positively is more powerful..
• Demand characteristics:figure out aim->behave in a way confirms it/according to interpretation
Demand characteristics: those aspects of the study that reveal the hypotheses being tested.
No longer act naturally--> reduce internal validity; more troublesome in with-in subject..
Devastating if affect some conditions but not others --> confound
P89
-Solutions: to reduce demand characteristics: aim not so obvious
• Effective deception: ask the question indirectly, use implicit measurements.
• Placebo control group: get medicine vs no but think they get, if different, placebo effect,
participants expectation..
• Manipulation checks: end of research, debriefing, aim? open-ended question,may remove.
Can also during experiment e.g. violent video games, also used to see if produce the effect it
supposed to produce..
• field research: do not know they’re being observed-->won’t guess research aim-->no bias
P90
True volunteer, interested in study vs reluctant, less interested
Exercise
Tutorial 6 Between & with-in subject design p.91-92
Difference between a between-subject design & with-in subject design
Advantage of between-subject design: naive with respect to the hypothesis

Lecture 7: Experimental Design I: Single-Factor Designs p.94-115


P94
Factor: single factor/multi factors (IV); level: conditions
Expose facto 可以 matching,但是和用 manipulated variable 实验中的 matching 不一样

P95-97
Four types of single-factor designs:
-Independent groups 1-factor design:
Study 1:
IV: type of note-taking (laptop, hand),
DV: memory: factorial questions/conceptual questions (understanding..)
Significant different for conceptual questions (hand is better than laptop), laptop write more
Study 2:
type everything, do not think, just copy the content lecturer said, too focused on typing..have
additional resources...
3 levels: Laptop (No Intervention), Longhand, Laptop (Intervention: not verbatim)

Handnote is beneficial, still try to write everything done


Study 3:
Study notes before take tests (ecological validity), still advantage for note
produces a deeper level of information processing and improves memory for the lecture.
Inter-rater reliability..
P98
-Matched groups 1-factor design:
Matching --> Random assignment when:
small subjects number, some attribute may affect dv, good way of measuring that attribute.
Study: Training on social skills of children with autism
Matching: standard scale for measuring autism-->equivalent group
IV: direct teaching/play activities (no training), DV: social skills (initiations, responses,
interactions)

significant improvement after 5 weeks for children with training

P99
-Ex Post Facto 1-factor design:
Traumatic brain injury (TBI), emotions (The Awareness of Social Inference Test (TASIT))
Show some video-->then emotion related questions...)
E.g. anger, sarcastic/sincere, sarcastic/diplomatic lies..
Matched on age, education, gender. Cannot random assign to different groups..
Make them more similar, but cannot say equivalent group..
TBI were significantly impaired in their abilities to recognize emotions.
P100-102
-Within-subjects 1-factor design (repeated measures 1-factor design):
ore sensitive to small differences between means
Single iv: complete counterbalancing...(if once)
First counterbalancing: stroop
Study: If two people share same experience, reactions might be amplified.
IV: shared experience/unshared experience: participant rate chocolate, confederate rate painting
DV: liking of chocolate..
No communication, but doing same task-->higher liking in the shared experience condition
Why? intensify or sharing make experience more pleasant? --> bitter chocolate, hate more..
P102
Single‐Factor Multilevel Designs: more than 2 levels, non-linear, if only 2 levels, linear..
E.g. Yerkes-Dodson Law
Wrong conclusions with 2 levels:
• Low & Moderate:
Anxiety improves the performance.
• Moderate & High:
Anxiety impairs the performance.
• Low & High: Anxiety has no effect.

Non-linear: better overall pic of arousal -performance is best at moderate levels of arousal, and
poor at either high or low levels of arousal
P103
Rule out alternative explanations: providing a framework helps memory, but only if the
framework is provided before the material to be learned
the multilevel designs include both between-and within-subjects designs of the same four types:
independent groups designs, matched groups designs, ex post facto designs, and within-subjects
or repeated measures designs.
P104
-Between‐Subjects, Multilevel Designs
Example: if young children would show a bystander effect (conceptual replication)
IV: Alone, Bystander, Bystander-Unavailable

because diffusion of responsibility, not shyness/social referencing


P105
-Within‐Subjects, Multilevel Designs
If 2-levels: limited counterbalancing options, >2, all counterbalancing options available..
Study: Do listening to “Mozart effect” improve memory? Partial counterbalancing - Latin square
IV: Listening to Mozart, rainstorm, not listening to anything, no difference between 3 groups
P106
Analyzing Data from Single-Factor Designs
-Presenting the data
-- Sentence: fine for 2-3 levels, strange if data increases
-- Table: means, SD
-- Graph (bar [Link]): Y-axis: DV, Never present same data in both table & graph (only use 1 form)
Graphs are striking if large differences/nonlinear effects/interaction
My own experience: Graph: important data, Sentence or table: less important data
P107:
Bar: compare between different groups, Line: track changes
If IV discrete: bar, line if continuous,
if iv is between-subject: bar (reflect separate groups)
If is with-in: line (experience all iv, connect..)
Bars may have error bar: SD, confidence intervals..
Most famous line graphs: Ebbinghaus forgetting curve, non-linear effects, Y axis: % saved

P108-109
-Analyzing single-factor, two-level designs
Difference caused by systematic variance (iv/extraneous) + error variance (nonsystematic)

Inferential statistical tests can be either parametric tests (assumptions)or nonparametric tests.
If nominal data --> chi-square test of independence
• t-test assumptions: Interval/ ratio, normally distributed (or close), Homogeneity of variance
• Independent samples t-test, for Independent groups designs, Ex Post Facto designs
• Paired sample t-test/dependent sample: Matched groups designs, Repeated measures designs
If some relationship, individuals that paired with similar scores essentially one person.
• One sample t-test: Whether an unknown population mean is different from a specific value
if don’t meet assumptions, alternate nonparametric test, Mann-Whitney U-test
P110
-Analyzing single-factor, multilevel designs
Multiple t-tests inappropriate, Increases chances of Type I error

Chances at least one Type I error, c = the number of comparisons being made
The more t-tests, greater chance
One-way ANOVA:
One-way=one independent variable, null hypothesis: level 1 = level 2 = level 3
Once overall significant effect found, then post hoc testing:
Tukey’s HSD test: equal numbers in each level.
Bonferroni correction: more conservative test, unequal sample sizes
Planned comparison: if not significant, but still do post hoc testing..
• ANOVA assumptions: Interval/ratio, normally distributed (or close), Homogeneity of variance
• One-way ANOVA for independent groups: Multilevel independent groups designs, multilevel ex
post facto design
• One-way ANOVA for repeated measures: Multilevel matched groups designs, Multilevel
repeated-measures designs
P111
Besides the typical control group situation in which a group is untreated, three other..
These are most informative when used in the context of multilevel experimental designs.
Special-Purpose Control Group Designs
• Placebo control groups
Treatment , Placebo: think are treated but not (expectation..), No treatment (baseline measure)

If treatment > placebo: effective, If no treatment = placebo: no participant bias, expectation..


P112
• Waiting list control groups
To insure equivalent groups in a study of program effectiveness
Example: two forms of therapy to treat clients who suffered from nightmares
Control: nightmare sufferers, no treatment, they will receive treatment later.. Ethical issues..
Another example: use subliminal approach for weight loss, Hawthorne effect
Can evaluate strength of placebo effect
P114
• Yoked control groups
Time each subject spend in the study is different.
Each member of control group is then matched, or “yoked”, procedural experience corresponds
Time spent keep constant for each group
Example: Eye movement desensitization and reprocessing (EMDR) therapy
Move eyes until low [Link]: same instructions, no eye movements..match session strength
therapy can control, but control group also
decrease... EMDR maybe due to placebo
effect

Tutorial 08: Data Analysis Exercise 1 Single Factor p. 116-121


Steps of data analysis: Data cleaning, Select the scale of the variable & Data label: nominal,
ordinal, Assumption check, Data analysis, Table or Figure, Report results
• Single Factor – Two level
-Between-subjects: Independent t-test:
Normality test (Shapiro-Wilk), Homogeneity of variance test (Levene’s)
Report:
Statistics, degree of freedom, p value, effect size, explanation in a non-mathematical way.
The main effect of the study method is significant, t(58) = -5.15, p<.001,
Cohen’s d = -1.33. The study method significantly impact the memory performance, participants
who studied individually recalled more items compared with participants who studied in a group.
-Within-subjects: Dependent t-test
• Single Factor – More Than Two level
-Between-subjects: One-way ANOVA,
If participants >30 per group, and similar group size, if violate for normality test, no big issue
-Within-subjects: Repeated measures ANOVA
Check assumption of sphericity: if significant result, violate assumption, Greenhouse-Geisser
correction or Huynh-Feldt correction is required
Main effect of time is significant, F(2.49, 72.09) = 188, p<.001. There are significant differences
among participants’ weight among four measures. (two degree of freedom, need report both..)
Bonferroni correction is used for the post-hoc analysis. Results indicate that participants’ weight
in the pre-training measure is significantly higher than that in the post-training measure (p<.001),
and the 1-month after training measure (p<.001). But no significant differences are found
between participants’ weight in the pre-training measure and in the 6-month after training
measure. Results suggest that this new training program is effective for losing weight in a short
period (i.e., 3 months), but the effect does not exist after 6 months.

Lecture 8: Experimental Design II: Factorial Designs p.122-143


P122
Factorial design = more than one independent variable (IV), factor = IV
Numbering system, identify no of iv and levels.
No. of digits = no. of IV, numerical values = no. of levels in each IV
2 x 4 x 4 factorial = 3 IVs, with 2, 4, and 4 levels, 32 total conditions
Factorial matrix: table that displays different combinations
Example: presentation rate, type of training (imagery/rote), gender: 2*2*2=8
Conditions =/= levels of IV,conditions=no. of cells in matrix = multiply numbers in notation system
P123
In factorial studies, two kinds of results occur: main effects and interactions.
Main Effects: Overall effect of a single independent variable.
difference between means of levels of any one IV. Combining all data for each level of that factor.
When yields both main effects and interactions, the interactions should be interpreted first

P124
Research example: closing time effect; time period: 3 levels, gender (2iv)
Ratings increase for both, but main effect of gender: males’ preference always higher
Possible confounding? Alcohol use? (no relationship with DV); reflect real attractiveness
differences (ppl different-->same photo, still..)
P125
Advantage of factorials over single‐factor designs: potential to show interactive effects.
Interactions: effect of one iv depends on the level of another iv.
E.g. 1st IV: teaching mode, 2nd IV: type of students...

no main effects, both mean = 75


Interactions can be described two ways:
1. Whether lab or lecture emphasis is better depends on which major is being evaluated
Science majors: better with lab (80>70), Humanities: better with lecture (80>70)
Effect of course type is depended on student major
2. Whether science or humanities majors do better depends on the type of course
If single factor, 75=75, no effects, factorial designs can be more informative than single ‐factor
P126
Research Example: 2 IV: study conditions (silent/ noisy), test conditions (silent/ noisy)

No mean effects, but interaction: best memory when study and test conditions match
P127
A similar study:
retrieved the most words when their encoding and testing
circumstances were the same

P128
Interactions Sometimes Trump Main Effects: main effects not matter, but interaction...
P129
Combinations of Main Effects and Interactions
In a simple 2 × 2 design, there are 8 possibilities:
1. a main effect for the first factor only
Main effect for memory strategy, no main effect for presentation rate, no interaction

2. a main effect for the second factor only


No main effect for memory strategy, a main effect for presentation rate (22>14), no interaction

3. main effects for both factors; no interaction


Main effect for memory strategy and presentation rate, no interaction

4. a main effect for the first factor plus an interaction


5. a main effect for the second factor plus an interaction
6. main effects for both factors plus an interaction
--Interaction and two main effects, but main effects of little importance

Ceiling effect, interactions can trump main effects, 4 s per item produces better recall than 2 s per
item is misleading: only true for the rote groups. the interaction is the key finding here.
--Interaction and two main effects, but main effects more important this time
There is a general trend: imagery is better than rote (in both 2s and 4s)
slowing the presentation rate improves recall for imagery group (23 is a bit better than 19), but
slowing rate improves recall considerably for the rote rehearsal group (15 is a lot better than 5).
At fast rate, imagery training is especially effective (19 is a lot better than 5—a difference of
14 items on the memory test). At the slower rate, imagery training still yields better recall, but
not as much as at fast rate (23 is somewhat better than 15—a difference of just 8)
7. an interaction only, no main effects
8. no main effects, no interaction
A standard feature of interactions:
if parallel, no interaction; If nonparallel, an interaction probably exists. Easier for line than bar
If interaction-->line, even if between, x-aixs discrete variables
Whether an interaction exists is a statistical decision, to be determined by an ANOVA.
P131
Creating Graphs for the Results of Factorial Designs
Error bars: standard deviations/standard errors (estimate of population standard deviation)/ CI
Variability

P132
Classic study: To Sleep, Perchance to Recall
An exercise
P133
Varieties of Factorial Designs
In most cases P variable is a between‐subjects factor because it is a subject variable. But it could
be a within‐subjects factor if the participants are tested over time, as in a developmental design.
Person * environment (P*E) interactions: A main effect for the P factor (i.e., subject variable)
indicates important differences between types of individuals that exist in several environments.
A main effect for the E factor (i.e., manipulated variable) indicates important environmental
influences that exist for several types of persons.
Environment factors: we can manipulate...
Mixed: 至少一个 between 一个 within
P*E: some manipulated some subject (only for designs with between-subject)
If mixed P*E: E-->within-subject variable..
P134
Mixed factorial designs: At least one IV a between, and one a within, All manipulated...
Deal with both the problems of equivalent groups and the problems of order effects.
But not always – no counterbalancing if the order is of research interest.
- Research Example - mixed factorial with counterbalancing
Terror Management Theory: if reminded of future death, impacts their coping mechanisms.
Salience of instructions: mortality/ exam (between)
leadership style of political candidates (within), counterbalanced (6 orders, complete):
Charismatic, Task-oriented: Relationship oriented
Task-oriented leaders were
generally favored. Reminding of
mortality increase evaluations of
charismatic leaders and
decrease their evaluations of
relationship‐oriented leaders.

P135
-Research Example - mixed factorial without counterbalancing
examines changes with the passage of time—a trials effect—no counterbalancing
work-as-exercise instruction: told/not(between), time: before/ after 4 weeks (within-subjects)

Significant interaction in all, except for the last one...


P136
P x E designs: Factorials with subject and manipulated variables
P = person factor (a subject variable), E = environmental factor (a manipulated variable)
If E is a within-subject factor --> mixed P x E factorial
“Columbia bible”, P: individual factors & E: situational factor powerful enough to influence the
behavior of many kinds of persons.
B (behavior)= f(P,E). Lewin’s formula, behavior is the joint function of person & environment P141
P137
Examples: P: Introverts & extroverts solve problems in E: a small room /a large room
– Main effect for P factor, no main effect for environment, no interaction...
individual differences apply to more than one
environment

– Main effect for E factor


environment (room size) produced the powerful
effect
this effect extended beyond a single type of
individual

– P x E interaction
For one type of individual, changes in the
environment have one kind of effect, while for
another type of individual, the same environmental
changes have a different effect.

P*E factorial designs popular in educational research/psychotherapy,


Aptitude-Treatment Interaction designs (ATI),
aptitude: subject (person) variable, treatment:manipulated, environmental variable.
P138
-Research example - P × E Factorial Design with Interaction
gender (subject variable P), math test with same-sex/opposite-sex (manipulated variable E)
Interaction: for men, performance was unaffected by who was
taking the test with them; for women, they did poorly when in a
room with men.

Another explanation: “tokenism”, but rule out..


-Research example - A Mixed P × E Factorial with Two Main Effects, no interaction
If if E is within-subject design --> mixed P * E factorial..
age: young/old (subject variable), driving with/without cell phones (within-subject variable)

young/no cell phones: RT faster


P141
Recruiting Participants & Analysis
Number needed depends on whether are tested between- or within-subjects

Ethics: 2 consent form, written protocol ahead (standardized procedure), if stress-->end,


debriefing is very important
P142
Analyzing data from factorial designs
factorial design: more than one F ratio. an F for each possible main effect & interaction.
In an A × B × C factorial, seven F ratios will be calculated: three for each of the main effects of A,
B, and C; three more for the two‐way interaction effects of A × B, B × C, and A × C; plus one for
the three‐way interaction, A × B × C.
Two between-subject IVs: Two-way ANOVA, Two within-subject IVs: Repeated measure ANOVA
if a mixed factorial design: A mixed ANOVA
subsequent (post hoc) testing may occur with factorial ANOVAs.
For example, in a 2 × 3 ANOVA, a significant main effect for the factor with three levels would
trigger a subsequent analysis (e.g., Tukey’s HSD) that compared the overall performance of levels
1 and 2, 1 and 3, and 2 and 3.
P143
Following a significant interaction--> simple effects analysis: break down interactions by
examining effect of each independent variable at each level of the other independent variable
A simple effects analysis would make these comparisons:
1. For science majors, compare lab emphasis (mean of 80) with lecture emphasis (70)
2. For humanities majors, compare lab emphasis (70) with lecture emphasis (80)
3. For the lab emphasis, compare science (80) with humanities majors (70)
4. For the lecture emphasis, compare science (70) with humanities majors (80)
Sir Ronald Fisher invented ANOVA

Tutorial 09 Data Analysis Exercise 2 Two Factors p.143-148

As long as one with-in, click repeated measures ANOVA

Lecture 9: Non-Experimental Design I Survey Methods p.149-170


P149
non‐experimental, correlational, descriptive research designs: not manipulate, description of relationship
Survey Research
Survey: a structured set of questions/statements to measure attitudes, beliefs, values, tendencies
to act.
Surveys vs. Psychological Assessment (reliability & validity, measure abstract construct, e.g. BDI)
P150
Sampling issues in survey research
• Biased vs. representative samples:
-Non-probability vs. Probability sampling, surveys most effective when representative sample.
-- Self selection bias (decision to participate in a study is left entirely up to individuals)
e.g., election of 1936 (Literary Digest), 25% ballots returned, own cars (unrepresentative, middle
class, more likely to support republicans...)
P151-155
Creating an effective survey
1. Types of survey questions or statement
Open-ended:wide range of responses,a sense of control, but hard to score,long time to complete
Use sparingly: at the end of closed questions, comment answers for extreme ratings..
As a pilot test: options, Partially open item: ..... other (please state)
a common open‐ended question is sometimes referred to as the “most important problem” item.
Combine closed ended and open-ended... in the end, explain answers, problem in survey
Closed questions: specific response, yes/no, likert scale: statement, different levels of agreement
Likert scale: odd: midpoint, neutral, means, sd, no clear ad of 5/7: depends on experimenter
5: sufficient discrimination among agreement, but de facto 3-point, avoid extreme
7: yield 5-points, but increase time to complete (add extra levels of discrimination)
Avoid mixing formats.

- Avoid response acquiescence (默認): a tendency to agree with statements.


Reach each carefully, make item-by-item decisions
- Start: unsensitive,non-personal,interesting questions,cluster questions with same theme
2. Assessing memory and knowledge
-Don’t overburden memory:
interval depends on the questions (if interval too long-->memory errors (month-->week), but if
infrequent one, zero)
Aid memory via providing list, aid retrieval
-use DK (“don’t know”) alternatives sparingly/moderate use
If DK: overuse, but if no: force participants to choose, truly ignorant
disguise knowledge questions: “Using your best guess..” /”Have you heard/have you read that...”
Use DK only when reasonable to expect some respondents have no idea of the answer might be.
[Link] demographic information:Age,gender,socioeconomic status,marital status,end of survey
Group results by demographic info
only demographic categories important for empirical questions
If too much-->longer survey time, irritated, invasion of privacy

4. A key problem: Survey wording


- Avoid linguistic ambiguity (pilot study helps): interpret different-->define the term
- Don’t ask for two things in one question (Double-barreled questions)
- Avoid biased and leading questions: Background info could misleading...
Avoid carry-over effects: previous items influence responds on later item
P155-158
Collecting survey data
• In-person interview surveys
Pros: comprehensive, detailed(follow-up questions, probes),clarify info, reduce unclear questions
Cons: 1) Representative samples, around the place of interview, poor, homeless
2) cost (training, standardized procedure), logistics (travel expenses), interviewer bias in face
to face (e.g., cross-race bias)-->Certain types of interviewers trained for specific purposes.
• Mailed written surveys
Pros: 1) different segments of the population, large population, representative
2) convenient (complete the survey at any time they want)
Cons: 1)Response rate (if low, unrepresentative, Return rate: Excellent: 85%; Very good: 70-85%;
Acceptable: 60 -70%; Bad: below 60%);
nonresponse bias: people who return surveys differ in some important way from those
who don’t return them (non-responder: older, unmarried male without a lot of education),
Nonresponse occurs when people have some attribute that makes the survey irrelevant for them.
Decent return rate if: brief and easy to fill out; start with interesting questions; before the
survey, participants are notified; nonresponse triggers follow-up reminders; professional, sign by
real person not machine; return postage; not a sales pitch; small gift
2) social desirability bias: attempt to create a positive picture, how they should respond
Ensuring anonymity can help reduce, but persistent..
• Phone surveys: popular 1980s, random‐digit‐dialing,telemarketing & cell phones lowpopularity,
sugging - Selling Under the Guise of a survey, general public not trust
To increase response rate, precede call with a letter/ email: Mixed-mode approach
Ad: combines the efficiency of a mailed survey with the personal contact of an interview.
Phone survey: unit of measurement: “household”; cell phones: “individual”
•Online surveys: Url via e-mail, purse list/by “search spiders”/post on listserv/social
media/software: SurveyMonkey, Qualtrics, Amazon’s Mechanical Turk (MTurk)
Pros: efficiency (large data, short time); min cost; open 24h; complete in less time
Cons: Sampling: able to use internet, but old people/young children; appear to be entryway for a
virus; delete messages without reading them; Don’t know who, same person many times..
Mixed mode: letter includes web address and password
P158
Ethical concerns: APA no need informed consent “anonymous questionnaires, confidentiality..”
But usually have, survey can be misused
P159-165
Analyzing Data from Non‐Experimental Methods:
• Correlation
statistical technique to determine degree two variables are related,Relationship, no casual
Can occur for all types of scale
Three types of [linear] correlations:Positive correlation (high associate with high), Negative
correlation (high associate with low), No correlation
-Scatterplots:
Visual representation of the relationship, Use parameters to determine how strong
Stronger relationship, closer to straight line,
Strength can be inferred from scatterplots & correlation coefficient & coefficient determination
Some relationship may not linear, cannot use statistical methods assume linearity
Each point represent individual
-Correlation coefficients:
Strength and direction, (r) from –1.00 to +1.00, perfect negative/ positive correlation
Statistical tests: Pearson’s r (interval/ratio), Spearman’s rho (ordinal), phi coefficient (nominal)
Numerical value: strength; sign --> direction
Can be relate to cohen’s d (effect size)
-Coefficient of determination: r2, always positive
Proportion of variability in variableA that can be accounted for (or explained) by variability in the
other variableB; how much variability is shared across both variables, shared variance
The remaining proportion can be explained by factors other than your variables
r = .60 --> r2 = .36
36% of the variability of one variable can be explained by the other variable (B)
64% of the variability can be explained by other factors
-Be aware of outliers: it distort r, could Type I error, extend beyond 3 sd, delete if established prior

• Regression: Making predictions


Regression analysis: Making predictions on the basis of correlations
if the correlation is not significant, regression analyses should not be done
-Regression line: straight line best summarizes a correlation, line of best fit, min distance to line
Y = a + bX, predictor variable to predict criterion variable
Y = criterion variable, variable being predicted; X = predictor variable, variable doing predicting
a = point where regression line crosses Y axis; b = the slope of the line
Higher correlation, closer points to regression line, smaller distance,more confident in prediction.
Confidence==>confidence interval, more confidence (strong correlation),narrower range CI
Beta coefficient(β): strength of predictor variable’s ability to predict changes in criterion variable
F/t-test can be conducted based on regression analysis: whether predictor variable is significant
- Bivariate (if one predictor variable..)
- Multiple regression: One criterion variable, More than one predictor variable
Relative influence of each predictor variable can be weighted
(B: relative importance of predictor = beta weights)

Multiple regression analysis yields a multiple correlation coefficient (R) and a multiple coefficient
of determination (R2).
R: correlation between combined predictors and criterion
R2: variation in the criterion variable that can be accounted for by the combined predictors.
Ad: when several predictor variables are combined (especially if predictors are not highly
correlated with each other), prediction improves compared to the single regression case.
The regression analysis, however, enables researchers to examine the unique contributions (i.e.,
unique variance) of a variable by controlling for shared variance across variables.
Predictions made only for ppl fall within the range of scores on which the correlation is based.
P166-169
Interpreting correlational results
• Directionality problem
A could cause B, or B could cause A, impossible to determine direction of causality
Reduced by Cross-lagged panel correlation: measure correlation at several time points, A1 is the
cause if significant correlation between A1 and B2 (A1 happens before B2)

2 variables at 2 time, 6 correlations, .31 is significant


Partially the cause..
But the correlation could be caused by .38 and .21...it could be aggressiveness in third grade
produced both a)preference for watching violent TV in third grade and b) later aggressiveness.
Problems of interpretation remains
• Third variables: uncontrolled third variable could cause both A and B to occur
Deal with Partial correlation measures remaining relationship between A and B, with third variable
partialed out or controlled /mediation‐moderation analyses, control for 3rd variable statically
Example 1: remaining relationship between reading speed & reading comprehension, with IQ
partialed out or controlled.
To complete a partial correlation, correlate (a) IQ and reading speed and (b) IQ and reading
comprehension. Incoporating all 3 correlations
If complete a partial correlation and original correlation does not change much, rule out third variable.
Example 2: 12 third variable, there is a correlation, and the preference is the cause

Two different types of third variables that may help explain a correlation:
Mediator: explains how or why a relationship between two variables exists
Third variable correlated with 2 variables..
If impulse control is statistically tested as a mediator, then original correlation between alcohol
use and sex is reduced, like what occurs with partial correlation techniques.
Moderator: explains under what conditions the relationship between two variables exist.
The correlation between alcohol use and risky behavior happens in only some conditions..

Partial correlation is not typically used to examine moderation because it focuses on controlling
for the association between variables rather than exploring the interaction effects between
variables. partial correlation can be used to examine the existence of a mediator but not a
moderator. (By AI)
P169-170
Combining non-experimental and experimental methods
Correlational study --> create casual hypothesis--> experimental studies
-Research example: lonely --> anthropomorphize objects
Study 1: correlational study
Studies 2 & 3: manipulated loneliness to tests its effects on likelihood to anthropomorphize
Iv2: false personal feedback (will be lonely); iv3: watching video clips
Converging operation (study 2&3 manipulate loneliness in different ways-->similar finding)

Tutorial 10 Data Analysis Exercise 3 Correlation & Regression p. 196-199


P196
Assumption of pearson correlation:
1. Both variable ratio/interval
2. Linear relationship
3. Both variables should be normally distributed
4. Each observation should have a pair of values
5. No outliers

P197
Linear regression:
1. Linear Relationship: There should exist a linear relationship between the IVs and the DV.
2. No Outliers: There should be no extreme outliers in the dataset.
3. No Multicollinearity: IVs are all linearly independent, if correlated--> lower statistical power

Difference between correlation and regression:


Correlation: bi-direction, don’t know which predict another X--Y, both direction is ok
Regression: clear direction, X (IV, predictor)--->Y (DV, outcome), X to predict Y

Example: moderator or mediator?


Mediator:
X--------------->y
---->m------->
X can predict Y, X to predict M, and M predict Y
1. X---------->Y 2. X---------->M 3. M--------->Y (mediator can significantly predict Y)
Moderator: Relationship between X & Y depends on levels of W
Whether interaction between X and W is significant or not

Lecture 10: Non-Experimental Design II Observational and Archival

Methods p.171-180
P171
Observational Research
Produce descriptive info
Varieties of observational research:
-Observe a variety of behaviors (global, e.g. Jan Goodall)/specific behaviors (snake selection, obesity)
-Structure of the setting: 0 structured (no influence) to highly structured (create structured
environment and observe)
0 structured: fighting for men/women; structured: usually in lab, aka lab observation studies, e.g.
child, helping, parent-children
-Involvement of experimenter: naturalistic observation (two-way mirror, video recorder, habituation
strategy for animals, not interact (appear reading), act in everyday environment, semi-artificial
environment sufficiently ‘natural’) and participant observation (involve, may become a member,
presence known to groups, only option if group closed to outsiders, common technique for qualitative
research, narrative analysis)
P172
Advantages of Observational Research
1)Can be a rich source of ideas for further study (questions/hypothesis)
2)theory testing (falsification): observation contradicts a theory-->theory is problematic.
E.g. animal aggression/ animal learning
P173
Challenges Facing Observational Methods
-Absence of control (draw conclusions carefully): imitate? Or simply attractive?
-Participant reactivity: behaviour influenced by awareness of being observed/recorded, a problem for
participant observation E.g. habituated-->hard to say the extent..
To reduce: Obstrusive measures: subject is unaware of the measurement
- Direct unobtrusive measures (e.g., hidden video or audio recordings )
- Indirect unobtrusive measures: record events one assumes resulted from certain behaviors even
though the behaviors themselves were not observed.
-Ethics: consent and privacy issues
Reducing reactivity raises the ethical problems of invading privacy and lack of informed consent
APA ethics code: naturalistic observation not require informed consent/debriefing. In public and strict
confidentiality and anonymity are maintained.
Sometimes, group not to consent to participant observer. informed consent and invasion of privacy.
-Observer bias: preconceived ideas what will be observed.. e.g. aggressive behaviors for girls/boys
What data select as relevant, omit, influenced by observer bias
To reduce it:
-Behavior checklists, predefined behaviors, operational definitions, training observers
-Inter-rater reliability (several observers), percentage they agree
-video recording, if procedure is mechanized
-sampling procedures: time sampling, event sampling
P174
Observational Research Example: A Naturalistic Observation in a science museum, inform consent
-Boys and girls did not differ in
terms of how much time they
spent at exhibits or in the
extent to which they actively
engaged the exhibits
- Parents explain science
concepts more to their sons
than to their daughters
Same pattern regardless of age, both dads & mums favored boys
Another Research—A Covert Participant Observation: homeless, maintenance strategies
P175
Analyzing qualitative data from non-experimental designs
some qualitative data may be transformed into quantitative data with techniques like coding.
Thematic analysis:
Identifying patterns of responses in qualitative data (inductive)
Based on theories, researchers may predict certain themes prior to data analysis (deductive),
coming to data with preconceived theme based on theory

any order
P176
Archival Research: use info already collected for other purpose, not collect new data.
Have iv, but not manipulated, so non-experimental research
Data from archival studies can be subjected to many statistical analyses: correlation, regression,
factor analysis, meta‐analysis.
Archives: records + places stored (database)
Archival data: census data, government records, credit histories, health history, educational
records, diaries, tweets data, etc.
Big data: vast amount of data available in databases that can be extracted and analyzed with
advanced data analytic tools.
Before statistical analysis: content analysis, systematic examination of qualitative information in
terms of predefined categories. Some subjectivity, multiple coders, inter-rater reliability estimate
P177
Dis: some info may missing, available data not representative..;
experimenter bias-->control, double blind
Ad: large amount of info, rule out reactivity, test hypothesis cannot for experiment manipulation
Not allow random assignment, but can control
Example: recovery from surgery could be affected by the type of room
Control: common type of gallbladder, Matching variables, nurse decide use which records
P178
Analyzing archival data – factor analysis & meta-analysis
-Factor Analysis
A multivariate technique in which a large number of measured variables are correlated with each
other. Then determined whether groups of these variables cluster together to form factors.
Example:
Pearson’s r’s could be calculated for all possible pairs of tests, yielding a correlation matrix
-some of the correlations cluster
together.
-Correlations between tests from
one cluster and from the second
cluster are essentially zero.
-2 mental abilities/factors: verbal
fluency & spatial skills
E.g. 16 major personality traits
Factor loading: correlations between each of the measures and each of the identified factors.
E.g. the first three measures would be heavily loaded on Factor 1 (verbal fluency), and the second
three would be heavily loaded on Factor 2 (spatial skills).
Factor analysis only identifies factors; what they should be called is left to researcher’s judgment.
Once such factors may be identified by a researcher, then decisions can be made on whether to
use those factors as predictors in a regression model.
P179
Meta-analysis
Systematically synthesize literature on some phenomenon and combine the results across studies
analyzes the effect sizes, special type of archival research, uses data from pre ‐existing sources
-Is the effect consistent across studies? -If the effect is consistent, what is the size of the effect?
P180
Meta-analysis example: Verbal Overshadowing Effect,25->12,direct replications from different lab
Verbalization of experience/perception impair subsequent visual recognition or memory recall.
Chapter 12. Small N Designs p. 181-196
P181
History: before the development of statistical analysis, study own behaviors/single individual
Separate in role not evident, subject more like “observers”, additional participants-->replication
Example: Ebbinhaus, little albert, Wundt, Dresslar studying facial vision (no descriptive data, just
individual data, no average.. ), lab-->field, basic-->applied..
One of origin of small N: cats in puzzle box
P182
Small N designs/ Single-Subject Designs:
one or few individuals, Data not summarized; data for each subject presented
Reasons for Small N Designs:
1. Occasional misleading results from statistical summaries of grouped data
A failure of individual-subject validity: the extent to a general conclusion from research with
large N applies to the individual participants in the study.
Summarize--> fail to characterize the behavior of the individuals/disguise the differences, group
data like a process, functional relation
-Example: concept‐learning experiment in children (rewards)
Grouped data --> continuity theory (gradual process), Individual data --> noncontinuity theory
Chance level-->if correct solution-->performance improve dramatically
(reach criterion at different rates), combine individual curves-->smooth curve
Large N designs need to examine the individual data

P183
2. Practical problems with large N designs
-Participants with a particular attribute are rare (expensive, time consuming, ethical..)
-Members of a specific animal species are rare, costly, or require much time for training
P184
Philosophical questions: Skinner: study individuals and derive general principles only after study
of separate cases. (reduce random variability, precise control)
Psychology, an inductive science, reasoning from specific cases to general laws of behavior.
The Experimental Analysis of Behavior:
Skinner: if a researcher is able to establish sufficient control over environmental influences, then
orderly behavior will occur and can be easily observed (i.e., without statistical analysis).
Dv: response rate, Operate conditioning..
To predict and control behavior:
(1) occasion which a response occurs, (2) the response, (3) reinforcing consequences.
The interrelationships among them are the ‘contingencies of reinforcement’
Cumulative recorder, response rate: slope
Schedules of Reinforcement: a rule that determines the relationship between a sequence of
behavioral responses and the specific occurrence of a reinforcer.
2 categories of scientists: 1) contemplative ideal, basic causes 2) technological ideal, use science
to control/change world
Applied Behavior Analysis: use operant principle solve real-life behavioral problems.
P185
Small N Designs in Applied Behavior Analysis
Small N designs to evaluate the effectiveness of these applied programs
• Simplest single-subject design: A-B design
Ideal if behavioral change when A changes to B, no control group
3 elements: 1)operational define 2) baseline level of responding (typical frequency) 3) begin
treatment, monitor behavior
Problem: Change in behavior can be due to confounds (history, maturation, regression to the
mean), Solution: Withdraw Designs
P186
• Withdrawal designs: A-B-A design/ A-B-A-B design
Behavior unlikely to return to baseline if due to confounds. If return, caused by B
If behavior changes accompany the introduction and removal of treatment, confidence is
increased that the treatment is causing the change. confidence is further strengthened if ABAB
ABAB: evaluate twice, replication, ethical ad: finish study with treatment
Research Example: A-B-A-B design: Peer reinforcement, on-task behaviors, ADHD children.
children solved more math problems during
treatment sessions than during baseline sessions.

P187
Problem of ABAB designs: A behavior does not always return to its baseline (learning a skill),
ethical problem if behavior changed is self-destructive
Solutions: Multiple Baseline Designs
• Multiple Baseline Designs
Several baseline measures are established, 3 baselines are typical.
behavior should only change when the training is put into effect
Used when withdrawal designs not feasible, ethical
Three varieties of multiple baseline:
-One behavior, two or more subjects: different time for different subjects
-Two or more behaviors, one subject: different time for different behaviors, If all three behaviors
change when the program is put into effect for the first behavior, then it is difficult to attribute
the behavior change to the program; the result of confounds.
-Two or more environments, one behavior, one subject:
Example: the third type, reduce drooling in individual..

P189
• Changing Criterion Designs
Used when target behavior shaped gradually, target behavior is hard/ complex to reach all at
once, shaped in increments
Shaping: develop behaviors by reinforcing gradual approximations to final behaviors.
Baseline-->treatment continue until initial criterion reached-->criterion increasingly stringent
Example: exercise behaviors of obese and non-obese boys, 15% higher than baseline
criterion not same for
everyone (Small N)
bell would ring and the light
would go on when they
pedaled at a rate that was,
on average, 15% higher
than their baseline rate.

Boys worked for reinforcers they valued, Social validity: (a) whether applied behavior analysis
program has value for improving society, (b) whether its value is perceived as such by the study’s
participants, (c) the extent to which the program is actually used by participant
P191
• Alternating treatments design
Evaluate/compare more than a single treatment approach
After the usual baseline is established, different treatment strategies (usually two) are then
alternated numerous times (a form of counterbalancing).
-Example: reducing stereotypy in children with autism
Two alternating play interventions (No AOC/with AOC)used in three targeted behaviors
it appears the AOC plan worked: Functional play increased, stereotypy decreased, and the only
increases in the screaming and falling behaviors occurred during non‐AOC.
Limitation: just a single child, replicate; withdrawal sessions..
P192
Evaluating Single‐Subject Designs
Ad:
helpful to assess the effectiveness of conditioning approaches, from behaviorist, if well
controlled, predictable, effective in therapeutic..
Study individuals in depth
Dis:
1. External validity: therapy generally effective for others/other settings? Sth unusual about the
individual. BUT replications are common/internal validity>external, problem for large N as well
2. Not using statistical analysis, just mere visual inspection
Defenders: conclusion only if effect large enough, time series analysis recently
3. cannot test adequately for interactive effects.
Interactive (P*E): multiple baseline across subjects designs, extensive replication: 2 种情形
4. DV replies on response rate
P194
Another small N research strategy:
Case study designs: A detailed description and analysis of a single individual.
Usually in narrative form, qualitative research, but also quantitative data, various methods
Example: Henry Molaison (1926–2008): surgery destroy medial temporal lobe, hippocampus
Yes: short-term memory, procedural memory, recall long-term memories before his surgery
No: new long-term memory
Distinction between declarative vs nondeclarative memory
Example: (2 ends of continuum)
(1) fighter (A.B.) with head trauma, ataxic dysarthria (2)S, remarkable memory (visual imagery)
(2) Some case studies investigate unique events/groups e.g. flashbulb memories
Case studies often have a broader meaning in disciplines other than psychology.
P195
Evaluate case study:
Ad:
1. detailed analysis not found in other research strategies, informative, various methods
2. well‐chosen cases can provide prototypical descriptions of certain types of individuals.
Comparison, enhance understanding...
3. inductive support for a theory, hypothesis, falsification
4. Sometimes only way to document extraordinary person/event
Dis:
1. External validity, generalization, could be atypical case (unique features)
2. Biases of investigator
3. Memory could be distorted, constructive
4. Lack of control

Chapter 11. Quasi-Experimental Designs p. 200-213


P200
Beyond the lab
Applied research:
1)easily recognizable problems 2)further knowledge of basic psychological process
Basic approach and applied approach:

Example: Trudel et al (3 experiment + field)


P201
History: American psychologist apply basic research methods to solve problems in education,
mental health, child...
Ad of applied research:
External validity (ecological), setting assemble real-life situations, address everyday problems
Dis:
-ethical: informed consent, privacy, coercion
-trade of between internal vs external, in field--> lose control
-for between: hard to random assignment, ex post facto, subject selection, maturation...
-for within: not always counterbalance properly, attrition for long-term
P202
Quasi-Experimental Designs
subjects cannot be assigned randomly, maybe because practical/ethical reasons..
(equivalent group: random assignment/matching...)
Can be considered as quasi-experimental designs:
- Single-factor Ex post facto designs
- Ex post facto factorial designs
- P x E (between) factorial designs; mixed P x E (within) factorial designs
- All of the correlational research
2 specific types, frequent, but other also exists: regression discontinuity design):
Nonequivalent control group designs, interrupted time series designs
P203
Nonequivalent Control Group Designs
evaluate the effectiveness of some treatment program, Specific example of ex post facto design
Nonequivalent groups at beginning, build-in confound

pretest --- treatment --- posttest


Not compare the posttest of 2 groups, but between the change scores (the difference between
O1 and O2) for each group.
Cannot draw conclusion simply because the increase for experimental group, look at the control
Example: flexible working schedule-->productivity

4 possible outcomes:
A) : both increase, may because of history, maturation..
B) (1)ceiling effect for control, cannot see the improvement (history/maturation), (2) regression
to the mean, extreme low at pretest for experimental group
C) Subject selection could interact with history... selection * history
D) Regression to mean could be ruled out, but still interaction between subject selection and
history...(hard to exclude) but it’s good evidence
P205
Research example: play street, physical activity

Not all nonequivalent control group has pretest, example: trauma-->nightmare


P206
Regression to the Mean and Matching
Could happen when two groups are sampled from populations that differ on the factor being
used as the matching variable.
Can enhance the influence of regression to the mean, appear program failed.
Matching on pretest may produce:
• Experimental group scores higher than population on pretest
• Control group scores lower than population on pretest
• Both groups may regress to mean on posttest, masking any real change due to treatment
Example: improve the reading skills of disadvantaged children is effective.

For experimental group: regression to the mean + improvement = no change


These children produce higher score in their population (pretest)
For control group: only regression to the mean (higher in the end)
These children produce lower score in their population (pretest)
Only 1 pretest...
Example: head start
P208
Interrupted Time Series Designs
No control group -->take several measures before and after treatment, see trend
The number of pre‐interruption and post‐interruption points does not have to be the same
The more measures, the better
Example: antismoking campaign can decrease teenagers’ smoking behaviour

A: reduction just part of general trend, this design can rule out alternative explanation
B: briefly dropped, effect but short-lived
C: another general trend, periodic fluctuation
D: effective and long-lasting, ideal, Relative steady baseline-->rule out regression effects
P209
Research example: incentive plan on productivity

stay high after treatment


BUT: history may affect it?
P210
Variations on the Basic Time Series Design
1. Add a control group
combining the best features of the nonequivalent control group design (a control group) and the
interrupted time series design (long‐term trend analysis). e.g. divorce rate and bomb
2. Add a switching replication (interrupted time series with switching replications)
Different location, different time points, 2 experimental groups, same treatment
Experimental group1: O1 O2 O3 T O4 O5 O6 O7 O8 O9 O10
Experimental group2: O1 O2 O3 O4 O5 O6 O7 T O8 O9 O10
3. Add a DV that is not influenced by treatment, e.g. crime rate
Logic is similar to small N group (no control group (we can have different baseline...)
P212
Program evaluation: Applied research attempts to assess effectiveness and value of public policy
or specially designed programs.
(a) if a need exists for a program and who would benefit if the program is implemented;
Need assessment...
(b) assessments of whether a program is being run according to plan, formative evaluation
(c) methods for evaluating program outcomes, summative evaluation
(d) cost-effectiveness analyses to determine if program benefits justify the funds expended.
P213
A Note on Qualitative Data Analysis:
Qualitative analysis also involve in program evaluation
Needs analysis: in-depth interview info
Ethics:
Longitudinal: repeatedly contact, coding system to protect identities
Questions:
P6
Pseudoscience replies heavily on anecdotal evidence, they ignored the examples that don’t fit
their theories. Does it mean they have confirmation bias?
P24
Why field research is more expensive?
P23
What is the difference between ecological validity and mundane realism?
P29
What is the difference between theory and construct? Why working memory is a theory but not
construct? It is a construct
P44
If a test measure the degree of depression, and assume the people scores low on this test score
high on a test measure the happiness. Is this concurrent validity or convergent validity?
Convergent validity
P44
Validity assumes reliability, but the converse is not true?
P121
If anova shows non-significant, does it mean the post hoc analysis wont be significant?
P121
What is assumption of sphericity?
The sphericity assumption is satisfied when the variance of the difference between scores for any
two levels of a repeated measures factor is constant. The sphericity assumption is violated when
the variance of the difference between scores for any two levels of a repeated measures factor is
not constant (online source)

Homogeneity of variance means equal variances between independent groups, Levene's Test of
Equality of Variances is used to assess this statistical assumption.

P135

why the last one is “no interaction”? the lines are not parallel.
P135: is period of time a manipulated variable or subject variable? Manipulated variable
P142:
Is the matched groups factorial using the repeated measures ANOVA?
P146: if 2-levels, no need to check sphericity?
P151: a common open‐ended question is sometimes referred to as the “most important
problem” item.

P160
Scatterplot: Using parameters to determine how strong?
P167 how to partial correlation?
P169
By AI:
Yes, partial correlation can be used to examine the existence of a mediator but not a moderator.
Here's an explanation of the difference between mediation and moderation and how partial
correlation is used in each case:

Mediation: Mediation analysis tests the hypothetical causal chain where one variable (X) affects a
second variable (M), and in turn, that variable affects a third variable (Y). Mediators explain the
"how" or "why" of a relationship between two other variables and describe the process through
which an effect occurs [2]. In mediation analysis, partial correlation is used to examine the
relationship between X and Y while controlling for the mediator M. By "partialing out" the
association between X and M and between Y and M, the partial correlation identifies the
association between X and Y that is unrelated to M [1]. This allows researchers to determine the
direct effect of X on Y after accounting for the mediator M.

Moderation: Moderation analysis, on the other hand, tests for the influence of a third variable (Z)
on the relationship between X and Y. Moderators can strengthen, weaken, or reverse the nature
of the relationship between X and Y under different conditions [2]. Partial correlation is not
typically used to examine moderation because it focuses on controlling for the association
between variables rather than exploring the interaction effects between variables. In moderation
analysis, researchers often use interaction terms and plot simple slopes to test and visualize the
moderating effects [2].

In summary, partial correlation is commonly used in mediation analysis to examine the direct
effect of X on Y after accounting for the mediator M. However, it is not typically used in
moderation analysis, where interaction terms and simple slopes are more commonly employed
to explore the influence of a third variable on the relationship between X and Y.

P172
Why habituated to animals not participant observation?

P165
What is the relationship between beta coefficient and b value in equation Y= a+ bX?
Why multiple regression allow researchers to estimate the relative strength of each predictor?
The size of b’s are the beta coefficient that reflect the relative importance of each predictor, beta
weights..... 是不是 the larger the b, the more important this predictor to the criterion variable?
b 越大,说明 r 越大吗?

converging operations. The idea is that the various operational definitions are “converging” on
the same construct. When scores based on several different operational definitions are closely
related to each other and produce similar patterns of results, this constitutes good evidence that
the construct is being measured effectively and that it is useful. (by ai)

What is time sampling? What is event sampling?


Time Sampling Example:
Imagine a researcher studying the behavior of children during recess. The researcher decides to
use time sampling to observe the children's play behavior. They divide the recess period into 10-
minute intervals and observe the children's play behavior for 1 minute at the beginning of each
interval. During this 1-minute observation, the researcher records whether the children are
engaged in cooperative play, solitary play, or aggressive behavior [3].

Event Sampling Example:


Let's consider a researcher studying the occurrence of aggressive behavior in a classroom setting.
The researcher decides to use event sampling to observe and record instances of aggression.
Whenever a student displays aggressive behavior, such as hitting or yelling, the researcher
immediately records the event. This method allows the researcher to focus specifically on
aggressive incidents and gather data related to those specific events.

Time sampling involves observing and recording behaviors at specific time intervals, such as
observing children's play behavior during recess [3].
Event sampling involves observing and recording specific events or occurrences, such as instances
of aggressive behavior in a classroom setting

Meta-analysis: why need to find if effect size consistent? Direct replication or conceptual
replication is also ok?
P202
What is the relationship between quasi-experimental designs and a lot of ex post....
P204
(C) Can regression to the mean explain the change of experimental group?

What is time sampling? What is event sampling?

Meta-analysis: why need to find if effect size consistent? Direct replication or conceptual
replication is also ok?

为什么 study with small N designs usually don’t involve a control group?

Chapter 1:
Limitation of using logic: it can be used to reach opposing conclusions
a priori method for acquiring knowledge: A belief develops as the result of logical argument,
before a person has direct experience with the phenomenon at hand (a priori translates from the
Latin as “from what comes before”).

Confirmation bias often combines with another preconception called belief perseverance.
Motivated by a desire to be certain about one’s knowledge, it is a tendency to hold on doggedly
to a belief, even in the face of evidence that would convince most people that the belief is false.
It is likely that these beliefs form when the individual hears some “truth” being continuously
repeated, in the absence of contrary information.
Strongly held prejudices include both belief perseverance and confirmation bias.

Refusing to give up on a theory, in the face of a few experiments questioning that theory’s
validity, can have the beneficial effect of ensuring that the theory receives a thorough evaluation.
Thus, being a vigorous advocate for a theory can ensure that it will be pushed to its limits before
being abandoned by the scientific community.

scientific thinking includes elements of the nonscientific ways of knowing: authority, logic,
experience (bias)

the traditional concept of determinism, as used in science, contends simply that all events have
causes. Some philosophers have argued for a strict determinism, which holds that the causal
structure of the universe enables the prediction of all events with 100% certainty, at least in
principle. Most scientists, influenced by 20th‐century developments in physics and the
philosophy of science, take a more moderate view that could be called probabilistic or statistical
determinism. This approach argues that events can be predicted, but only with a probability
greater than chance. Research psychologists take this position and use this definition of
determinism in their science.

Free will, researchers can: (a) the extent to which behavior is influenced by a strong belief in free
will, (b) the degree to which some behaviors are more “free” than others (i.e., require more
conscious decision making), and (c) what the limits might be on our “free choices”.

Of course, in order to repeat a study, one must know precisely what was done in the original one.
This is accomplished by means of a prescribed set of rules for describing research projects.
These rules are presented in great detail in the Publication Manual of the American Psychological
Association (American Psychological Association, 2010), a useful resource for anyone reporting
research results or writing any other type of psychology paper.

If psychology was to be truly “scientific,” Watson argued, it needed to drop introspection and
measure something that was directly observable and could therefore be verified objectively (i.e.,
by two or more observers).

The problem with introspection was that although introspectors underwent rigorous training that
sought to eliminate bias in their self‐observations, the method was fundamentally subjective—I
cannot verify your introspections and you cannot verify mine.

Beliefs rooted in scientific methodology are always subject to change based on new data.

Falsification: theories must generate hypotheses producing research results that could come out
as the hypothesis predicts (i.e., support the hypothesis and increase confidence in the theory) or
could come out differently (i.e., fail to support the hypothesis and raise questions about the
theory).

research psychologists can be described as “skeptical optimists.” They are open to new ideas and
optimistic about using scientific methods to test these ideas, but at the same time they are
tough‐minded—they won’t accept claims without good evidence. Also, researchers are
constantly thinking of ways to test ideas scientifically, they are confident that truth will emerge by
asking and answering empirical questions, and they are willing (sometimes grudgingly) to alter
their beliefs if the answers to their empirical questions are not what they expected.

Even not scientist, All of us could benefit from using the attributes of scientific thinking to be
more critical and analytical about the information we are exposed to every day.

Hypothesis: a prediction about the study’s outcome, can be deducted from a theory

伪科学:
In some cases, the origins of a pseudoscience can be found in true science; in other instances, the
pseudoscience confuses its concepts with genuine scientific ones.
E. g. Phrenology, graphology
Phrenology originated in legitimate attempts to demonstrate that different parts of the brain
had identifiably distinct functions, and it is considered one of the first systematic theories about
the localization of brain function (Bakan, 1966). Phrenologists believed that (a) different
personality and intellectual attributes (“faculties”) were associated with different parts of the
brain (see Figure 1.3), (b) particularly strong faculties resulted in larger brain areas, and (c) skull
measurements yielded estimates of the relative strengths of faculties. By measuring skulls,
therefore, one could measure the various faculties that made up one’s personality.
Even if a theory is discredited within the scientific community, then, it can still find favor with the
public.
Graphology has an intuitive appeal, because handwriting styles do tend to be unique to the
individual, so it is natural to assume that the style reflects something about the person.
Advocates for graphology try to associate with true science in two different ways. First, there
is a fairly high degree of complexity to the analysis itself, with measurements taken of such
variables as slant, letter size, pen pressure, spacing between letters, etc. With actual physical
measurements being made, one gets the impression of legitimacy. After all, science involves
measuring things. Second, graphologists often confuse their pseudoscience with the legitimate
science of document analysis, performed by professionals called “questioned document
examiners. This latter procedure is a respected branch of forensic science, involving the
analysis of handwriting for identification purposes. That is, the document analyst tries to
determine whether a particular person wrote or signed a specific document. This is accomplished
by getting a handwriting sample from that person and seeing if it matches the document in
question. The analyst is not the least bit interested in trying to assess personality from the
handwriting. Yet graphologists sometimes point to the work of document examiners as a
verification of the scientific status of their field.
anecdotal evidence:
there may be some thieves with a particular skull shape, but in order to evaluate a specific
relationship between skull configuration and thievery, one must know (a) how many people who
are thieves do not have the configuration, and (b) how many people who have the configuration
aren’t thieves. Without having these two pieces of information, there is no way to determine if
there is anything unusual about a particular thief or two with a particular skull shape.
Anecdotal evidence involves using specific examples to support a general claim (they are also
known as testimonials); they are problematic because those using such evidence fail to report
instances that do not support the claim.
A law is a regularly occurring relationship.

Sidesteps the Falsification Requirement


As you learned earlier in this chapter, one of the hallmarks of a good scientific theory is that it is
stated precisely enough to be put to the stern test of falsification. In pseudoscience this does not
occur, even though on the surface it would seem that both phrenology and graphology would be
easy to falsify. Indeed, as far as the scientific community is concerned, falsification has occurred
for both. As you know, Flourens effectively discredited phrenology (at least within the scientific
community), and the same has occurred for graphology. Professional graphologists claim to have
scientific support for their craft, but the studies are inevitably flawed. For example, they typically
involve having subjects produce extensive handwriting samples, asking them to “write something
about themselves.” The content, of course, provides clues to the person. The graphologist might
also interview the person before giving the final personality assessment. In addition, the so‐
called
Barnum effect can operate. Numerous studies have shown that if subjects are given what they
think
is a valid personality test (but isn’t) and are then given a personality description of themselves,
filled with a mix of mostly positive traits, they will judge the analysis to be a good description of
what they are like. This occurs even though all the subjects in a Barnum effect study get the exact
same personality description, regardless of how they have filled out the phony personality test!
So
it is not difficult to imagine how a graphologist’s analysis of a person might seem fairly accurate
to that person. However, the proper study of graphology’s validity requires (a) giving the
graphologist several writing samples about topics having nothing to do with the subjects
participating in the
study (e.g., asking subjects to copy the first three sentences of the Declaration of Independence),
(b) assessing the subjects’ personality with recognized tests of personality that have been shown
to be reliable and valid (concepts you’ll learn more about in Chapter 4), and then (c) determining
whether the graphologist’s personality descriptions match those from the real personality tests.
Graphology always fails this kind of test (Karnes & Leonard, 1992).
Phrenologists sidestepped falsification by using combinations of faculties to explain the apparent
anomaly.

Another way that falsification is sidestepped by pseudoscience is that research reports in


pseudoscientific areas are notoriously vague and they are never submitted to reputable journals
with stringent peer review systems in place. As you recall, one of science’s important features is
that research produces public results, reported in books and journals that are available to
anyone. More important, scientists describe their research with enough precision that others can
replicate the experiment if they wish. This does not happen with pseudoscience, where the
research reports are usually vague or incomplete and, as seen earlier, heavily dependent on
anecdotal support.

Description also involves classification, as when someone attempts to classify various forms of
aggressive behavior.

A Passion for Research in Psychology


Ivan Pavlov, what it takes to be a great scientist. The first was to be systematic in the search for
knowledge and the second was to be modest and to always recognize one’s basic ignorance. The
third thing, was “passion.”
-Eleanor Gibson (1910–2002), developmental psychology, “visual cliff” studies. the unwillingness
of eight‐month‐olds to cross the “deep side”
-B. F. Skinner (1904–1990)
His work on operant conditioning created an entire subculture within experimental psychology
called the experimental analysis of behavior.
In both the Gibson and Skinner quotes, the concept of beauty appears.

Chapter 2 Ethics
Little Albert:
Despite serious methodological weaknesses and failed replication attempts.
Watson and Rayner made no attempt to remove the fear, although they made several
suggestions for doing so.
Little stimulations: slow in language development

Nuremberg Code of ethics (1949), which emphasized the importance of voluntary consent from
individuals involved in medical research. Nazi..Psychologists in the United States published their
first formal code of ethics in 1953 (APA, 1953), and it was influenced by the Nuremberg code.
The document was the outcome of about 15 years of discussion within the APA, which had
created a temporary committee on scientific and professional ethics in the late 1930s. This soon
became a standing committee to investigate complaints of unethical behavior (usually concerned
with the professional practice of psychology) that occasionally were brought to its attention. In
1948, this group recommended the creation of a formal code of ethics. As a result, the APA
formed a Committee on Ethical Standards for Psychologists, chaired by Edward Tolman (Hobbs,
1948).
Critical incidents: Although most concerned the practice of psychology (e.g., psychotherapy),
some of the reported incidents involved the conduct of research (e.g., research participants not
being treated well).
The first version of code: Although it was concerned mainly with professional practice, one of its
sections in this first ethics code was called “Ethical Standards in Research.”
Belmont:
The Civil Rights Act (1964) and the Voting Rights Act (1965) were passed by Congress and
contributed to growing movement toward civil rights for all Americans. In part, the civil rights
movement in American culture emphasized basic equality and dignity among individuals and
created a heightened awareness of instances of inequality and mistreatment of vulnerable
groups. One such group whose plight had finally come to light in the 1970s included poor Black
men from the area around Tuskegee, Alabama, who were diagnosed with syphilis, but
deliberately left untreated so that researchers could study the development of the disease over
time (see Box 2.2). The revelation of the Tuskegee study in the early 1970s is one factor that led
to the United States Congress to enact the National Research Act in 1974, which created the
National Commission for the Protection of Human Subjects of Biomedical and Behavioral
Research. In 1979, the commission published what came to be called the Belmont Report, which
includes three basic principles for research with human subjects: Respect for persons,
Beneficence, and Justice.
There are many similarities between the principles of the Belmont Report and those of the APA
Code, including some very similar language.

APA ethics code has been revised several times, most recently in 2002. It currently includes a set
of 5 general principles and 89 standards, the latter clustered into the 10 general categories. The
general principles are “aspirational” in their intent, designed to “guide and inspire psychologists
toward the very highest ideals of the profession” (APA, 2002, p. 1062), while the standards
establish specific rules of conduct and provide the basis for any charges of unethical conduct.
The five general principles reflect the philosophical basis for the code as a whole.

Ethical Guidelines for Research with Humans: In the 1960s, a portion of the original ethics code
was elaborated into a separate code of ethics designed for research with human participants.

As you might guess, there are gray areas concerning decisions about exempt, expedited, and
full review. Hence, it is common practice for universities to ask that all research be given some
degree of examination by the IRB. Sometimes, different members of an IRB are designated as
“first step” decision makers; they identify those proposals that are exempt, grant approval (on
behalf of the full board) for expedited proposals, and send on to the full board only those
proposals in need of consideration by the entire group. At medium and large universities, where
the number of proposals might overwhelm a single committee, departmental IRBs are sometimes
created to handle the expedited reviews

Limitations of IRB
...

Historical cases of lack of inform consent:


brief descriptions of cases in which (a) children with severe intellectual disabilities were infected
with hepatitis in order to study the development of the illness; (b) southern Black men with
syphilis were left untreated for years and misinformed about their health, also for the purpose of
learning more about the time course of the disease; and (c) Americans, usually soldiers, were
given LSD without their knowledge.

First, potential volunteers agree to participate after learning the general purpose of the study
(but not the specific hypotheses), the basic procedure, and the amount of time needed for the
session. Second, participants understand they can leave the session at any time without penalty
and with no pressure to continue. A third major feature of consent forms is that participants are
informed that strict confidentiality and anonymity will be upheld. Fourth, if questions linger
about the study or if they wish to complain about their treatment as participants, there are
specific people to contact, including someone from the IRB. Finally, participants are informed of
any risk that might be encountered in the study, and they are given the opportunity to receive a
summary of the results of the study, once it has been completed. When writing a consent form,
researchers try to avoid jargon, with the aim of making the form as easy to understand as
possible.

A new feature of the 2002 revision of the ethics code is a more detailed set of provisions for
research designed to test the effectiveness of a treatment program that might provide benefits
but might also be ineffective and perhaps even harmful (Smith, 2003)—a program to treat post‐
traumatic stress disorder, for instance. This revision is found in Standard 8.02b, which tells
researchers to be sure to inform participants that the treatment is experimental (i.e., not shown
to be effective yet), that some specific services will be available to the control group at the end of
the study, and that services will be available to participants who exercise their right to withdraw
from the study or who choose not to participate after reading the consent form.

the Society for Research in Child Development (SRCD) follows a set of guidelines that expand
upon some of the provisions of the code for adults
assent occurs when “the child shows some form of agreement to participate without necessarily
comprehending the full significance of the research necessary to give informed consent” (SRCD,
1996, p. 337). Assent also means the researcher has a responsibility to monitor experiments with
children and to stop them if it appears that undue stress is being experienced.
In addition to the assent provision, the SRCD code requires that additional consent be obtained
from others who might be involved with the study in any way. For example, this would include
teachers when a study includes their students.
researchers should not use the potential rewards as an inducement to gain the child’s assent.
legal guardians must give truly informed consent for research with people who are confined to
institutions (e.g., the Willowbrook case).
cookies (tools used to track information about Internet users) are not
left on the participant’s computer as a result of taking the survey.

Use of animals, reasons:


1. Genetic and life‐span developmental studies can take place quickly
2. animals can be subjected to procedures that could not be used with humans

In “The Value of Behavioral Research on Animals”, Miller (1985) argued that (a) animal activists
sometimes overstate the harm done to animals in psychological research, (b) animal research
provides clear benefits for the well‐being of humans, and (c) animal research benefits animals as
well.

The Animal Welfare Act (AWA) enacted in 1966 is the only federal law in the United States that
regulates the treatment of animals used in research. Part of the AWA’s mandate is that
institutions where animal research is conducted should have an Institutional Animal Care and Use
Committee (IACUC). Like an IRB, the IACUC is composed of faculty from several disciplines in
addition to science, a veterinarian, and someone from outside the university.
Often, the IACUC will use guidelines put forth in the Guide for the Care and Use of Laboratory
Animals (National Research Council, 2011) in its evaluation of the ethical treatment of animals in
research. In addition, psychologists rely on Standard 8.09 of the 2002 APA ethics code, which
describes the ethical guidelines for animal care and use.

Justifying the Study


The scientific purpose of the study should fall within one of four categories. The research should
“(a) increase knowledge of the processes underlying the evolution, development, maintenance,
alteration, control, or biological significance of behavior, (b) determine the replicability and
generality of prior research, (c) increase understanding of the species under study, or (d) provide
results that benefit the health or welfare of humans or other animals”

Data falsification
the pressure to “publish or perish” overwhelms the individual and leads the researcher (or the
researcher’s assistants) to cut some corners.

Chapter 3
a major advantage of basic research is that the principles and procedures (e.g., shadowing)
developed through basic research can potentially be used in a wide range of applied situations,
even though these uses may not have been considered when the basic research was being
done.
IRBs sometimes favor applied over basic research; nonpsychologist members of an IRB in
particular often fail to see the relevance of basic laboratory procedures.
the lab‐field correspondence is higher in some areas than others. For example, Mitchell (2012)
reported a high degree of similarity between laboratory and field studies in
industrial/organizational psychology and in personality psychology, but somewhat lesser
similarity in social psychology and consumer psychology, and little similarity in developmental
psychology.
It can happen when a scientist is wrestling with a difficult research problem and a chance event
accidentally provides the key, or it might occur when something goes wrong in an experiment,
such as an apparatus failure. Skinner’s experience with extinction curves following an
apparatus breakdown, described in Chapter 1, is a good example of a serendipitous event.
Another involves the accidental discovery of feature detectors in the brain.

The Clever Hans case illustrates two other points besides the falsification strategy of Pfungst. By
showing the horse’s abilities were not due to a high level of intelligence but could be explained
adequately in terms of the simpler process of learning to respond to two sets of visual
cues (when to start and when to stop), Pfungst provided a more parsimonious explanation of the
horse’s behavior.
Second, if von Osten was giving subtle cues that influenced behavior, then perhaps
experimenters in general might subtly influence the behavior of participants when the
experimenter knows what the outcome will be. --experimenter bias.

Parsimony
Lloyd Morgan’s Canon,” was that “[i]n no case may we interpret an action as the outcome of the
exercise of a higher psychical faculty, if it can be interpreted as the outcome of the exercise of
one which stands lower in the psychological scale”.
The need for parsimonious explanations also guards against one of the social cognition biases
described in Chapter 1—confirmation bias. Hearing an interesting anecdote about a dog that
appears to be using logical reasoning, the listener with a preconceived bias about dog
intelligence
might see this as a confirming example of dog brilliance, while ignoring other instances in which a
dog’s behavior did not seem so smart. The point was made nicely by another animal researcher,
Edward Thorndike who stated, somewhat sarcastically.

Given recently noted cases of blatant scientific fraud in psychology, such as that by Diederik
Stapel discussed in Chapter 2, it is important to also notice that there are other forms of scientific
misconduct that may go unnoticed, referred to as Questionable Research Practices (QRPs) by
John, Loewenstein, and Prelec (2012). In the wake of cases of scientific fraud, concerns about
QRPs are mounting across disciplines, including medical and psychological research. John et al.
conducted a study in which they anonymously surveyed over 2,000 psychologists about ten
different types of QRPs ranging in severity from not reporting all measures used in a study, to
“selectively reporting studies that ‘worked,’” to outright falsifying data

Creative thinking
a thorough knowledge of one’s field may be a prerequisite to creative thinking in science, but the
blade is double‐edged. Such knowledge can also create rigid patterns of thinking that inhibit
creativity. Scientists occasionally become so accustomed to a particular method or so
comfortable with a particular theory that they fail to consider alternatives, thereby reducing the
chances of making new discoveries.
focus on an existing apparatus can limit creative thinking in science. The origins of scientific
equipment such as mazes may reveal creative thinking at its best (e.g., Sanford’s idea to use the
Hampton Court maze), but innovation can be dampened once an apparatus or a research
procedure becomes established.

What is the mixed mode approach in conducting survey?


Combining methods of survey delivery, such as sending an email or mailing a letter prior to a
phone survey.

linear regression there is one predictor variable, whereas in multiple regression, there is more
than one predictor variable; in both, there will be just one criterion variable.

the psychologist doing field research faces serious risks that do not occur in the laboratory.

Research Example 32—Meta-analysis and Psychology’s


First Registered Replication Report (RRR)
Recall from Chapter 3 the crucial role of replication in psychological science. Direct replications
are exact reproductions of research studies, using similar samples and essentially the same
procedures. Conceptual replications change some procedures and/or samples to confirm a
finding but also to extend it in some fashion (e.g., show that the original finding also applies to a
different population). In general, successful replication allows researchers to be more confident in
their results. Furthermore, failures to replicate can signal the possibility of scientific fraud or the
kinds of questionable research practices (QRP’s) described in Chapter 3. Recognizing that direct
replications are not always popular with researchers—new findings have a better chance of being
published—the Association for Psychological Science (APS) began an initiative in 2013 to support
and fund direct replications of well‐known findings. These replications would then be combined,
using meta‐analysis, with the final product to be known as a Registered Replication Report (RRR).
Registered Replication Reports (RRRs) and meta-analysis are closely related in the context of
scientific research. RRRs often include a meta-analysis as part of their methodology and
reporting. (BY AI)
APS’s first RRR was based on a 1990 finding by Schooler and Engstler‐Schooler that examined
accuracy in eyewitness identification. Imagine you witness a crime, and police then ask you
to provide a verbal description of the perpetrator. Later, you try to identify the perpetrator from a
lineup. Your chances of making the correct identification will depend on a number of factors, but
the fact that you first described the perpetrator verbally can actually reduce the accuracy of
identifying that person in a lineup. This phenomenon was discovered by Schooler and Engstler‐
Schooler and is known as the verbal overshadowing effect. The basic idea is that when individuals
first verbally describe a face, they then relied on their recollection of their verbal description,
rather than their visual memory for the face. Thus, the verbal description overshadowed their
visual memory. In their initial study, Schooler and Engstler‐Schooler reported that individuals
were 25% worse at identifying the culprit from a lineup if they first verbally described him than
if they did not. The finding was a surprise and has some obvious practical implications for police
procedures. However, given that the result was unexpected, combined with the fact that sample
sizes were small in the original research, Schooler and others recognized the need for replication,
which indeed occurred over the next several years.
About a decade later after the original publication, Meissner and Brigham (2001) published a
meta‐analysis of 29 studies on verbal overshadowing. The effect was replicated, but it did not
seem to be as strong as originally thought (i.e., an estimated 12% impairment in lineup
identification instead of 25%). The studies included in Meissner and Brigham’s meta‐analysis,
however, varied in the procedures used in testing the verbal overshadowing effect.
That is, Meissner and Brigham conducted a meta‐analysis of conceptual replications of
overshadowing. But what would happen to the overshadowing effect if direct replications were
attempted? Enter APS and its first [Link] in 2013, APS funded direct replications of the
original Schooler and Engstler‐Schooler (1990) research. There were two replication projects,
RRR1 and RRR2, which varied slightly in the timing of the procedures (Alonso et al., 2014). Our
focus will be on RRR2, which duplicated the methodology of the first of six experiments
completed by Schooler and Engstler‐Schooler the experiment that yielded the 25% drop in
accuracy. Figure 10.3 outlines the sequence of events in the procedure.
As you can see, participants viewed the crime, had a 20-minute filler task, spent
5 minutes writing a verbal description of the criminal, and then tried to identify the criminal out
of an 8‐person lineup. For both replication projects, samples sizes were greater than those used
by Schooler and Engstler‐Schooler A total of 22 different labs conducted direct replications of the
original Schooler and Engstler‐Schooler experiments. The results of the 22 different studies were
subjected to a meta‐analysis, and the overall outcome was that the verbal overshadowing effect
did indeed occur, although it was not quite as robust as in the original Schooler and Engstler‐
Schooler (1990) study.
Instead of a 25% reduction in accuracy, the reduction for all the studies combined was 16%
(Alonga et al., 2014). Although the effect was not as large in the replication, given that the effect
still consistently occurred across multiple replications, it can be concluded that the verbal
overshadowing effect has strong empirical support.
the APS replication project is that it addresses a problem, the file drawer effect. all the
participating labs were told their studies would be published regardless of the outcome. Hence,
file drawer effects cannot occur in the published RRRs.

Thematic analysis is a qualitative research method used to identify, analyze, and interpret
patterns or themes within qualitative data. It involves systematically organizing and categorizing
data to uncover meaningful patterns, concepts, or ideas that emerge from the data [1].
Coding in thematic analysis is the process of assigning labels or codes to segments of data that
represent different themes or patterns. It is a crucial step in the analysis process as it helps to
organize and categorize the data into meaningful units. Coding allows researchers to identify and
analyze the recurring ideas, topics, or concepts within the data.
There are different approaches to coding in thematic analysis, including:

1. Open Coding: In this approach, the researcher starts with an open mind and examines the data
to identify initial codes or labels that capture the essence of the data. The codes are generated
directly from the data without any preconceived categories or themes.

2. Axial Coding: This approach involves organizing the initial codes into broader categories or
themes. The researcher looks for relationships and connections between the codes to develop a
more comprehensive understanding of the data.

3. Selective Coding: In this final stage of coding, the researcher focuses on refining and
consolidating the themes. The researcher selects the most relevant and significant codes and
develops a coherent narrative or explanation of the data.
By ai:
Moderated mediation analysis is a statistical technique used to examine the relationship between
an independent variable, a mediator, and a dependent variable, while also considering the
influence of a moderating variable. It allows researchers to investigate how the mediation
process might differ for different levels or conditions of the moderating variable.

In a typical mediation analysis, the researcher examines whether the effect of the independent
variable on the dependent variable is mediated through the mediator variable. However, in
moderated mediation analysis, the researcher explores whether the strength or direction of the
indirect effect (mediation) varies depending on the level or condition of the moderating variable.

Chapter 12
a 36‐year‐old fighter (“A.B.”) who had suffered repeated head trauma during his 15 years of
boxing,
A.B. had been working as a physical education instructor in a substance abuse rehabilitation
center. He was referred for assessment to a speech pathologist because his speech was often
slurred and inarticulate; by his own admission he sounded drunk, which was affecting his
credibility at the rehab center. His supervisor described him as an excellent worker, but his
speech was
a problem. Because repeated brain trauma is known to produce cognitive impairment, A.B. was
first given a battery of cognitive tests that assessed both short‐term and long‐term memory,
judgment, spatial cognition, reasoning, and problem solving. He scored at the 91st percentile or
higher on all the tests—so, no cognitive impairment. A.B. was then given several tests for specific
language impairment having to do with comprehension; he passed these as well. Physical
dexterity was also screened by using a field sobriety test (e.g., touch your finger to your nose). No
problem here either. Note the use of falsification thinking here—trying to arrive at a diagnosis by
ruling out, or disconfirming, alternatives.
The problems for A.B. occurred with tests of articulation. For example, he had trouble producing
the normal rhythm in a sentence; he would stress the wrong words or the wrong syllables
in a word. He also failed to make clear pauses between words, thereby producing the slurred
speech that led others to think he was inebriated. He was diagnosed with ataxic dysarthria.
Ataxic generally refers to a lack of coordination in muscle movements, while dysarthria is a
broad term referring to articulation failure. Ataxic dysarthria, articulation failure due to lack of
control over the motor components of articulation, is usually thought to result from damage in
the area of the cerebellum, a form of damage that could easily result from the sharp head
twisting motions that accompany a sharp blow to the head. And A.B. had lots of experience with
sharp blows to the head.
Having identified the most likely diagnosis, speech therapists then developed an articulation
training program for A.B. that got him to focus on individual words and to increase the volume
of his speech. Another component of the therapy was a form of operant shaping: Training
started with individual words, then very brief sentences, and then longer, more typical sentences.
At the end of treatment, he had improved significantly on a variety of measures. For
instance, one test measured “perceptual intelligibility” (how his speech was understood by
others) on a scale from 1 (unintelligible) to 7 (perfectly intelligible). A.B. scored 3.7 during the
assessment period and 5.3 after treatment. The brain damage would probably prevent him from
ever scoring a 7, but by becoming more “mindful” of how to articulate words and how to build
normal rhythm into sentences, A.B. showed significant improvement and was able to return
successfully to work. Nice outcome!
You will notice that the case study of A.B. involved a case history, inclusion of various
psychological and neurological tests, the implementation of a treatment (i.e., speech therapy),
and an assessment of the effectiveness of the treatment. The comprehensive nature of case
studies allows both researchers and practitioners to effectively use data to understand and treat
individuals in need of help. This harkens back to our earlier discussion in Chapters 1 and 11
about translational research, whose goal is to transform information to improve physical and
psychological well‐being.

S:
Case histories often document lives that are classic examples of particular psychological types. In
abnormal psychology, for example, the case study approach is often used to understand the
dynamics of specific disorders by detailing typical examples of them. Case studies also can be
useful in experimental psychology, however, shedding light on basic psychological phenomena. A
classic example is the one compiled by Alexander Romanovich Luria (1902–1977), a Russian
scientist famous for his studies of Russian soldiers who were brain‐injured during World War II
and for his work on the relationship between language and thought (Brennan, 1991).
The case involved one S. V. Sherashevsky, or “S.,” as Luria referred to him, whose remarkable
memory abilities gave him a career as a stage mnemonist (yes, people actually paid to watch him
memorize things) but also caused him considerable psychological distress. The case is
summarized in Luria’s The Mind of a Mnemonist (1968). Luria studied S. for more than 20 years,
documenting both the range of S.’s memory and the accompanying problems of his being
virtually unable to forget anything. Luria first discovered there seemed to be no limit on how
much information S. could memorize; more astonishing, the information did not seem to decay
with the passage of time. He could easily memorize lists of up to 70 numbers and could recall
them in either a forward or a reverse order. Also, “he had no difficulty reproducing any lengthy
series . . . whatever, even though these had been presented to him a week, a year, or even many
years earlier” (Luria, 1968, p. 12).
This is an unbelievable performance, especially considering that most people cannot recall more
than seven or eight items on this type of task and that forgetting is the rule rather than the
exception. As a student, you might be wondering what the downside to this could possibly be.
After all, it would seem to be a wonderful problem to have, especially during final exam week.
Unfortunately, S.’s extraordinary memory skills were accompanied by severe deficits in other
cognitive areas. For example, he found it almost impossible to read for comprehension. This was
because every word evoked strong visual images from his memory and interfered with the overall
organization of the ideas conveyed by the sentences.
Similarly, he was an ineffective problem solver, found it difficult to plan and organize his life, and
was unable to think abstractly. The images associated with his remarkable memory interfered
with everything else.
Is S. anything more than an idle curiosity, a bizarre once‐in‐a‐lifetime person who doesn’t really
tell us anything about ourselves? No. He was indeed a very rare person, but the case sheds
important light on normal memory functioning.
In particular, it provides a glimpse into the functional value forgetting from short‐term memory.
We sometimes curse our inability to recall something we were thinking about just a few minutes
before, but the case of S. shows that forgetting allows us to clear the mind of information that
might be useless (e.g., there’s no reason to memorize all the items on the menu we just read in a
restaurant) and enables us to concentrate our energy on more sophisticated cognitive tasks such
as reading for comprehension. Because S. couldn’t avoid remembering everything he
encountered, he was unable to function at higher cognitive levels.
One final point: It turns out S. was not a once‐in‐a‐lifetime case. Another person (“V.P.”) with a
similarly remarkable memory was studied by the American psychologists Hunt and Love (1972).
Oddly enough, V.P. grew up in a city in present‐day Latvia that was just a short distance from the
birthplace of Luria’s S.

What is the difference between case of boxer and S?


The boxer is an example of case involving a common problem, while S. illustrates an extremely
rare case (that nonetheless sheds light on normal behavior).

“facial vision,” the ability to detect the presence of nearby objects even when they cannot be
seen. At one time, blind people were believed to have developed this as a special sense to
compensate for their loss of vision.

a confounding variable co-varies with IV and could influence DV.

A two-way mirror, also known as a one-way mirror or a two-sided mirror, is a special type of glass
that allows for unidirectional visibility. It is commonly used in observational research settings to
observe people or subjects without them being aware that they are being observed.

Classical conditioning:
It involves the process of learning through the association of a neutral stimulus with a biologically
significant stimulus to elicit a reflexive response.
Unconditioned Stimulus (UCS), Unconditioned Response (UCR), Conditioned Stimulus (CS),
Conditioned Response (CR)

Operant Conditioning:
It focuses on how behavior is shaped and modified by the consequences that follow it. Operant
conditioning involves learning through the association between voluntary behaviors and their
consequences.
Example in reporting interaction:
"A two-way mixed ANOVA revealed a significant interaction effect between gender and
temperature on memory accuracy, F(1, 50) = 6.82, p = 0.012. The interaction effect indicated that
the impact of temperature on memory accuracy differed between genders. Post hoc tests
revealed that for males, memory accuracy was significantly higher at higher temperatures (M =
85.2, SD = 3.7) compared to lower temperatures (M = 79.6, SD = 4.1), t(25) = 3.89, p = 0.001.
However, for females, there was no significant difference in memory accuracy between higher
temperatures (M = 81.1, SD = 3.9) and lower temperatures (M = 80.3, SD = 3.5), t(25) = 0.81, p =
0.426."

You might also like