Research Methodology Overview
Research Methodology Overview
Ilide - Rfuhvhjhb
Purpose of Research: The primary purpose of research is to advance understanding and solve problems. It
can be applied (solving practical issues) or basic (expanding theoretical knowledge). As one source notes,
basic research “advances fundamental knowledge” and often generates new ideas or refines theories 3 .
Applied research, in contrast, aims to solve specific real-world problems (e.g. improving educational
methods or treatments). Overall goals include generating evidence, testing theories, and informing practice
or policy. For example, clinical research applies research to improve patient care, while basic psychological
research might investigate underlying cognitive processes 3 1 .
• Purpose/Use: Basic vs. Applied. Basic research (also called pure or fundamental) seeks to expand
knowledge without immediate practical applications 3 . Applied research targets specific practical
problems, often using theories to improve practice.
• Design/Purpose: Exploratory vs. Descriptive vs. Explanatory. Exploratory research investigates new or
not well-understood phenomena (e.g. pilot studies, case observations). Descriptive research aims to
describe characteristics of variables (e.g. surveys reporting prevalence). Explanatory (analytic)
research tests hypotheses about causal relationships 1 4 .
• Time Dimension: Cross-sectional vs. Longitudinal. Cross-sectional research collects data at one point
in time (a “snapshot”), whereas longitudinal studies collect data over extended periods to detect
change.
• Methodological: Qualitative vs. Quantitative (discussed later), and Experimental vs. Non-experimental.
Experimental manipulates variables under controlled conditions; non-experimental (e.g.
correlational) observes variables without manipulation.
• Scope/Discipline: Research may also vary by discipline (e.g. clinical, social, cognitive) and scope
(single study vs. programmatic research).
In summary, research is a deliberate, systematic process aiming to answer questions or test ideas. It serves
purposes from generating new theories (basic research) to solving applied problems. Its dimensions include
its goals, design features (time, control), and methodological orientation 3 1 .
therapy reduce anxiety levels in college students?”). A problem statement identifies what is unknown or
questionable and frames the study’s focus 5 . Good research problems are clear, specific, and
researchable, often pointing to the relationships or differences to be investigated. For instance,
research guides define a problem as “an area of concern, difficulty to be eliminated… or a troubling
question” in theory or practice 5 .
• Variables: A variable is any characteristic or factor that can vary across individuals or situations.
Variables can be quantitative (measured numerically, e.g. height, IQ score) or categorical/
qualitative (measured by categories, e.g. gender, type of therapy) 6 7 . In hypotheses, variables
often take the roles of independent (predictor) and dependent (outcome) variables. For example, in an
experiment testing a drug, the dosage is an independent variable and the health outcome is a
dependent variable 8 . It is essential to explicitly define each variable.
• Hypotheses: A hypothesis is a tentative, testable prediction about the expected outcome of a study.
It usually states the expected relationship between variables, derived from theory or prior research.
For example: “College students who receive cognitive-behavioral therapy will report lower anxiety scores
than those who receive no therapy.” Hypotheses often come in null (H₀) and alternative (H₁) forms. The
null hypothesis typically asserts no effect or no difference (e.g. “There will be no difference in anxiety
levels between groups”), whereas the alternative hypothesis asserts the expected effect (e.g.
“Therapy will reduce anxiety”). In significance testing, we use statistical tests to assess whether data
allow us to reject H₀ in favor of H₁. As one guide notes, H₀ essentially answers “No effect in the
population,” while H₁ posits “Yes, there is an effect” 10 . Hypotheses can be directional (specifying a
direction, e.g. higher or lower) or non-directional. They form the basis for designing analyses (e.g.
choosing t-tests if comparing means).
• Sampling: Sampling refers to selecting a subset of individuals from a larger population for study.
The population is the full set of cases the researcher is interested in (e.g. all college students in
India). A sample is a subset chosen for the actual study. Sampling methods are crucial:
• Probability sampling ensures each member of the population has a known, usually equal, chance of
selection. Examples include simple random sampling, stratified sampling, cluster sampling. For
instance, simple random sampling might involve drawing names from a hat, giving every individual
an equal chance 11 . Proper random sampling helps generalize findings to the population.
• Non-probability sampling does not use random selection; it includes methods like convenience
sampling (selecting whoever is easiest to reach) or purposive sampling. For example, recruiting
nearby students as participants is convenience sampling 12 . This can introduce bias, as the sample
may not represent the population. For instance, convenience samples tend to over-represent
accessible or volunteer individuals and under-represent harder-to-reach groups 12 .
• Sample size is also critical: too small a sample may lack power or representativeness; very large
samples cost more. Researchers often use statistical formulas or power analysis to determine
adequate sample sizes (see below).
• Sampling bias (like selection bias) occurs if the sample systematically differs from the population.
Researchers must plan carefully to minimize such biases, or at least acknowledge them.
Citations for this section include definitions of variables and operational definitions 6 9 , explanation of
null/alternative hypotheses 10 , and sampling concepts 11 12 .
• Informed Consent and Voluntary Participation: Participants must be given full information about
the study’s purpose, procedures, risks, benefits, and funding, and must voluntarily agree to
participate 13 . They should know they can withdraw at any time. Coercion or undue inducement
must be avoided.
• Anonymity and Confidentiality: Data should be collected and stored so that participants’ identities
remain unknown (anonymity) or not linked to data (confidentiality). Anonymity means even
researchers cannot identify individuals from data. Confidentiality means researchers know identities
but promise not to reveal them. Both protect privacy. For example, subjects might be assigned codes
instead of using names. As one ethics guide notes, anonymity means “you don’t know the identities of
those in your study”, and confidentiality means “data from participants are kept secret” 14 .
• Avoiding Harm: Researchers must minimize potential physical, psychological, or social harm. This
includes minimizing stress, embarrassment, or other negative effects. For example, vulnerable
populations (e.g. prisoners, children) require additional protections. Debriefing (explaining purpose
after study) is often required, especially if deception was used.
• Integrity in Data Collection and Reporting: Researchers must not fabricate, falsify, or selectively
report data. According to official definitions, research misconduct includes “fabrication, falsification,
or plagiarism” in proposing, performing, or reporting research 15 16 . Fabrication means making up
data, falsification means manipulating data/analysis, and plagiarism means using others’ ideas or
words without credit 15 16 . For example, one source quotes U.S. policy that misconduct is
“fabrication, falsification, or plagiarism in proposing, performing, or reviewing research” 15 . These are
serious ethical violations.
• Ethical Reporting: Apart from honesty in data, researchers should report findings fully and
accurately, including non-significant results and study limitations. They must credit collaborators and
prior work properly (avoid self-plagiarism and duplicate publication). Authorship should reflect
actual contributions.
• Regulatory and Institutional Oversight: Research involving humans typically requires approval by
an Institutional Review Board (IRB) or ethics committee, which evaluates risks and consent
procedures. International guidelines (e.g. Declaration of Helsinki, APA Ethics Code) provide
frameworks. Researchers must follow relevant laws and codes.
Overall, ethical research is characterized by respect for persons (autonomy), beneficence (maximizing
benefits/minimizing harm), and justice (fair participant selection). Maintaining high ethical standards
protects participants and lends credibility to the scientific enterprise 13 15 .
Qualitative Research: In contrast, qualitative research deals with non-numerical data (words, images,
observations) to understand meanings, experiences, or social processes. It is typically inductive, generating
theories from data, and emphasizes context, depth, and participants’ perspectives. Methods include
interviews, focus groups, ethnography, narrative analysis, case studies, etc. Qualitative research allows deep
exploration of phenomena that are not well understood or are complex (e.g. subjective experiences). Its
advantages are rich, detailed data and flexibility. Limitations include less ability to generalize due to typically
small, non-random samples, and reliance on researcher interpretation (potential subjectivity). As one source
summarizes, qualitative research “is expressed in words… used to understand concepts, thoughts or
experiences… and gather in-depth insights” 18 .
Mixed Methods Research: This paradigm combines quantitative and qualitative approaches in a single
study (or set of studies) to leverage the strengths of both. Mixed methods can be sequential (e.g. qualitative
exploration followed by quantitative testing) or concurrent. It provides a more comprehensive
understanding. For example, a researcher might measure outcomes quantitatively and also interview
participants for context. According to NIH guidance, “mixed methods strategically integrates or combines
rigorous quantitative and qualitative research methods to draw on the strengths of each… to gain a more
complete understanding of research problems” 19 . Another source defines it simply as using “both qualitative
and quantitative data collection and analysis to answer a research question” 20 . Advantages include richer
insights and the ability to validate findings across methods. Challenges include increased complexity, need
for multiple skills, and integration of findings.
5. Methods of Research
Observation
Observation involves systematically watching and recording behavior in natural or controlled settings. It is
non-experimental (no manipulation of the independent variable) and can be either naturalistic
(unobtrusive recording in a natural setting) or participant (researcher joins the group or situation) 21 22 .
Observational data can be qualitative (notes) or quantitative (tallied counts, categories).
- Naturalistic Observation: Conducted in real-life settings without intervention. It offers high ecological
validity (real-world relevance) 21 . Example: Jane Goodall’s observation of chimpanzees in the wild. However,
limitations include lack of control over extraneous variables and difficulty inferring causality.
- Participant Observation: The researcher actively engages in the setting (often covertly or overtly) 22 .
This can yield deeper insight into the context and subjects’ perspectives. For example, a researcher living in
a community to study cultural practices. A major drawback is potential observer bias (the researcher’s
presence may influence behavior, known as reactivity) 23 , and ethical concerns if covert. As noted, one
benefit is that “the researcher can better understand the viewpoint of participants,” but a limitation is that “the
presence of an observer could affect behavior” 23 .
- Structured vs. Unstructured Observation: In structured observation, the researcher identifies specific
behaviors or categories and records them in a systematic way 24 . This yields quantitative data (e.g.
number of instances of a behavior). In unstructured observation, the researcher makes open-ended notes of
whatever seems relevant. Structured observation increases reliability and focus, but may miss unexpected
phenomena. Unstructured is more exploratory, but harder to analyze.
- Other Variants: Observation can be disguised (subjects unaware of being observed) or undisguised.
Disguised observation reduces observer effect but raises ethical issues.
Key points: Observational methods are excellent for describing behaviors as they naturally occur. They
cannot establish causality, because no experimental manipulation or random assignment is involved 21 .
However, they can generate hypotheses for later testing.
Surveys collect information by asking participants questions. They are used to gather self-reported data on
attitudes, opinions, behaviors, or attributes from a sample. Surveys can be administered via paper/online
questionnaires or through interviews (face-to-face, telephone, or online).
• Interviews: Can be structured (fixed set of questions in a set order, like a questionnaire read aloud),
semi-structured (guide with topics but flexibility in wording/order), or unstructured (open-ended,
conversational). Structured interviews ensure consistency across respondents, aiding comparability.
Semi-structured allow exploring interesting responses while maintaining focus. Unstructured
interviews (depth interviews) allow free expression and discovery, but are time-consuming and
harder to analyze quantitatively. For example, a structured survey might ask, “On a scale of 1–5, how
stressed do you feel?”, whereas an unstructured interview might simply ask “How do you feel about the
current semester?” and probe as needed 25 .
• Questionnaires: Standardized sets of written questions (often closed-ended with fixed response
options, sometimes with a few open-ended items). Advantages include ease of administration to
large samples and anonymity. Response rates can be an issue, and fixed responses may not capture
nuance. Well-constructed questionnaires are tested for reliability and validity.
• Focus Groups: These are group interviews where 6–10 participants discuss topics led by a
moderator. Focus groups elicit interaction and dynamics (participants may spark ideas in each other).
They are often used in exploratory research or to develop survey items. For example, exploring
consumer opinions on a product concept. As defined in research guides, “a focus group is a research
method that brings together a small group of people to answer questions in a moderated setting” 26 .
Advantages include gathering a range of views quickly; limitations include possible domination by
outspoken individuals and moderator bias.
Survey and interview methods allow gathering large amounts of data relatively efficiently. They rely on self-
report, so issues include social desirability bias and recall error. They assume respondents interpret
questions similarly.
Experimental Research
In experimental studies, the researcher manipulates one or more independent variables and observes the
effect on a dependent variable, while controlling extraneous factors. The hallmark of experiments is
random assignment of participants to conditions, which helps ensure that groups are comparable and
allows causal inference. Key features:
- Manipulation: The researcher chooses at least two levels of the IV (e.g. drug vs. placebo, or different
teaching methods).
- Control: Through experimental control (e.g. lab conditions, standardized procedures), other variables are
held constant.
- Random Assignment: Participants are assigned to conditions by chance, so confounding variables are
(theoretically) evenly distributed. This is why experiments allow causal conclusions: if groups differ on
outcome, the manipulated IV is inferred as the cause 27 .
- Design Variations: Simple designs include one IV with two conditions (treatment vs. control). More
complex designs include factorial experiments (multiple IVs, see below).
For example, to test a new teaching method, students could be randomly assigned to either the new
method or a standard method (control), and their test scores compared. According to an open-text source,
the experimental method “is the only method that allows causal conclusions” 27 . Indeed, experiments
typically precede claims of cause and effect in psychology.
Limitations: Experiments can lack external validity (artificiality of lab settings), and some research questions
cannot be ethically or practically manipulated. They also assume randomization is feasible.
Quasi-Experimental Research
Quasi-experiments resemble experiments but lack full random assignment. They often occur in natural
settings or with intact groups. For example, a researcher might compare two classrooms (one taught with a
program, one without) where students were not randomly assigned. This limits causal inference, because
pre-existing differences might explain outcomes. Quasi-experimental designs include nonequivalent control
group designs, one-group pretest-posttest, time-series designs, etc.
Definition: As one introduction notes, “Quasi-experimental design is a research methodology used to study the
effects of independent variables on dependent variables when full experimental control is not possible or
ethical” 28 . It “investigates cause-effect relationships in real-world settings” 28 . For example, evaluating a
therapy program rolled out by a school district (without random student assignment) would use a quasi-
experiment. Researchers attempt to increase validity by using statistical controls or matching, but some
confounding may remain. Thus, quasi-experimental studies can suggest causal relationships, but with
caution about internal validity.
Field Studies
Field studies are conducted in natural environments rather than in the lab. They can be experimental
(field experiments) or non-experimental (field observations or surveys). In field experiments, the researcher
manipulates an IV in a real-world setting (e.g. a social psychology experiment in a public place). In field
observations, behaviors are recorded in natural settings. Field research sacrifices some control but gains
ecological validity. As one source explains, “laboratory experiments usually have high internal validity but may
have low external validity, whereas field experiments have high external validity but lower internal control” 29 .
For instance, measuring customer behavior in an actual store versus a mock store can yield more realistic
data. Field studies often involve quasi-experimental elements (since full control/randomization may be
harder).
Cross-Cultural Studies
Phenomenology
Grounded Theory
A qualitative method where theory is inductively derived from data. Researchers collect data (e.g.
interviews) and iteratively code them to identify categories and relationships, until a theory “grounded” in
the data emerges. Used when existing theories are inadequate. Grounded theory involves constant
comparison, theoretical sampling, and memo-writing. It prioritizes theory development over hypothesis
testing. (General research methods knowledge.)
Focus Groups
As noted above, focus groups gather 6–10 people to discuss research topics in a moderated setting 26 .
They are a qualitative method (or early stage of mixed methods) to explore opinions and group interactions.
For example, focus groups can test how potential survey questions are interpreted. They allow participants
to build on each other’s ideas. Limitations include group dynamics overshadowing some voices and
potential moderator bias.
Narrative research is a qualitative approach where the researcher studies stories or personal accounts. It
treats participants’ life stories or narratives as data. For example, collecting autobiographical stories from
refugees and analyzing how they construct meaning over time. Researchers might conduct open-ended
interviews and analyze themes or plot structures. Narrative analysis assumes that personal stories reveal
how individuals make sense of experiences. Advantages: captures personal meaning over time; challenges:
requires careful interpretation and usually small samples 30 . As one source states, narrative research
“always includes data from individuals telling the story of their experiences” 30 .
Case Studies
A case study is an in-depth investigation of a single person, group, or event. It uses multiple data sources
(interviews, observations, documents) to capture complexity. For example, famous psychological case
studies include Phineas Gage or “patient HM” in memory research. Case studies provide detailed insight and
can generate hypotheses. However, they have limited generalizability and are subject to researcher
interpretation. According to research literature, “a case study is an in-depth study of one person, group, or
event… the goal may be to generalize across subjects”, but often generalization is not possible 31 .
Ethnography
Ethnography comes from anthropology: it involves long-term immersion of the researcher in a culture or
community to understand its practices and meanings. The researcher typically lives among the group
(participant-observation) and collects data on social interactions, rituals, language, etc. The goal is to
describe a cultural group “from the native’s point of view.” As one source defines, ethnography is “the
recording and analysis of a culture or society, usually based on participant-observation” 32 . For instance, an
ethnographer might live for a year in a foreign village and document social structure. Ethnography
produces rich, contextual data, but is time-consuming and findings may be idiosyncratic to that context 32 .
Summary of Methods: These diverse methods illustrate the range of approaches in psychology and social
science. Quantitative methods (experiments, surveys, structured observations) focus on measurement and
generalization; qualitative methods (interviews, focus groups, ethnography) focus on depth and meaning;
mixed or specialized methods (quasi-experiments, phenomenology, case studies) combine elements to suit
the research question. Each method has specific strengths, limitations, and typical applications, as outlined
above.
6. Statistics in Psychology
When summarizing quantitative data, central tendency and dispersion are fundamental descriptive
statistics.
• Measures of Central Tendency: These describe the “center” of a data set. The mode is the most
frequent score (useful for nominal data). The median is the middle score when data are ordered
(splits the data into two equal halves). The mean (arithmetic average) is the sum of all scores divided
by N 33 34 . For symmetric distributions, mean = median = mode; if skewed, the median often
better represents the “typical” value. The mean uses all data and is sensitive to extreme values
(outliers), whereas the median is robust against outliers 35 . For example, for data [1,2,3,4,19], the
mean is 5.8 but the median is 3, making the median more representative when data are skewed 35 .
Examples: If test scores are [70, 75, 80, 85, 90], the mode is 70 (if it repeats), the median is 80, and
the mean is (70+75+80+85+90)/5 = 80 33 34 .
• Measures of Dispersion: These describe variability around the center. The range is the simplest
(max – min). The variance and standard deviation (SD) indicate spread: SD is the average distance
of scores from the mean. A smaller SD means scores are clustered near the mean; a larger SD means
more spread. Psychologists often use SD because it is in the same units as the data. The
interquartile range (IQR) (middle 50% range) is used for skewed data. As one source explains,
range and SD “tell us how much variability there is in the data” 36 . For example, if group A scores [50,
52, 53, 55, 57] and group B scores [30, 40, 50, 60, 70], both have median 53, but group B has a much
larger range and SD, indicating higher variability. Tables and bar charts of means or medians often
pair with measures like SD or standard error to show spread.
The normal distribution (or Gaussian distribution) is a continuous probability distribution that is symmetric
and bell-shaped. Many psychological variables (IQ, measurement errors, etc.) approximate normality. Key
features: symmetry about the mean, and mean = median = mode at the center. The distribution is
described by its mean (μ) and standard deviation (σ). About 68% of values lie within ±1σ of the mean, ~95%
within ±2σ, and ~99.7% within ±3σ (the Empirical Rule). The tails approach the horizontal axis but never
touch it (asymptotic).
Figure: The normal (Gaussian) distribution curve. It is symmetric, with peak at the mean, and inflection points at
μ±σ.
This curve underlies many statistical tests. Raw scores (X) can be transformed to z-scores by subtracting the
mean and dividing by the SD: z = (X – μ)/σ. A z-score indicates how many SDs a value is from the mean. For
example, a z=+1 means one standard deviation above the mean. In a standard normal distribution (μ=0,
σ=1), these percentiles are easy to compute. The normal curve’s importance lies in the Central Limit
Theorem: the means of repeated samples tend to follow a normal distribution even if the underlying data
are not normal, given large N. In statistics, test assumptions often include normality; for example, t-tests
and ANOVA assume normally distributed residuals.
Statistical tests can be parametric (making assumptions about data distribution) or nonparametric
(distribution-free tests). Common parametric and nonparametric tests in psychology include:
• Parametric Test (t-test): The t-test compares means between groups. A one-sample t-test checks if a
sample mean differs from a known value. An independent-samples t-test compares means of two
independent groups (e.g. treatment vs control) 37 . A paired-samples t-test compares means of two
related samples (e.g. pre-test vs post-test). The t-test assumes interval/ratio data, normally
distributed populations, and (for independent t-test) equal variances. For example, in testing a new
teaching method, a t-test could compare students’ test scores between the method and control
groups. As scribbr notes, “a t-test is a statistical test that compares the means of two samples” 37 .
• Nonparametric Tests: These do not assume normality or interval data, often using ranks or signs.
Key nonparametric tests include:
• Sign Test: Tests whether the median of a single sample differs from a hypothesized value, or
whether two related samples differ in median. It simply counts the number of positive vs. negative
differences, ignoring magnitude. Useful for paired data with unknown distribution.
• Wilcoxon Signed-Rank Test: The nonparametric analogue of the paired t-test. For related samples
(or one sample vs. median), it ranks the absolute differences and considers sign, testing for a
median shift. According to references, “the Wilcoxon signed-rank test is a non-parametric rank test…
used to compare the locations of two populations using two matched samples” 38 .
• Mann-Whitney U Test: The nonparametric analogue of the independent-samples t-test. It compares
two independent groups by ranking all scores and comparing rank sums. It tests whether the
distributions (medians) differ. (Also called Wilcoxon rank-sum test.)
• Kruskal-Wallis Test: The nonparametric equivalent of one-way ANOVA for more than two groups. It
tests whether independent groups come from the same distribution by comparing ranks across
groups. For example, comparing median reaction times across three medication groups. One source
explains that “Kruskal-Wallis test can be thought of as the non-parametric equivalent to ANOVA” 39 ,
determining if groups have the same median.
• Friedman Test: The nonparametric equivalent of one-way repeated measures ANOVA. It tests for
differences across multiple related measures (e.g. same subjects measured under three conditions)
by ranking within each subject. As Wikipedia notes, “the Friedman test… similar to the parametric
repeated measures ANOVA” for detecting treatment differences across attempts 40 .
10
The choice between parametric and nonparametric depends on data type and distribution. Parametric tests
generally have more power when assumptions are met, while nonparametric tests are more robust to
violations of normality or ordinal data.
• Effect Size: This quantifies the magnitude of a phenomenon, independent of sample size. Common
effect sizes include Cohen’s d (difference in means/SD), Pearson’s r (correlation strength), and η² or
R² (proportion of variance explained). As one guide states, effect size is “a standardized index…
independent of sample size and quantifies the magnitude of the difference between populations or the
relationship” 41 . For example, Cohen’s d = 0.2 (small), 0.5 (medium), 0.8 (large) by convention. Effect
sizes help interpret practical significance beyond p-values.
• Power Analysis: Statistical power is the probability that a test will correctly reject the null hypothesis
when the alternative is true (i.e., detect a real effect). It is defined as 1 – β, where β is the Type II error
rate. According to statistical sources, “power is the probability of detecting a given effect (if that effect
actually exists)” 42 . Power depends on effect size, sample size, significance level (α), and variability.
High power (commonly 0.8 or 80%) is desired so that studies have a good chance of finding true
effects. Researchers conduct power analyses when designing studies to determine needed sample
size for expected effect sizes. For instance, a large effect can be detected with smaller N, whereas
small effects require larger N.
7. Correlational Analysis
Correlation measures the strength and direction of association between variables. In psychology, we
distinguish between Pearson (product-moment) correlation and Spearman (rank-order) correlation, as
well as more complex forms.
• Pearson Product-Moment Correlation (r): Measures the linear relationship between two
continuous (interval/ratio) variables. It ranges from –1 (perfect negative linear) through 0 (no linear
relation) to +1 (perfect positive linear). For example, height and weight often correlate positively. The
Pearson r is sensitive to outliers and only captures linear association. A value of r=0.75 indicates a
strong positive linear relationship 43 . As noted by statistics guides, “the Pearson correlation
coefficient (r) is the most common way of measuring a linear correlation… between two variables” 43 . The
formula standardizes covariance by the product of standard deviations. Hypothesis tests can assess
whether r is significantly different from zero.
11
• Partial Correlation: Partial correlation quantifies the relationship between two variables while
statistically controlling for one or more other variables. For example, one might examine the
correlation between stress and health while controlling for age. Partial correlation effectively
removes the influence of the control variable. One tutorial defines it as measuring “the strength of a
relationship between two variables, while controlling for the effect of one or more other variables” 45 .
Partial correlations are used in regression diagnostics and to examine the unique association
between two factors apart from shared influences.
• Multiple Correlation: Refers to the correlation between an observed variable and the linear
combination of multiple predictors (as in multiple regression). The multiple correlation coefficient R
(and its square, R²) indicates how well the set of independent variables collectively predict the
dependent variable. For example, if weight, diet, and exercise jointly predict blood pressure, the
multiple R quantifies their combined predictive strength. In regression context, R² is “the square of
the coefficient of multiple correlation” and represents the proportion of variance explained 46 . For
one predictor, R = |r|; with multiple predictors, R² can exceed any single r.
Note: Correlation does not imply causation. Even a high r simply indicates association, not directional
influence. However, partial and multiple correlations help disentangle relationships among several
variables.
• Point-Biserial Correlation: Used when one variable is dichotomous (binary) and the other is
continuous. For example, correlation between gender (male=0, female=1) and test score. The point-
biserial coefficient is mathematically equivalent to a Pearson r computed after coding the dichotomy.
By definition, “the point biserial correlation coefficient… is used when one variable (e.g. Y) is
dichotomous” 47 . It measures the difference in means between the two groups normalized by
variability.
• Biserial Correlation: Similar to point-biserial, but assumes the dichotomous variable is artificially
split (underlying continuous). Biserial uses an adjustment because the dichotomy is not “true” (e.g.
pass/fail on a test assumed from an underlying ability). The formula is more complex, involving the
proportion above/below cut. It estimates the correlation as if the data were continuous.
• Tetrachoric Correlation: Used when both variables are dichotomous, but assumed to arise from
underlying normal distributions. It estimates what the Pearson correlation would be between the
latent continuous variables. For example, a survey asking yes/no on two symptoms: the tetrachoric
correlation estimates the association of the underlying symptom severity. It is not the same as phi
(below).
• Phi Coefficient (φ): A simple measure of association for two binary variables. It is equivalent to
Pearson correlation on two 0/1 coded variables. For example, in a 2×2 table of gender (M/F) and
smoking status (Yes/No), φ measures their association. Wikipedia notes: “the phi coefficient… is a
12
measure of association for two binary variables” 48 . Its value ranges from –1 to +1, and is essentially
the Pearson r for dichotomies.
In sum, special correlation methods allow researchers to handle categorical data appropriately. For
instance, the point-biserial is used when comparing a continuous outcome across two groups (the
dichotomy), while phi and tetrachoric handle two dichotomies. These coefficients are important when data
do not fit simple continuous-variable assumptions.
9. Regression
Simple Linear Regression: A statistical method for modeling the relationship between one continuous
independent variable (X) and one continuous dependent variable (Y). The model assumes a linear
relationship: Y = β₀ + β₁X + ε, where β₀ is the intercept, β₁ the slope (regression coefficient), and ε the error
term. The slope β₁ indicates the expected change in Y for a one-unit change in X. For example, predicting
exam score (Y) from study hours (X) might yield score = 50 + 5×(hours), meaning each extra hour adds 5
points on average. Simple regression also provides R² (proportion of variance in Y explained by X) and tests
significance of β₁ (whether slope≠0). In practice, after estimating the regression, we check assumptions
(linearity, homoscedasticity, normality of errors). Regression output includes coefficients, their standard
errors, t-values and p-values 49 . Writing up results typically reports the estimated effect (β), its significance,
and R². As scribbr guides note, “simple linear regression is a statistical method that allows us to summarize and
study relationships between two quantitative variables” 50 .
Multiple Regression: An extension where multiple independent variables predict a dependent variable. The
model: Y = β₀ + β₁X₁ + β₂X₂ + … + βₖXₖ + ε. It allows assessing the unique effect of each predictor while
controlling others, and estimates overall model fit. For example, predicting job performance from education
level, work experience, and test scores. Multiple regression answers questions like: How strongly does each
factor relate to the outcome, and how much variance is explained by the set? According to guides, “multiple
linear regression is used to estimate the relationship between two or more independent variables and one
dependent variable” 51 . This method reports each β, their significance, and the multiple R² (shared
explained variance). It assumes no perfect multicollinearity and maintains the other linear-model
assumptions. Multiple regression is widely used for prediction and theory testing in psychology. It can also
include categorical predictors by coding them (dummy variables).
In both simple and multiple regression, the coefficient β (slope) is of primary interest. For example, if
β₁=0.71 (from [158]), it means for each unit increase in X (e.g. income), Y (e.g. happiness) increases by 0.71
units 52 . Regression coefficients are tested by t-tests (null β=0), and R² is tested by ANOVA (F-test).
Diagnostics must check for outliers and assumption violations.
13
Assumptions and Preliminaries: Factor analysis generally assumes: a sufficiently large sample size
(common rule: at least 5–20 observations per variable) 54 ; roughly continuous and normally distributed
variables (though Likert items are often used); linear relationships among variables; no severe outliers
(which can distort correlations); and that the data are “factorable” (variables correlate enough to form
factors). Tests like Bartlett’s test of sphericity and Kaiser-Meyer-Olkin (KMO) measure check adequacy. For
example, a KMO > 0.6 and significant Bartlett’s test support factorability.
Methods (Extraction): In EFA, factors are extracted from the correlation matrix via mathematical
transformations. Eigenvalues and eigenvectors determine factor loadings. An eigenvalue represents the
variance accounted for by a factor (an eigenvalue >1 rule is common to retain factors) 55 . The first factor
has the largest eigenvalue, capturing the most variance, the next factors capture remaining variance. A
scree plot (plot of eigenvalues) helps decide how many factors to keep: one looks for the “elbow” or the
point where eigenvalues level off 56 . After extraction, each variable has a factor loading on each factor
(correlation between the variable and factor), which aids interpretation (loadings > ~0.3 are considered
meaningful 57 ).
Common extraction methods include Principal Components Analysis (PCA) (often called just principal
axes extraction) and Principal Axis Factoring (PAF), as well as Maximum Likelihood. PCA is technically a
dimension-reduction technique (components), but is often used in psychology under “factor analysis.” PAF/
ML estimate common factors accounting only for shared variance (more aligned with true factor models).
Rotation: To aid interpretation, factors are rotated. Rotation redistributes variance across factors but does
not change total variance explained. There are two main types: - Orthogonal rotation (factors are
constrained to be uncorrelated). Examples: Varimax, Quartimax, Equamax. These often simplify loadings by
making some large and others near zero.
- Oblique rotation (factors are allowed to correlate). Examples: Promax, Direct Oblimin. These are used
when underlying factors are expected to relate.
Rotation makes it easier to interpret factors as distinct constructs. For instance, an orthogonal Varimax
rotation might load different items cleanly on each factor, allowing us to label each factor. As a factor
analysis guide explains, orthogonal rotations keep factors at 90° (uncorrelated), while oblique allow “any
position” (correlated factors) 58 . The choice depends on theory and goal: if factors theoretically relate (e.g.
in personality), oblique is appropriate. After rotation, researchers examine the rotated factor loadings
matrix to interpret factors (looking for items loading strongly on one factor).
Interpretation: Each factor is interpreted based on the high-loading variables. For example, a factor on
which only anxiety and depression items load might be labeled “negative affect.” The factor loadings
indicate variable-factor correlations. Researchers also consider communalities (proportion of each variable’s
variance explained by all factors) and total variance explained by retained factors. A common guideline is to
retain factors that explain a meaningful amount of variance (e.g. cumulative R²) and make theoretical sense.
In reporting, one would specify the number of factors extracted, rotation method, factor loadings, and the
percentage of variance explained by each factor (and cumulatively).
References: Factor analysis is defined as techniques revealing underlying constructs 53 , with assumptions
including sample size, measurement level, normality, linearity, and factorability 54 . Rotation procedures
(orthogonal vs oblique) are recommended to improve factor interpretability 58 .
14
• One-Way ANOVA: A statistical design (analysis) for comparing means across more than two groups
of one factor. For example, comparing test scores across three teaching methods (factor=Method
with 3 levels). It partitions variance into “between-group” and “within-group” components and tests if
group means differ overall. As noted, “one-way ANOVA is used to compare the means of more than two
groups when there is one independent variable and one dependent variable.” 59 . If significant, post-hoc
tests determine which groups differ.
• Factorial ANOVA: When there are two or more independent variables (factors). For example, a 2×3
factorial design might manipulate teaching method (2 levels) and study time (3 levels)
simultaneously. Factorial ANOVA tests for main effects of each factor and their interaction. This
design is efficient (tests multiple hypotheses in one study) and can detect interactions (e.g. Method A
works better only at high study time). It typically requires random assignment to each condition
(cell).
• Randomized Block Design: Here, subjects are first grouped into blocks based on some nuisance
variable (e.g. gender, age group), and within each block randomly assigned to treatments. This
controls block effects by comparing treatments within homogeneous blocks. For example, in a drug
trial, patients might be blocked by severity of illness, then randomly given drug or placebo. The block
design reduces variance due to the blocking variable, increasing test sensitivity. It assumes the
blocking factor is known and discrete.
• Repeated Measures (Within-Subjects) Design: The same subjects receive all conditions. For
example, each participant tries all drug doses in random order. This eliminates between-subject
variability (each person is their own control) and requires fewer participants. Analysis uses repeated
measures ANOVA. Key considerations: counterbalancing order to avoid carry-over effects. Limitation:
potential order effects (fatigue, learning).
• Latin Square Design: A special design for controlling two sources of nuisance variability without
having each subject experience all conditions. In an n×n Latin Square, n treatments appear exactly
once in each row and column. For example, with 4 treatments and 4 time slots and 4 groups, each
treatment appears once per group and once per time. This controls for two extraneous factors (rows
and columns). It assumes no interaction between treatment and blocking factors. Latin squares are
efficient when full counterbalancing (with repeated measures) is infeasible. 60 describes arranging
treatments in a square so each appears once in each row/column, controlling two extraneous
variables.
• Cohort Study (Observational Design): Not a randomized design, but included here as a design to
study causal hypotheses over time. In a prospective cohort, a group sharing a characteristic (e.g.
smokers) is followed over time to measure outcomes (e.g. incidence of cancer), often compared to a
non-exposed cohort. It is used when experiments are unethical/impractical (cannot randomly assign
smoking). Prospective cohort studies record exposure first, then track outcomes 61 . Retrospective
cohorts use existing records. Cohort studies are common in epidemiology and can suggest causal
15
links (though confounding must be considered). As summarized, “a cohort study… follows a group of
participants over time, examining how certain factors (like exposure) affect outcomes” 61 .
• Time Series Design: A design in which the dependent variable is measured repeatedly at many time
points, often before, during, and after an intervention. For example, measuring crime rates monthly
for several years before and after a new policing policy. This allows analysis of trends and
intervention effects. Time series analysis accounts for autocorrelation (each observation correlated
with previous ones). In experimental context, an interrupted time series can be quasi-experimental by
analyzing the “interruption” of a treatment introduction.
• MANOVA (Multivariate ANOVA): Extends ANOVA when there are multiple dependent variables. For
example, testing whether two groups differ on a set of psychological scales together. MANOVA
assesses whether the multivariate means differ across groups, considering the pattern of
relationships among DVs. It controls Type I error inflation from multiple ANOVAs, at the cost of
requiring multivariate normality and more complex interpretation (e.g. Wilks’ Lambda statistic). If
MANOVA is significant, follow-up analyses determine which DVs show effects.
• ANCOVA (Analysis of Covariance): Combines ANOVA and regression by including one or more
covariate (continuous) variables. ANCOVA adjusts group means on the DV for differences on the
covariate, effectively removing its variance. For example, comparing test scores between teaching
methods while controlling for IQ. ANCOVA increases precision (reduces error variance) and can
remove confounds. It assumes linear relationship between covariate and DV and homogeneity of
regression slopes across groups.
These experimental and quasi-experimental designs form the toolkit for rigorous research in psychology.
Each design has specific purposes, advantages, and assumptions (e.g. repeated measures designs assume
no carryover effects; Latin square assumes no interaction with blocks; cohorts assume reliable follow-up).
Choosing the appropriate design depends on the research question, practical constraints, and ethical
considerations.
Note: Many of the above methods are illustrated with additional examples and diagrams (e.g. ANOVA
tables, design matrices) in standard research methodology texts, which can aid understanding.
Sources: Key definitions and principles for experimental designs and statistical analyses are drawn from
methodological guides and statistical references 59 60 61 42 41 , ensuring a robust academic
presentation of the material.
16
3(PDF) Dimensions Of Research - Characterizing Research, Research Design, and Research Studies: A
Comprehensive Faceted Classification
[Link]
_Characterizing_Research_Research_Design_and_Research_Studies_A_Comprehensive_Faceted_Classification
17
18
19