International Business Research
1. The Research Process
A. Why business research?
Business scientists are empiricists
Business Research
A systematic process of testing hypotheses through carefully executed
data analyses that are aimed to help a manager solve or minimize a
problem.
1. BR is systematic process
2. BR tests hypotheses
3. BR entails collecting and analyzing data
4. BR helps manages make better decisions
a. Evidence based decisions - decisions that rely on a thorough and
painstaking assessment of empirical data.
Intuition:
Should never be a substitute for research
Managers (like all humans) are prone to cognitive biases.
Cognitive Biases - are unconscious thinking errors.
They are an attempt of our brain to simplify the complex world and
speed up decision-making.
o Thinking Fast and Slow by Nobel-prize winner Daniel Kahneman.
Confirmation Bias
refers to the tendency only to consider information that agrees with
("confirms") our preexisting beliefs.
Availability bias
bias in which we decide based on readily available information, even
though it may not be the best information to inform our decision.
Evaluating Research Evidence
Judging academic-journal quality
Predatory journals - sole purpose is to make money: they ask authors
to publish for a fee without providing a peer review
When evaluating journals check:
o Peer-reviewed - if they are not, the journal is likely to be
predatory.
o Impact factor - While this is not a perfect measure, a journal with
an impact factor of at least 1.0 is less likely to be predatory.
o Consult the list of quality journals
Judging popular-press articles
Most science journalists do a good job. However, sometimes a popular-
press article is unreliable because it is based on flawed academic
research.
You should be Knowledgeable about business research to:
Evaluate BR
Delegate BR – interact with research departments or agencies
Perform BR
B. Stages of the research process
1. Deductive vs Inductive research
a. Inductive research approach - collect data, find a pattern,
develop theoretical framework.
Goal: developing a theory
b. Deductive research approach - hypothesize relationships
between variables based on theory, test hypotheses using data.
Goal: Testing a theory
These approaches often are used in combination.
2. The 7-step deductive research process:
1. Demarcate the business problem
2. Formulate research questions
3. Develop the theoretical framework
4. Choose a research strategy
5. Collect the data
6. Analyze the data
7. Write a report
2. The Business Problem
A. Demarcating a business problem
2. When does a business problem arise?
Problem arises when a company faces either:
a. Threat - a difficulty to be overcome
b. Opportunity - a situation with the potential for improvement
3. Demarcating a business problem
Before diving into research study, a business problem must be:
a. Demarcated
b. Narrowed down
Example of poorly demarcated business problem:
a. “Pfizer wants to boost its profits”
Example of well demarcated business problem:
b. "Pfizer wants to know the impact of advertising spending on the
number of prescriptions written by doctors for Pfizer's products."
B. Problem relevance
1. Two types of relevance
Conducting research requires significant investment
Therefore, research should focus on relevant business problems:
a. Academic relevance
b. Managerial relevance
2. Academic relevance
When the business problem has already been researched, new
research study into the same problem offers little added value, not
worth the investment.
How a research study can contribute to the existing literature:
a. New topic
No prior research exists – the topic is important
b. New context
Prior research exists but in a different context
c. Integrate scattered findings
Prior studies focus on different variables in isolation, hence
their relative impact is unclear
d. Reconcile conflicting findings
Prior studies report different findings (small vs large,
positive vs negative), and the conditions under which these
findings hold are unclear.
3. Managerial relevance
Research of Business problem should benefit someone (company or set
of companies)
Study is Managerially relevant when one or more parties benefit from
the research into the problem:
1. Managers from:
1. One company
2. One industry
3. Multiple industries
2. End users
1. Consumers
2. Investors, etc.
3. Public policy makers
1. Governments, unions, etc.
3. Research Questions
A. Formulating research questions
1. From business problem to research questions
Demarcating a business problem is the first step in the deductive
research process.
The second step is formulating the study’s problem statement and
research questions.
2. The central question or problem statement
A good problem statement is:
1. An open-ended question
2. Identifies the study’s unit of analysis
3. Is expressed in terms of (i) variables and (ii) relationships
2.1 The problem statement is an open-ended question
An open-ended question is a question that cannot be answered
by a simple yes or no.
Problem statement has to be open-ended question to avoid
jumping to conclusions before the research has been conducted.
Ways of starting an open-ended question: what, how, to what extent.
2.2 A problem statement identifies the study’s unit of analysis
The unit of analysis is the focus of the study, entity that the study
wishes to say something about.
A problem statement should be clear about the unit of analysis of the
study.
Subjects – the entities being studied.
Examples of unit of analysis in a BR study:
Individuals: consumers, investors…
Firms: publicly listed companies, multinationals…
Groups: board of directors, alliances, industries.
Things: products, brands or shares.
Geographical units: cities, regions, countries.
A study's unit of analysis can be at a lower or a higher level of
aggregation:
Student < class < university
Country > industry > firm > brand > consumer
2.3 The problem statement: expressed in terms of variables
Variables are the core of every research study.
Variable must have at least two values or levels in a study.
Variables can vary:
1. Across subjects
2. Over time
3. Across subjects and over time
Constant – value that does not vary, shared characteristic.
2.4 The problem statement: Expressed in terms of relationships
A problem statement expresses the relationship between at least two
variables.
Moderating effect – how the relationship between these two
variables depends on a third variable.
Conditions – environment under which X is related to Y
3. From problem statement to research questions.
Research Question - Instead of formulating a single overly
complex question (the problem statement), researchers
formulate multiple subquestions.
Problem statement and research questions should be formulated using
correlational rather than causal claims.
B Writing a background section
The structure of the background section:
1. Framing the research (1 paragraph)
2. Identifying the problem (1)
3. Formulating the study’s aim (1)
4. Academic and managerial relevance (2 or more)
5. Outline of the study (1)
1. Framing the research:
Putting problem into context:
Through numbers (indicating the percentage of firms facing a
problem A)
Through examples (naming firms facing problem A)
Sources of numbers and examples:
Newspapers, business press, company websites and
annual reports, etc.
2. Identifying the problem:
Introduction of the problem you will address (1 or 2 paragraphs).
Option to wrap up this paragraph with problem statement or
research question.
Make sure the paragraph focuses on the variables of the study.
3. Academic relevance:
Argue how the study contributes to existing research.
Paragraph per contribution.
Ways in which a study can contribute to the academic literature:
New topic – important but has not been studied.
o Solely pointing out a gap in the literature is insufficient to
justify a research study, as the gap might be there for good
reasons.
New context – prior research exists in a different context.
o Merely applying an existing conceptual model to a new
industry does not constitute a substantial academic
contribution.
o Essence of contributing by studying a new context lies in
the reevaluation of the very bedrock of the conceptual
model.
Integrate scattered findings – knowledge is scattered across
multiple articles and disciplines.
Reconcile conflicting findings - The main-effect relationship may
have been studied before, but the moderators may still be
unknown.
o Introducing moderators is particularly interesting when
they might explain why prior studies reported contradictory
main-effect findings.
4. Managerial relevance:
Indicate who can benefit from your research and in what way.
4. Theoretical Framework
A. Literature review & Conceptual model
Theoretical framework includes:
1. Providing Literature review.
2. Presenting/visualizing the conceptual model.
3. Formulating expectations (research hypotheses) for the
relationships between the variables.
1. Literature review:
Provide a condensed overview of the key studies on a particular
topic.
2. The conceptual model:
Summarizes the new study.
Explains which variables are included and how they relate to
each other.
o Variables displayed by boxes: dependent, independent,
mediating, moderating, control.
o Relationships are visualized by arrows:
Main Effects (x effect on y)
Moderating Effects
o Dependent variable:
captures the phenomena we are trying to explain or
predict
Depicted by y
AKA: criterion variable, DV
o Independent variable:
influences the dependent variable
Denoted by x
AKA: predictor variable, IV
o Mediating variable:
Explains the process underlying x and y relationship
Explains how or why x affects y
AKA: mediator, intervening variable
o Moderating variable:
Changes the relationship between x and y (strength
or direction)
Explains when/under which conditions x affect y.
AKA: moderator, interaction variable.
Moderated mediation:
o Control variable:
Not the focus of the research
May confound the relationship between x and y
Moderator variable can sometimes also have a direct effect on
the DV, in addition to the moderating effect on the relationship
between the IV and the DV.
B. Writing a conceptual model section:
1. Literature review
For a good literature review:
Structure is the key to an insightful literature review.
Structured around relevant themes.
Controversies and gaps are pointed out.
Synthesize existing studies rather than summarizing them.
Never describe studies chronologically.
2. Conceptual model
Before visually showing the conceptual model:
Briefly discuss the general thrust:
o Include a formal definition of each variable.
o Variable definition should be based on the literature.
o If many definitions are present:
Acknowledge the major differences between them,
you can focus on shared meaning across them.
Pick one definition and justify why you will use it.
o Recommended: include a figure representing all the
variables and relationships that are the subject of the
study.
For a good Conceptual model:
Never define your variables one after the other
o Integrate the definitions in the text that briefly describe
your conceptual model.
Never use synonyms for your variable names
o Using the same variable names throughout your report
provides clarity.
Never define a variable by copy-pasting the first definition you
encounter in the literature
A variable cannot be defined by using examples
Never strive for complexity
C. Research hypotheses:
1. What is a research hypothesis?
Research hypothesis - is a tentative statement about the coherence
between two or more variables.
1. Research study will test, using data, whether this statement is
sound.
2. Is about the coherence, or the relationship, between variables.
3. Pertains to two or more variables:
A main-effect hypothesis concerns the relationship between
two variables.
A mediator and a moderator hypothesis concern the
relationship between three variables.
1.1 Link with statistics:
The alternate hypothesis in statistics equals a research
hypothesis.
1.2 Directional vs. non-directional hypothesis.
Directional hypotheses indicate the expected direction of the
relationship; is the expected association positive or negative?
Non-directional hypotheses expect a relationship, but they do not
indicate the direction.
Use non-directional hypotheses sparingly. Only use them
when the theory points in two equally likely directions.
2. Why start with research hypotheses?
Same study starting with research hypotheses or without it, would
produce the same empirical findings.
However, without research hypotheses, these empirical findings could
be a mere coincidence.
3. What makes a good research hypothesis?
1. A research hypothesis must be testable, it should be phrased in
terms of (measurable) variables.
2. A research hypothesis must be justified using logical
arguments based on prior (high-quality) research studies.
4. Finding support for a hypothesis.
Hypothesis can never be proven, it can only be supported by
(or consistent with) or not supported by (or not consistent with)
the data.
Supporting findings increase confidence that hypothesis is true, but not
proves it
D. Writing a hypothesis section.
1. Choosing variable names.
Variable names must:
Not overpromise.
Leave no room for ambiguity.
Be short.
2. Guidelines to formulate research hypothesis:
Research hypothesis proposes a relationship between two (or more)
variables.
The correct formulation of a hypothesis differs depending on the type of
variables involved:
Metric variable captures a quantity.
Categorical variable has different "levels" or "categories" that are
not ordered along an underlying dimension.
3. How to word a main-effect hypothesis.
3.1 A main-effect hypothesis when both the DV and IV are metric.
If two metric variables are related, you can expect the
relationship/association between these two variables to be either positive or
negative.
When you expect that an increase in X corresponds to an increase in Y,
the relationship between X and Y is positive.
It can be expressed:
H: When X increases, Y increases.
H: X is positively related to Y.
H: X is positively associated with Y.
When you expect an increase in X to correspond to a decrease in Y -- is
a negative relationship.
It can be expressed:
H: When X increases, Y decreases.
H: X is negatively related to Y.
H: X is negatively associated with Y.
When one of the variables is categorical rather than metric, hypotheses need
to be worded slightly differently.
3.2. Main-effect hypotheses when the IV is categorical and the DV is metric:
H: Men earn more than women.
3.3. Main-effect hypotheses when the DV is categorical and the IV is metric:
H: When downsizing increases, the likelihood of bankruptcy
increases.
4. How to word a mediator hypothesis.
The first hypothesis pertains to the relationship between the
independent variable X and the mediator MED.
The second hypothesis concerns the relationship between the
mediator MED and the dependent variable Y.
To emphasize the mediating role of MED, you can label the
hypotheses as a and b components of one and the same
hypothesis (e.g., H1a and H1b).
For three metric variables, the hypotheses could be formulated
as follows:
o H1a: When X increases, MED increases/decreases.
o H1b: When MED increases, Y increases/decreases.
Alternatively, a mediator hypothesis can be formulated as
follows:
o H: The relationship between X and Y is indirect through
MED.
Full mediation:
o Med fully explains the relationship between x and y
Partial mediation:
o MED partially explains the relationship between x and y
5. How to word a moderator hypothesis:
A moderator changes the strength of a relationship between two variables.
A moderator can make:
5.1 Positive relationship stronger (more positive)
5.2 Positive relationship weaker (less positive and possibly even
negative)
5.3 Negative relationship stronger (more negative)
5.4 Negative relationship weaker (less negative and possibly even
positive)
5.1 Positive relationship becoming stronger:
When variables (DV, IV, and MOD) are metric, and the MOD is expected to
strengthen the positive relationship between the DV and the IV, the
moderator hypothesis can be expressed as follows:
H: The positive relationship between X and Y strengthens when
MOD increases.
Positive relationship is denoted by +
5.2 Positive relationship becoming weaker:
When variables (DV, IV, and MOD are) metric, and the MOD is expected to
weaken the positive relationship between the DV and the IV, the moderator
hypothesis can be expressed as follows:
H: The positive relationship between X and
Y weakens/decreases when MOD increases.
Weakening relationship is denoted by –
5.3 Negative relationship becoming stronger:
When variables (DV, IV, and MOD) are metric, and the MOD is expected to
strengthen the negative relationship between the DV and the IV, the
moderator hypothesis can be expressed as follows:
H: The negative relationship between X and
Y strengthens when MOD increases.
5.4 A negative relationship becoming weaker
When (DV, IV, and MOD) are metric, and the MOD is expected to weaken the
negative relationship between the DV and the IV, the moderator hypothesis
can be expressed as follows:
H: The negative relationship between X and Y weakens when
MOD increases.
5.5 Moderation hypotheses when either the IV or the MOD is categorical
1. IV is metric and the MOD is categorical (2 levels), you can formulate
the following hypothesis:
H: The relationship between X and Y
is stronger/larger (weaker/smaller) for Level_1_of_MOD than
for Level_2_of_MOD.
2. The IV is categorical (2 levels) and the MOD is metric, you can
formulate the following hypothesis:
H: The difference (or: gap) in Y between Level_1_of_IV and
Level_2_of_IV is larger/smaller (or:
increases/decreases) when MOD increases.
3. The IV and MOD are both categorical (2 levels), you can formulate
the following hypothesis:
H: The difference in Y between Level_1_of_IV and Level_2_of_IV
is larger (or smaller) for Level_1_of_Mod than
for Level_2_of_Mod.
6. How to justify hypotheses?
You justify a hypothesis by providing logical arguments, based on the
existing literature.
These arguments collectively make the hypothesis plausible.
After you have built up your line of reasoning, you conclude with your
hypothesis, using:
This leads to the following hypothesis:
We therefore hypothesize:
As such:
Therefore:
Thus:
1. Do not simply claim that author x said so:
Provide reasons / explain underlying mechanisms (as to why X is
related to Y, or why MOD strengthens or weakens the
relationship between X and Y).
2. Do not summarize one article after the other:
Create your own synthesis of the literature at hand.
3. Use correlational rather than causal language
5a. Data Collection
A. Research strategies
1. Choose a research strategy
Addressing how you are going to reach the conclusion of the study
through the use of data.
Research strategies can be divided into:
Quantitative – archival research, survey research and
experimental research.
Qualitative – interviews, focus groups and case studies.
2. Three quantitative research strategies.
2.1. Survey research – collecting data from respondents by asking them
questions.
2.2. Experimental research – research strategy in which one or more
independent variables are manipulated.
Manipulation means that the researcher creates different
levels/categories of the independent variable(s).
A variable is measured when its levels are recorded as they occur
naturally.
2.3 Archival research – research that capitalizes on archival data.
Archival data are data that already exist, initially collected and
stored for purposes other than addressing the business problem.
Archival data are also referred to as secondary data.
3. Causal vs. correlational research.
Correlational research strategies: survey and archival research.
Causal research strategy: archival research.
To claim causality a study must meet four requirements:
1. Correlation of cause with effect.
2. Cause needs to come before the effect.
3. Control for confounding factors.
4. Good explanatory theory is required.
4. Why experimental research is a causal research strategy.
Causal research strategies test whether changes in one variable
actually result in (cause) changes in another variable.
In experiments, researchers actively change one or more X
variables (manipulate one or more variables)
The key to ensuring that participants differ only in terms of X lies
in how the experiment is carried out.
Random assignment of participants to the levels of the
manipulated variables automatically controls for all third factors
Z.
5. Why not always run experiments?
Random assignment is often either impossible or undesirable in
business settings.
Correlational studies cannot state that the association identified is a
cause-and-effect relationship but often they are the only way, thus, you
must control for third factors that you can identify.
B. Structuring a dataset
1. Introduction.
We collect data from subjects like consumers, firms, brands, etc.
The data offer insights into how these subjects score on various
characteristics we are interested in: the variables of interest.
These characteristics are variable in nature (hence the name variable):
their values may vary from one subject to another.
We say that we observe the scores on the variables for each subject.
Variables represent a subject's characteristics,
while observations denote the specific values associated with these
characteristics.
2. Datasets
A dataset is a collection of values in rows and columns.
The rows in a dataset capture the observations.
The columns in a dataset capture the variables.
In cross-sectional studies (studies at one point in time), the number of
rows equals the number of subjects (consumers, firms, ...) in the data
set.
Identifier – variable identification name.
Variables can take on textual values but manually entering textual data
in a data set is often impractical.
Researchers typically recode textual values: they replace text-based
answers with a numerical code.
Dummy variables - only take on the values 0 or 1 are called.
3. Ensuring your variables match your unit of analysis.
The unit of analysis is the entity that you are analyzing
The variables in a data set must match the study’s unit of
analysis.
The identifier must ensure each observation can be uniquely
distinguished in the dataset.
The dependent variable has to be measured at the level of the
unit of analysis. So must the mediator variables.
Independent and moderator variables have to be measured
at the level of the unit of analysis or at a more aggregate level.
C. Sampling
1. Sampling vs population.
Population – entire group of people, firms, events or things of
interest.
Sample – subset of the population of interest.
Sampling – process of selecting a number of elements from the
population.
Sampling error – the difference between the sample and population
values.
Sampling is used since often it is impossible or too difficult to evaluate the
entire population.
2. Steps in the sampling process.
1. Define the population of interest.
2. Determine the availability of a good sampling frame.
a. Sampling frame is the physical representation of the
population through which one can reach out to that
population.
3. Decide on the sampling design.
2.1 Population and sampling frame
Coverage error - sampling frame unequal population
Under-coverage: true population members are excluded.
Miss-coverage: non-population members are included.
Solutions:
o If small, recognize but ignore.
o If large, redefine the population in terms of the sampling
frame.
2.2. The sampling design.
If a good sampling frame exists, you can opt for a probability sampling
design.
if there is no (good) sampling frame available, you will need to resort
to a nonprobability sampling design.
Probability sampling – each element of the population has a known
chance of being selected as a subject.
Results generalizable to population.
More time resource intensive.
Nonprobability sampling – the elements of the population do not have
a know chance of being selected as a subject.
Less time and resource intensive.
Results not generalizable to population.
Sampling designs:
Probability:
o Simple random sampling
Each element has an equal chance of being chosen.
High generalizability
Might be costly
o Systematic sampling
Select random starting point and then picking every
ith element.
High simplicity
Low generalizability
o Stratified sampling
Divide the population in meaningful (homogenous)
groups, then apply SRS within each group.
All groups are adequately sampled, allowing for
group comparisons.
More time consuming.
Requires homogenous subgroups.
o Cluster sampling
Divide the population in heterogeneous groups,
randomly select a number of groups an select each
member within these groups.
Geographic clusters
Subsets of naturally occurring clusters are
typically more homogeneous than
heterogeneous.
Nonprobability:
o Convenience sampling
Select subjects who are conveniently available.
Convenient (inexpensive and fast)
Lower generalizability
o Quata sampling
Population is first segmented into mutually exclusive
subgroups, then, judgment is used to select the
subjects or units from each segment based on a
specific proportion.
When minority participation is critical
Lower generalizability
o Judgement sampling
Select subjects based on their
knowledge/professional judgment.
Convenient (inexpensive and fast) when a
limited number of people have the info you
need.
Lower generalizability
o Snowball sampling
Small initial group of respondents is selected
randomly, then subsequent respondents are selected
based on information provided by initial respondents.
This process in carried out in waves by obtaining
referrals from referrals.
Useful for rare characteristics (experts).
First participants strongly influence the sample.
D. Measurement
1. What is operationalization?
Operationalization or measurement is the process of turning
conceptual variables into measurable observations.
Operationalization almost inevitably involves some loss of
meaning because measures seldom reflect all that we mean by a
concept.
Therefore, we have to seek measures that capture as much of
the concept's meaning as accurately as possible.
2. The process of operationalization.
How do we operationalize conceptual variables?
1. Start with a precise definition of the conceptual variable.
2. Turn to prior empirical work in your area of study. This will
show you how the conceptual variable has been measured in the
past.
a. It is recommended that you use the same measurement
instruments as (high-quality) research studies before
yours.
b. These measurement instruments have already been tested
and refined.
How do you find existing measurement instruments?
Read the methods sections of the articles in your literature
review.
Search handbooks on measurement scales
Sometimes the measurement instruments will need to be
adapted.
o Remember that changing even small parts of a
measurement instrument can influence how well it
measures the construct.
o Only make adaptations when really needed and do this
with care.
You may want to develop your own measurement
instrument from scratch.
3. Single vs multiple indicators.
Single-item scale - measurement instrument that consists of one
indicator.
Multi-item scale - measurement instrument that consists of multiple
indicators.
Concrete variables are typically physically verifiable, such as
consumer age or the number of products purchased.
For concrete conceptual variables, a single indicator can be
sufficient to capture the variable's entire meaning.
Abstract variables cannot be directly observed because they are
mental constructs, such as a person's attitudes, opinions, or
intentions.
For abstract conceptual variables, multiple indicators are
preferable.
4. Levels of measurement.
Variables can be operationalized at various levels of measurement.
Nominal variables – represent two or more categories without any
inherent order or ranking.
Ordinal variables involve categories that do have a logical order, but
the distance between them is not consistent or measurable.
Interval-ratio variables have ordered categories with consistent
differences between them.
Because of this property, interval variables can be meaningfully
subtracted or added. It allows for more statistical analyses. For
example, the central tendency can be measured by the mode,
the median, or the mean; standard deviations can also be
calculated.
Interval variables and ratio variables are similar in that they have
consistent differences between categories.
They differ in a few meaningful ways:
Ratio variables include a natural zero point, where zero means
the absence of the characteristic being measured. (Profit,
number of products sold)
Interval variables have no natural zero point. (Time,
temperature)
Differences between variable levels can be meaningfully
calculated for both interval and ratio variables
For ratio variables, also the ratios between variable levels are
meaningful and can be compared quantitatively.
3. The measurement level of 5- and 7-point scales.
5- and 7-point scales are ordinal.
However, these scales are typically treated as interval scales:
the scores are considered to have even spacing between them.
Nominal and ordinal variables are jointly called categorical variables.
Interval and ratio variables are also referred to as metric variables.
The nominal measurement level is the least precise and
informative
The interval/ratio level is the most precise and informative.
Given a choice, choose an interval or ratio measurement level
over a nominal or ordinal level.
5b. Data analysis
A. Choosing statistical tests/techniques (R version)
1. introduction
Data becomes valuable only after the application of the proper analysis tools
Important to choose the appropriate statistical tests/techniques
The measurement level of variables determines which statistical
test/technique is suitable to assess their relationship.
In essence, two factors guide the selection of a suitable statistical
test/technique:
a. The number of independent variables in the conceptual model:
one vs. multiple
b. The measurement levels of these variables: metric vs.
categorical
2. Overview of statistical tests/ techniques
2.1 Statistical test in the case of a conceptual model with one independent
variable
Pearson's correlation coefficient is most suitable for uncovering
relationships between two variables when:
Both the dependent variable and the (single!) independent
variable in the conceptual model are metric (interval or ratio).
Chi-square test is most appropriate where:
Dependent variable and the independent variable are
categorical (nominal or ordinal).
T-test or One-way ANOVA is suitable when:
Dependent variable is metric, but the (single) independent
variable is categorical.
The selection between these two hinges on the number of levels of
the independent variable:
o When the independent variable has just two levels, a t-test is
the appropriate choice.
o When the independent variable comprises three or more
levels, a one-way ANOVA is the appropriate test.
2.2 Statistical techniques in the case of a conceptual model with multiple
independent variables
When dealing with a metric dependent variable, the tool of choice is
either ANOVA or linear regression analysis.
Although both techniques are mathematically equivalent, they report
the results in a slightly different way:
Regression analysis focuses on how changes in continuous
independent variables affect the dependent variable (but can
also deal with categorical IVs)
o The independent variables in a linear regression analysis
are typically metric, but regression can also deal with
categorical variables.
ANOVA focuses on uncovering group differences (but can also
deal with continuous IVs).
o The independent variables in an ANOVA are typically
categorical, but ANOVA can also deal with metric variables.
Experimental studies typically use (variations of) ANOVA
Archival and survey studies typically use (variations of)
regression analysis.
Logit analysis is the appropriate technique when the dependent variable is
categorical. Once again, the independent variable(s) can encompass metric
and/or categorical variables.
3. Pearson’s correlation coefficient: a metric DV and IV
Pearson's correlation coefficient measures the strength of the linear
relationship between two metric (interval or ratio) variables.
The possible range of values for the correlation coefficient is -1.0 to 1.0.
Note that a correlation of zero between two variables does not mean that
there is no relationship between the two variables at all.
It means that there is no linear relationship. The relationship could still be
non-linear.
To calculate a correlation coefficient between two variables, X and Y, in R,
use the cor() function.
data ← [Link]("[Link]")
correlation ← cor(data$X, data$Y, method = “pearson”)
print(correlation)
To test if the correlation is statistically significant, use [Link]()
[Link](data$X, data$Y, method = “pearson”)
4. Chi-square tests: a categorical DV and IV
Chi-square test tests whether there is a relationship between two
categorical variables (nominal or ordinal).
Essentially, it checks whether the frequencies observed in the sample
differ significantly from the frequencies one would expect.
Thus, the observed frequencies are compared with the expected
frequencies and their deviations are examined.
In R, the chi-square test can be conducted using the [Link] function.
First, we construct a contingency table that summarizes the counts of
each category combination. Then, we run the chi-square test:
data ← [Link]("[Link]")
table_data ← table(data$Gender, data$Education)
[Link](table_data)
This is how the R-output for a χ2χ2 (chi-square) test looks like:
data: table_data
X-squared = 0.487, df = 3, p-value = 0.485
With a p-value of .485, the χ2χ2 value of .487 is insignificant. This
implies that the variables are not related.
5. T-test: a metric DV and a categorical IV with 2 levels
T-test is appropriate, when the dependent variable is metric and the
independent variable is categorical with two levels (just two, no more).
Measures if the difference between the means of two groups is
significant
The groups compared can be: independent or paired:
o Independent groups:
Unrelated
In search groups the subject from the first group differs from
the subjects in the second group.
Independent samples T-test is used to compare the
independent groups
The data set needs to contain one nominal variable with
two levels
One metric variable is needed (interval or ratio) to
calculate the mean
o Paired groups:
Related
The same subjects are present in both groups
Paired samples T-test:
Two metric variables that are measured for the same group
Hypotheses:
Null hypothesis:
No mean difference exists between the groups
Alternative hypothesis:
Non-directional – a mean difference exists between the groups
Directional – The mean of group 1 is larger than the mean of
group 2
Interpretation:
Output:
T-statistic
P-value – how likely the difference between the two groups could
have happened by accident
< 0.05 = statistically significant
T-test in R:
The T-test requires packages:
“tidyverse” package for running a T-test
“car” package to check the variance assumptions of the T-test
T-test for independent groups:
T-test for paired groups:
6. One-way ANOVA: a metric DV and a categorical IV with 3 or more levels
One-way ANOVA is appropriate when the dependent variable is metric, and
the independent variable is categorical with three levels (or more).
Aims to establish if the difference between the means of three or more
groups is significant.
One-way ANOVA is an extension of the independent samples t-test for
more than two groups.
Hypothesis:
Null Hypothesis:
No significant differences between the means of the groups
Alternative Hypothesis:
At least two group means are significantly different from each
other.
Post hoc tests is used to determine which group means differ
One-way ANOVA in R:
Results:
A one-way ANOVA is the extension of a t-test for independent samples for
more than two groups.
If we want to test whether there is a difference between three or more
dependent samples, we use the analysis of variance with repeated
measures.
This is the case, for example, when the same group of subjects is
surveyed at three different times.
In R, a one-way analysis of variance with repeated measures can be run
using the function aov() or lme().
7. ANOVA or regression analysis
When the dependent variable is metric, an ANOVA or a regression analysis
can be used.
Regression analysis focuses on how changes in continuous
independent variables affect the dependent variable (but can also deal
with categorical IVs)
ANOVA focuses on uncovering group differences (but can also deal with
continuous IVs).
7.1 ANOVA
One-way ANCOVA
When a study contains one manipulated IV and one metric IV, it is called a
one-way ANCOVA (analysis of covariance). You can think of a one-way
ANCOVA as an extension of the one-way ANOVA that incorporates
a covariate.
Like the one-way ANOVA, the one-way ANCOVA is used to determine
whether there are any significant differences between two or more
independent (unrelated) groups on a dependent variable.
However, whereas the ANOVA looks for differences in the group means,
the ANCOVA looks for differences in adjusted means (i.e., adjusted
for the covariate).
As such, compared to the one-way ANOVA, the one-way ANCOVA has
the additional benefit of allowing you to "statistically control" for a third
variable (also known as a "confounding variable" or "control variable"),
which you believe will affect your results.
This third variable that could be confounding your results is called
the covariate; hence the name analysis of covariance.
You can have more than one covariate. Although covariates are
traditionally measured on a metric scale, they can also be categorical.
Two-way ANOVA or ANCOVA
The "one-way" part of one-way ANOVA refers to the number of
manipulated independent variables.
If you have two independent categorical variables rather than one,
you run a two-way ANOVA.
Again, you can add one or more control variables, in which case we
would refer to a two-way ANCOVA.
7.2 Regression analysis:
Linear regression analysis can be used when the dependent variable
is metric
The independent variables in a linear regression can be metric
and/or categorical.
Logit analysis is appropriate when the dependent variable is
categorical.
the independent variable(s) in a logit analysis can also be metric
and/or categorical.
Regression analysis seeks explain the relationship between a
dependent variable and one or more independent variables.
Can be used to make predictions
Simple linear regression contains one independent variable
Multiple linear regression contains multiple independent variables
In regression analysis:
Dependent variable must be metric
Independent variable can be metric or categorical
A categorical IV’s with 2 levels can be included in a regression
analysis as a dummy variable.
o Dummy variable is a variable that takes on a value of 1 or 0
Since linear regression analysis can accommodate multiple independent
variables, it is highly suitable to test moderating effects
Linear regression in R:
For human behavior R^2 of 0.30 is very good
Testing for moderating effect:
Results:
For those seeking in-depth knowledge about these tests/techniques and
others (e.g., in the context of writing a bachelor's or master's thesis), we
recommend exploring the excellent tutorials on DATAtab.
Reliability & Validity
A. Measurement reliability and validity
1. Introduction
Measurement reliability is like consistency, the degree to which
multiple measurements give the same result.
Measurement validity is about whether a measurement instrument
actually measures what it is supposed to, the degree to which the
scores on a measure represent the variable they are intended to.
Construct validity – when measurement instrument is both reliable
and valid, desired scenario.
4. Demonstrating measurement reliability
We should think about the different ways in which we can "repeat" our
measure to see if the results are similar:
A. Test-retest reliability:
o Test-retest reliability is the degree of agreement between the
results when the same measure is repeated sometime later
(under the same conditions).
B. Inter-rater reliability:
o Inter-rater reliability is the degree of agreement between the
results when (at least) two people ("raters") administer the
measure to the same subject (under the same conditions).
C. Internal consistency:
o Internal consistency is the degree of agreement between a single
measurement instrument's different questions (also referred to
as items).
o High internal consistency can be measured by calculating
Cronbach’s alpha
Demonstrating internal consistency via Cronbach's alpha:
Cronbach’s alpha is a statistic derived from pairwise correlations
between items that are supposed to measure the same conceptual
variable.
Cronbach’s alpha is relevant when multiple items are used to
measure the same construct
Evaluates to what extent these items are interrelated.
Formula:
Ranges from 0 (not correlated) to 1 (perfectly correlated).
Cronbach’s Alpha should not be below 0.6
In case of a negative a:
o Check whether items need to be reverse coded.
Extremely high a might indicate that items are too similar,
therefore, useless.
Calculating Cronbach’s alpha in R:
1. Check if any of the items measure the opposite of what we are
trying to capture.
a. This can happen when the item is phrased negatively.
b. We need to reverse code these items.
i. Read in the CSV
ii. Create a “name_rev” column to store the reverse
coded values. For ex, take maximum values and
subtract the data values.
iii. Check head data to see if the recoding worked
correctly.
2. Calculate item-total correlations – help us see if each item fits
well with the rest.
a. Shows us how strongly individual item is correlated with
the sum of the other items in the scale.
b. If the items correlation with the total is below 0.3 it is a
concern. Suggests that the item might not be strongly
connected to the overall construct.
3. Check internal reliability:
a. We lode the “psych” library.
b. We specify the columns that contain our scale items.
c. We use the function “alpha” to calculate the Cronbach’s
alpha and the item total correlations.
4. Interpreting results:
a. Is Cronbach’s alpha acceptable (>0.7)?
b. Are the item-total correlations above 0.3?
i. If any of these conditions is not met, we can check
the output to see if dropping an item will improeve
the Cronbach’s alpha.
ii. Do not delete items just to increase the Cronbach’s
alpha by a small margin. Creates content loss.
iii. Never drop more than one item at once. Drop items
sequentially.
c. Looking at the output we see two values for Cronbach’s
alpha:
i. The first one should be used when all values are
measured on the same scale.
ii. The second one is for cases when items are not on
the same scale.
d. Checking the item total correlations.
i. Look at [Link] column in “Item statistics” for item-
total correlation below 0.3.
ii. Look at “Reliability if the item is dropped table”
“Raw_alpha” column to see the potential Cronbach’s
alpha if the item is dropped.
e. Look at the output after reteining the useful items
i. Cronbach’s alpha is high.
ii. All item-total correlations above 0.3.
iii. Cronbach’s alpha cannot be substitantially increased
by dropping an item.
5. After we know which items to retain, we can combine them into a
single scale value.
a. To combine the three items into a single score in R, we
calculate the average of those items.
i. If items are measured on the same scale, we use the
“rowMeans” function to do so.
ii.
If
items are measured on different scales:
1. Standardize each item:
a. Transforming each item, so that it has a
mean of zero and SD of 1
b. Once standardized, all of the items will
be measured on SD – the same scale.
c. In R, first create standardized item values
using the “scale” function.
2. Average the standardized item values into a
single scale value.
a. Once the items are standardized you can
then average them into a single scale
value using “rowMeans” function
Demonstrating measurement validity:
To evaluate (or demonstrate) the validity of measurement
instruments, assess (or argue) to what extent the items in the measurement
instrument adequately represent the conceptual variable they are
intended to measure.
One way: by providing precedence - reviewing the literature and
referring to other (high-quality) studies that have used the same
measurement instrument to measure this particular variable.
When precedence does not exist (e.g., because an entirely new
conceptual variable is studied), measurement validity can be assessed
through expert judgment.
o A number of experts in the field can be provided with the
construct definition and asked to evaluate the appropriateness of
the measurement items.
Single-item measures for abstract conceptual variables tend to have
low measurement validity. Abstract variables are typically multifaceted
and can have various dimensions or aspects.
B. Internal and external validity
Ensuring the quality of a study (presuming that we have already ensured
that the measures are of high quality):
Two critical criteria are study’s internal validity and external
validity:
Internal validity refers to the extent to which a study can
eliminate alternative explanations (confounding variables)
apart from the independent variable for a caused change in the
dependent variable.
o The less chance there is for confounding in a study, the
higher the internal validity and the more confident we can
be in the findings.
o Make sure there is a control group that receives everything
the experimental group gets, just without the experimental
intervention.
External validity - how well the findings of a study
are generalizable to a broader population.
o Do the findings generalize to:
Other subjects
Other settings
o Good sampling is the key to good external validity
1. Sample should be representative of the wider
population of interest.
2. Sample should be large enough to impart adequate
statistical power.
3. Exclusion criteria should relate to the question of
interest.
Balancing internal vs external validity:
Trade-off between internal and external validity:
As a study becomes more applicable to a broader context (external
validity), it becomes increasingly difficult to control for all
extraneous factors (internal validity).
The optimal study design has both internal and external validity.
However, increasing one without decreasing the other is often not
possible.
We must evaluate to what degree a study performs in terms of
both.
7. Survey Research
A. Surveys vs polls
In survey research, concepts are operationalized through questions,
and observations consist of recording respondents' answers to these
questions.
1. Survey vs. poll
Survey consists of many questions, usually across a wide range of
question types, same as questionnaire.
Respondents - those who answer the questions.
Surveys can be cross-sectional or longitudinal:
Cross-sectional surveys capture a snapshot of a population at a
specific point in time, freezing one moment for analysis.
Longitudinal surveys track the same group of respondents over
time, help to identify trends.
Polls contain just a single or a few questions, can thus be thought of
as a special type of short survey.
Polls only allow for descriptive research - research that draws
a detailed picture of the current state of affairs.
Surveys allow for explanatory research: deep diving into the
reasons behind certain outcomes.
2. Surveys vs. census
Census is a survey of the entire population. It does not use a
sampling method.
Countries’ governments generally carry out censuses.
Census is usually carried out several years apart.
Surveys differ from censuses in their use of sampling a population.
Surveys do not attempt to collect data for every member of the
population.
3. When to use survey research?
In survey research:
All concepts are operationalized through questions
Observations consist of recording respondents' answers to
these questions.
Particularly suited for studies in which individuals (consumers,
managers, investors, etc.) are the unit of analysis.
Survey research is particularly useful for discovering
individuals' perceptions, opinions, attitudes, and behaviors:
Perceptions pertain to what individuals (know or think they
know).
Opinions pertain to people's preferences or judgments.
Attitudes are relatively stable, subjective orientations towards
events, objects, or ideas.
Behaviors are measured through statements about how people
act.
B. Developing survey questions
1. Introduction
Item is a term used to refer to questions in a questionnaire, because
survey questions can be questions as well as phrases.
2. Guidelines for question wording
Use simple words to increase your respondents' understanding.
o Wrong “What is your top priority when selecting stocks for your
portfolio?”
o Better “What is most important when selecting stocks for your
portfolio?”
Avoid jargon/abbreviations unless your respondents widely understand
these.
o Wrong “To what extent do you believe that a company's
sustainability investments affect its EBITDA?”
o Better “To what extent do you believe that a company's
sustainability investments affect its profitability?”
Avoid long sentences
o Longer questions are more likely to confuse and bore
respondents.
o Wrong “If you were asked to evaluate electric cars, which of the
following brands do you think you like best?”
o Better “Which of the following electric car brands do you like
best?”
Avoid using ambiguous terms that may have individually defined
meanings
o Wrong “During the winter, do you regularly eat ice cream?”
o Better “Regularly is ambiguous term”
Avoid double negatives:
o Questions that are negatively phrased may confuse respondents.
o Double negatives are even worse and may cause respondents to
answer just the opposite of what they meant.
Wrong ”Do you not oppose not allowing the board to pass
Article 10?”
Do you allow the board to pass Article 10?
Avoid double-barreled questions:
o Double-barreled questions are questions that ask about two or
more different topics or aspects in one.
Wrong “Do you like our products and customer service? “
Respondents cannot answer with "yes" or "no" if they
like the product but do not like the service (and vice
versa).
When respondents answer, you do not know whether
they might be responding to the first half of the
question, the second half, or both.
Better:
Do you like our products?
Do you like our customer service?
Avoid leading questions:
o Leading questions are questions that influence or persuade
("lead") the respondent to answer in a certain way.
Example of a leading question: “Would you rather go
through the hassle of shopping at the store or use this new,
improved app on your phone?”
Better:
Do you favor shopping at the store or shopping
through this app?
3. Open-ended versus closed-ended questions
The first decision you have to make is: are you going to ask an open-
ended or a closed-ended question?
Open-ended questions provide respondents with a question
prompt and a blank answer space in which they can write down
their responses, may ask to give a longer description,
explanation or a short response.
Closed-ended questions provide a question prompt and ask
respondents to choose from a list of possible responses.
3.1 Guidelines for creating open-ended questions
When asking for longer descriptions, provide statements to impress
upon the respondents the importance of their response:
Phrases such as:
o this question is very important
o please take your time answering this question
have been found to increase the length of the response.
Overuse of these phrases will only reduce their effectiveness
When asking for numerical responses, indicate the specific unit
desired in the question stem and provide unit labels with the answer
space.
3.2 Guidelines for creating closed-ended questions
Three common format types:
Rating questions, where the respondent is asked to rate a
statement.
Comparative questions, where the respondent is asked to rank
order something.
Categorical questions, where the respondent's answer can fit
only one category.
3.2.1 Rating questions
Rating questions ask respondents to rate their beliefs and perceptions
on a numerical scale.
Likert scales and semantic differential scales are two rating
question types that are often used in survey research:
Likert scale question asks respondents to agree or disagree
with a statement.
o Example the question ‘How much do you agree with a
statement?’ is provided alongside a 1-5 scale, where 1 is
strongly disagree, while 5 is strongly agree.
o When respondents answer multiple statements on the
same scale, these are usually organized as a matrix.
Semantic differential scales use a pair of polar-opposite
adjectives or phrases at the extremes of the scale, on the left
and the right, and respondents are asked to indicate their
attitudes on what may be called a semantic space toward a
particular individual, object, or event.
Guidelines on using rating scales
Include a middle point.
Considering the number of scale points:
o Too few scale points can oversimplify responses, potentially
losing valuable nuances in participants' opinions.
o Too many scale points, however, may lead to respondent
confusion and lower response quality.
o Using 5- or 7-point scales seems ideal, can be treated as
interval scales when analyzing the data.
How should I label the response options?
o Some surveys only label the endpoints, while others also
label the midpoint.
o The most accurate surveys will have a clear and specific
label that indicates the exact meaning of each point, no
room for different interpretations.
3.2.2 Comparative questions
Comparative questions are used to tap preferences between two or
more objects.
Rank ordering scales and constant sum scales are the most
common comparative questions.
Rank ordering scales - Respondents rank objects relative to
one another, among the alternatives that are provided.
o For example: fitness tracker manufacturing company wants
to know what features are ranked most important by their
customers.
Constant sum scale - Respondents divide a budget of points
(often 100 points) amongst a set of options according to their
personal preferences.
o For example, you can ask respondents to allocate 100
points on how they spend their income. Say they spend 40
on groceries, 20 on entertainment, 30 on utilities, and 10
on miscellaneous expenses.
o Allows respondents to show the magnitude of their
preferences, not just their ranked order.
Guidelines on using comparative scales
Limit the number of things to rank:
o Longer lists, especially those exceeding 10 items, typically
yield non-meaningful data, since respondents won’t feel
strongly about the middle rank of long lists.
All of the things in the list must be things the respondents are
familiar with.
Don't overuse comparative scales:
o It is difficult and high effort for respondents to decide
which things they prefer, especially if they like two things
equally.
3.3 Categorical scales
A categorical scale is a scale where respondents choose from a limited
number of discrete answer categories.
These can be ordered or unordered,
as in the examples below:
Guidelines on using categorical scales.
Make responses mutually exclusive.
Provide exhaustive answers:
o Include a list of all reasonably
possible answers.
o In situations where an
exhaustive list is not available
or practical, you can include a
response option that states
"Other (please specify)".
Avoid "check all that apply" – present many analysis issues.
C. Designing a survey
1. Introduction
The sequence of designing a survey follows a logical progression:
1. Decision on the survey mode.
2. Obtaining participants' consent.
3. Presenting the questions, ordered in a particular way.
4. The conclusion of the survey, should express gratitude for
participation and provide any follow-up instructions.
2. Picking a survey mode
A survey mode or survey method is the way you decide to
administer or distribute your survey.
The most common survey modes are:
Online
Paper
Telephone
Face-to-face
2.1 Online Surveys
Online Survey where the respondents are recruited by an online
method and the questionnaire is completed online.
Benefits of Online Surveys:
Reach large audience
Affordable
Templates are available
Limitations of online surveys:
Coverage bias - Certain portions of the population may not
have easy access to the internet. To eliminate coverage bias,
those without internet access can be asked to complete the
survey via other means.
Survey fatigue
There are two major ways to recruit participants for online
surveys: river sampling and panel sampling:
River sampling is the simplest approach where respondents are
recruited by inviting them to follow a link to a survey placed on a
web page, email, or somewhere else where it is likely to be
noticed by members of the target population.
o The name refers to the idea of researchers dipping into the
traffic flow of a website, catching some of the users
floating by.
Panel sampling - researchers select members of a
preassembled panel to take part in their survey.
o Two types of panels exist:
Online probability panels
Online non-probability panels
Online probability panels select individuals for their panels
through probability-based sampling methods.
These panels are relatively expensive to set up and
maintain, hence their usage remains rare.
Online non-probability panels typically select individuals for
their panel through river sampling or individuals voluntarily sign
up to complete surveys on the website.
Some of these sites offer considerable quota sampling
capabilities.
They have many commercial providers.
Typically, the cheapest method for online recruitment.
2.2 Paper surveys
Respondents receive paper surveys via traditional mail, usually with a
postage-paid envelope for their return.
Benefits of paper surveys:
Paper surveys are an excellent alternative when your target
audience's internet access or internet knowledge is limited.
Respondents give more honest answers when compared to other
modes.
Respondents trust paper surveys more than online surveys.
Limitations of paper surveys:
Cost of printing, postage, etc.
Respondents may only answer certain questions, leaving an
incomplete response.
If your study requires an alternating question order, paper
surveys may be too costly to support this requirement.
2.3 Phone surveys
In phone surveys, an interviewer follows a script in asking a specific set
of questions to the respondents on the phone, and data-entry software
is used to record the respondent’s answers.
Benefits of phone surveys:
Extensive geographic access since most people have a phone.
Easy access to a sampling frame since phone numbers can easily
be purchased from sample companies.
Interviewers can encourage respondents to answer all questions.
They can provide assistance in case the respondent is confused
about (any part of) the survey.
Limitations of phone surveys:
Intrusive, since telephone surveys are usually done without
notice.
Interviewers may be perceived as telemarketers and,
consequently, turn off respondents.
There is a high risk of respondents not being completely honest—
giving brief answers to end the call sooner or changing their
responses because they are speaking to someone directly.
2.4 Face-to-face surveys
Face-to-face surveys are one of the oldest and most widely used
survey types.
The researcher typically interviews in the home, office, hangout place,
etc. of the target respondent.
Benefits of face-to-face surveys
Interviewers can encourage respondents to answer all questions.
They can assist in case the respondent is confused about one or
more questions.
Interviewers can take advantage of the five senses of their
respondents. Aside from offering audio and visual stimuli, the
researcher can let respondents touch, taste, and smell materials
to support the interview.
Limitations of face-to-face surveys
Face-to-face surveys can take longer. They can last for weeks,
depending on the number of respondents needed.
Face-to-face surveys are considerably more expensive than
paper, online, and phone.
2.5 Mixed-mode surveys.
Mixed-mode surveys combine different ways (modes) of collecting
data for a single research project.
You may use mixed-mode survey designs to address problems
associated with the under-coverage of key groups of interest or
to improve response rates.
3. Informed consent
Certain laws and regulations (such as the GDPR) require that
respondents clearly agree upfront to participate in a survey study.
Therefore, all surveys must start with informed consent.
The following information must be communicated before a person
starts the survey:
The purpose of the study
What happens with the data after receiving them
Whether the data are confidential
Respondents' right to terminate the survey at any time
How and where respondents can obtain the results from the
study once it is completed
When conducting an online survey, you can include the informed
consent form on the first page the respondents see. On the landing
page, you can ask respondents to click “Yes, I agree” to give their
consent.
For a paper survey, you can include the informed consent information
in the cover letter that accompanies your survey.
If you are conducting a phone or a face-to-face survey, you should
have a script with the information that you read aloud and then ask if
respondents agree to participate.
4. Question order.
When constructing your questionnaire, the questions should not per
definition follow the order in which they are listed in your
operationalization table.
Some general principles about the order in which your questions
appear best:
1. Start with the most easy, straightforward questions.
2. Move to questions that require more thought.
3. Leave demographic and personal questions until the end.
5. Closing the questionnaire
At the end of your questionnaire:
you thank the respondent for completing the questionnaire
you restate who they may contact (name, email, phone) for any
questions they may have
in the case of a paper survey, you give instructions on how to
return the questionnaire
Sometimes, you may indicate that you will make a summary of your
research findings available, do not forget to follow up on it!
It is usually a good idea to leave an open-ended comment question at
the end to express any extra information the respondents want to
share.
D. Reliability and validity in survey research.
1. Measurement reliability in survey research
1.1 Single-item measures
Single-item survey measures for abstract constructs tend to have low
measurement validity.
Single-item measures often fail to capture the full breadth and
depth of an abstract construct.
Using only one item may overlook important nuances within the
construct, leading to low measurement validity.
1.2 multi-item measures
The measurement reliability of multi-item survey measures is
evaluated using Cronbach’s alpha.
2. Measurement validity in survey research
Three measurement validity threats that are specific to survey
research:
Response sets
Social desirability bias
Survey-mode bias
2.1 Response sets
Response sets are a shortcut people can take when answering a
series of survey questions.
For example, towards the end of a long survey, people might
answer all questions positively, negatively, or neutrally.
o Acquiescence bias or yea-saying occurs when people
consistently say "yes" or "strongly agree" to every question
instead of thinking carefully about it.
o Nea-saying occurs when people consistently say "no" or
"completely disagree" to every question instead of thinking
carefully about it.
o Fence-sitting occurs when people consistently choose the
middle neutral option.
In all three instances, measurement validity is hampered because the
survey does not measure the constructs it was intended to measure.
2.2 Social desirability bias
Social desirability bias is the tendency of survey respondents to answer
questions in a manner that others will view favorably.
Respondents may over-report "good behavior" or under-report "bad
behavior."
Solving social desirability bias – to minimize socially desirable
responding:
Deliberatively leading and/or loading the question to make the
sensitive “normal”
o “Everybody does it”
“Even the most truthful people may sometimes not
declare all income for taxes. Has this happened to
you?”
o “Assume-the-behavior”
“How often have you overeaten yourself in the past
week?”
o “Authorities-recommend-it”
“Doctors generally acknowledge that drinking wine in
moderation is beneficial. Did you drink wine
yesterday?
o “Reasons-for-doing-it”
“Did things happen, so that you could not go to the
dentist for that regular check-up, or did you go?”
In all other instances, avoid leading & loading
questions!!!
2.3 Survey-mode effects
A survey-mode effect is a systematic difference that is attributable to
the survey mode chosen.
It occurs when respondents answer at least some questions differently
depending on the survey mode used
Expert judgment & pilot testing to assess and improve measurement
validity:
Expert judgment - You can ask one or more experts to
comment on the extent to which your measures capture your
construct definitions. Based on their comments, you can amend
your questions before the next step: pilot testing.
Pilot testing - Second, before sending out your questionnaire to
collect data, you should ALWAYS pilot-test it. Through a pilot test,
you can ensure that respondents understand the questions as
they are intended.
o You should ensure that the number of people in the pilot is
sufficient to include major variations in the data that might
affect responses.
o As a rule of thumb, for most student surveys, the minimum
number of people to include in a pilot is 10.
3. Internal validity in survey research
Internal validity is the extent to which a study can rule out alternative
explanations.
To establish the true relationship between variables (e.g., X and Y), a
researcher needs to remove the influence of extraneous variables. The
less chance there is for "confounding" in a study, the higher the
internal validity and the more confident one can be in the findings.
The internal validity of survey research can be improved by including
questions related to control variables in your questionnaire!
This way, one can "filter out" or "isolate" the control variables' effects
from the relationship between the variables of interest.
4. External validity in survey research
If a survey's response rate is extremely low, doubts about the study's
external validity (generalizability) may arise.
4.1 Calculating the response rate
The survey response rate is the percentage of people in your sample
who successfully completed your survey.
To calculate your survey's response rate, you can use the following
formula:
4.2 What is a good response rate?
Typically, a response rate falls between 20% and 30%.
A survey response rate below 10% is considered very low.
A good survey response rate is anything above 50%.
4.3 Increasing the response rate
1. Send a gentle reminder:
o When you haven’t heard from a respondent, send one to
three reminders, using refreshed language each time so
you’re not simply repeating the original.
2. Offer to provide feedback on the results:
o It’s rewarding for respondents to see how their input is
utilized effectively.
3. Use incentives:
o Larger incentives for survey completion will generally
produce higher response rates. In general:
i. A small incentive for each respondent is better than
a large incentive for a few.
ii. The possibility to win a large reward produces lower
response rate than certainty in winning a small
reward.
4. Use a survey panel
o Panel providers manage a massive community of pre-
screened survey respondents who are ready to take
surveys.
It is important to realize that a low response rate is not always a
problem.
Non-response is only a problem if there are systematic differences
between the characteristics of respondents and non-respondents, and
if such differences affect the findings.
Problem:
How to compare respondents with non-respondents?
Solution:
Compare the characteristics of early respondents with those
of late respondents (e.g., respondents who only filled out the
survey after the final reminder).
o The idea is that the characteristics of these last-minute
respondents will resemble those who did not bother to
respond at all.
If you can conclude that early respondents do not differ
significantly from late respondents, you can infer that it
is unlikely that non-response bias will have biased your findings.
8. Archival research
A. Internal vs. external archival research
1. What is archival research?
Archival research is research based on archival data.
Nowadays most archival data are digitalized.
Archival research relies on internal or external archival data.
1.1 Internal archival data
Internal archival data are generated within the company that
conducts the research.
1.2 External archival data
External archival data are data that are generated by sources
outside a company that can be used by anyone.
o Publicly available archival data are available for free for
the entire world to use:
Government publications
World Bank
OECD
Annual reports
o Commercially available archival data has to be paid for
by someone.
Compustat, ORBIS (accounting data)
Datastream (stock price data)
SDC (alliance data)
Execucomp (CEO data, incl. executive compensation)
Nielsen data (consumer & retail panel data)
2. When is archival research used?
Archival research has four potential strengths, viz. its ability to:
learn from past successes and failures in the industry ("industry
wisdom").
Examine effects over time.
Examine effects across countries.
Examine socially sensitive phenomena unobtrusively.
B. Piecing together archival data
1. Piecing together archival data
Often, you need to piece together multiple archival data sets into one
new data set that you can subsequently use to address your research
questions.
When piecing together different archival datasets, make sure
that your unit of analysis corresponds.
Archival data can be and often is collected over time
o These longitudinal data is ideal to examine changes over
time.
C. Reliability and validity in archival research
1. Measurement reliability in archival research
Three sources of measurement unreliability that plague any type of
research, but are particularly common in archival research:
1. Missing data
2. Inaccurately recorded data
3. Fake data
Solutions:
Missing data (cross-sectional):
o Listwise deletion:
If missing = “close to 0” recoding missing to 0
o Mean-substitution:
Replace missing value for observation i and
variable j with average value on variable j for
all other observations.
Missing data (longitudinal):
o Interpolation:
Replacing the missing data with values based
on the data we have.
Inaccurately recorded data:
Inaccuracies that turn up as extreme data points
(outliers):
o Remove observation:
Run analyses with and without observation.
o Trim/truncate (in large data sets):
Remove a fraction of observation, e.g. 1% most
extreme observations.
o Inaccuracies that are not extreme.
Fake observations:
o Be critical, check:
Who collected the data?
When? Where?
For what purpose?
1.1 Composite measures
When you need to measure a less concrete conceptual variable for
which multiple potential measures exist, then multi-item measurement
instruments can be desirable in an archival research study.
Composite measure combines multiple measures into one,
provided that they have a high Cronbach's alpha (that is, they
should be interrelated as you consider them indicators of the
same underlying construct).
Combining multiple measures:
o Standardize each indicator:
Subtract the mean and divide by SD
o Average the standardized indicators
2. Measurement validity in archival research
Measurement validity pertains to whether measurement instruments
capture what they are supposed to measure.
An archival measure may only be a Proxy of the underlying construct
(approximation of construct)
How to validate that and archival measure is a good measure rather
than a bad proxy?
1. Provide precedence but use high quality studies.
2. Provide sound logic to support that considerable conceptual
overlap exists between construct and measure.
3. Provide evidence of a substantial correlation between your proxy
and a valid survey measure for a small subsample of data.
4. Provide evidence of substantial correlation (r > 0.3) with related
constructs (“nomological validity”)
3. Internal validity in archival research
Internal validity is the extent to which a study can rule out
alternative explanations.
Including control variables improves internal validity of archival
research.
4. External validity in archival research
You cannot assume that someone else took care of the external
validity.
You must always look up how the archival data were collected (e.g.,
what population was the data collected from?)
Always read the documents explaining the methodology behind
archival data bases.
9. Experimental Research
A. Lab vs. field experiments
1. Experiments
In an experiment, researchers manipulate at least one variable and
measure at least one other variable.
Two major types of experiments are:
Lab experiments
Field experiments
1.1 Lab experiments
A lab experiment is an experiment in an artificial environment
1.2 Field experiments
Field experiments are conducted in a natural environment.
In a field experiment the researchers manipulate the independent
variable(s) but:
the setting,
the participants,
the manipulation(s)/treatment(s), and
the outcome measures
are all authentic.
Field experiments are typically carried out unobtrusively, i.e., without
participants realizing they are participating in an experiment.
A/B testing is a special type of online field experiment.
2. When to use lab and field experiments?
An experiment is not equally suitable for every research question.
Experiments (lab or field) are particularly suitable when:
the number of independent (and moderator) variables is limited,
and
at least one of these variables can be manipulated.
A lab experiment is preferred when researchers want maximum control
over the research environment to rule out alternative explanations.
Because of this control, lab experiments, on average, show
higher internal validity than field experiments.
Field experiments are preferred when it is essential to measure real-
world behavior in real-world situations - high external validity is crucial.
o When studying the long-term effects of manipulation(s). The
immediate effect of a manipulated variable may differ from the
long(er)-term effect. A field experiment can run for a longer
period of time to study these long-term effects.
B. Experimental designs
1. Terminology
Measured variable is the dependent variable.
Manipulated variable is the independent variable.
Conditions are the levels of the manipulated variable.
Control group is a level of the independent variable that represents a
neutral condition.
o When a study includes a control group, the other levels of the
variables are called the treatment groups.
2. Within-subjects vs. between-subjects designs.
One of the most basic distinctions between experiments is within-
subjects designs and between-subjects designs.
o Within-subjects design - each subject (participant) is
presented with all levels of the independent variable.
o Between-subjects design - different groups of subjects are
assigned to different levels of the independent variable.
3. Two basic forms of between-subjects designs:
Posttest-only design – the simplest between-subjects design:
o Subjects are randomly assigned to the levels of the independent
variable.
o The dependent variable is then measured once.
Pretest/posttest design:
o Participants are randomly assigned to the levels of an
independent variable.
o The dependent variable is measured twice: once before and once
after exposure to the independent variable.
Pretest/posttest design offers the strongest form of experimental
control.
Why not always use the Pretest/posttest design?
o In rare circumstances, it may be problematic to measure the
dependent variable beforehand, as it may influence the second
measurement.
o If the pretest makes participants change their subsequent
behavior/reaction, a pretest should be avoided.
4. The factorial design:
Adding an additional independent variable allows researchers to look
for an interaction or moderator effect
When testing for interactions, the factorial design is used.
In a factorial design, researchers combine the two independent
variables: they study each possible combination of the independent
variables.
4.1 Within-subjects, between-subjects, and mixed factorial designs.
Researchers can manipulate each independent variable in a factorial
design as within-subjects or between-subjects.
In a between-subjects factorial design, both independent variables
are studied as between-subjects. Therefore, if the design is a 2x2
design, there are four different groups in the experiment. In other
words, there are different subjects in each cell, each of whom is only
subjected to one treatment.
In a within-subjects factorial design, both independent variables
are manipulated as within-subjects. Therefore, if the design is a 2x2
design, there is only one group in the experiment, but they participate
in all four cells (or combinations) of the design.
In a mixed factorial design, one independent variable is
manipulated as between-subjects, and one independent variable is
manipulated as within-subjects.
4.2 More than two levels of an independent variable.
The notation for a factorial design with two independent variables is "a
x b", where:
o a indicates the number of levels of the first independent variable
o b indicates the number of levels of the second independent
variable
4.3 More than two independent variables
Sometimes, research studies have three independent variables.
o Such a design is called a three-way design.
o For example, in a 2x2x2 factorial design, there are two levels of
each of the three independent variables.
o This leads to eight cells or conditions in the experiment (2x2x2 =
8).
Three-way factorial designs are complex to interpret. A three-way
interaction means that the two-way interaction between two of the
independent variables depends on the level of the third independent
variable.
C. Reliability and validity in experiments
1. Measurement reliability in experimental research
In experimental studies, the dependent variable is measured, and at
least one independent variable is manipulated.
The measured dependent variable is either:
Concrete (e.g., the number of M&Ms eaten)
o A single-item measure is typically sufficient.
Abstract (e.g., the perceived tastiness of the M&Ms)
o Multi-item measures must be used.
To demonstrate the internal consistency (reliability) of multi-item
measures, Cronbach's alpha is calculated.
If Cronbach's alpha is acceptable (>.70), the items in a measurement
instrument are internally consistent and therefore the measurement
instrument is reliable.
One can then average the scores on the items to create a construct
score for the dependent variable.
2. Measurement validity in experimental research.
The validity of measured variables (such as the dependent variable
in an experiment) can be demonstrated by
Providing precedence (has this measure been used before?)
Using sound arguments (why does this measure capture the
variable?)
The validity of manipulated variables can be demonstrated by
Providing precedence (has this manipulation been used before?)
Using sound arguments (why does this manipulation capture the
variable?)
Manipulation checks - tests used to determine the effectiveness
of manipulation in an experimental design and ensure that the
participants understood the manipulation as the researcher
intended.
3. Internal validity in experimental research
To assess the internal validity of an experimental study, you must
understand the three types of threats that can occur:
1. Design confounds:
2. Demand effects:
Demand effects arise when the participants guess the
purpose of the experiment and change their behavior
accordingly.
3. Experimenter bias:
Experimenter bias happens when the experimenters
(intentionally or unintentionally) influence the data,
participants, or results because they can't stay completely
objective.
3.1 Design confounds
Design confounds occur when an experiment is poorly designed, and
another variable happens to vary systematically along with the
independent variable. This provides an alternative explanation for the
observed effects.
1. Maturation effect:
Refers to natural (biological or psychological) changes in
participants during a study, like becoming tired, bored, or
hungry, or learning and forgetting things over time. These
changes occur simply because time passes, there is no
outside intervention.
Does maturation always threaten a study's internal validity?
o The internal validity is not threatened if the natural
over-time changes do not provide an alternative
explanation.
o Internal validity is threatened if the maturation effect
provides an alternative explanation.
Solution: introduction of control group that is measured at the
same time but is not exposed to the treatment.
2. History effect:
Refers to an (unforeseen) “historical” or external event that
occurs during the course of a study, between the pretest and
posttest.
This could be a large-scale event or a small-scale event
(something that goes wrong during data collection).
To have a history effect (or a history threat), external factors
must systematically affect most members of the treatment
group at the same time as the treatment itself.
o Solution: Introduction of control group.
3. Testing effect:
Refers to administering participants a test or a measurement
procedure.
A testing effect (or testing threat) happens when the
measurement procedure affects participants and thereby
provides an alternative explanation for the results - a change
in participants' responses as a result of taking a test more
than once.
Participants become more practiced at taking the test (leading
to higher scores over time) or because they become fatigued
or bored (leading to lower scores over time).
The act of being tested can sensitize participants to the
material, methods, or stimulus in further tests, thereby
altering their responses in ways that do not reflect the true
impact of the experimental condition.
o Solution: introduction of control group
o Adding a control group is not always enough, though. In
some cases, there is a risk that the pretest sensitizes
people in the control group differently than people in
the experimental group.
o One solution is to add an experimental and control
group that weren't given a pretest or abandon a pretest
altogether and use a posttest-only design.
4. Instrumentation effect:
Instrumentation effect (or threat) occurs when
a measurement instrument is changed during the course
of the study.
This change in the instrumentation can lead to differences in
how participants respond (the dependent variable) that are
not related to the experimental manipulation.
Example: switching survey formats (e.g., from paper to online)
Solution:
o Researchers should aim to maintain consistency in the
measurement instruments throughout the study.
o If a change is necessary, it is crucial to document and
account for it in the analysis.
o Including a control group exposed to the same changes
in instrumentation
5. Selection bias:
Selection bias occurs when the characteristics of participants
in one group are systematically different from those in
another group.
Solution:
o Random selection.
6. Mortality effect:
Mortality (also referred to as attrition) refers to participant
drop-out before the end of the study.
Mortality effect (or threat) occurs when this dropout leads to
systematic (i.e., non-random) differences between those who
remain and those who drop out.
Solution:
o Researchers must always ensure that attrition is not an
explanation for their results; for example, by checking
whether dropouts differ from completers.
o If the attrition is not systematic the researchers can
conclude that attrition is not a threat to internal validity.
o If the attrition is systematic, In most cases, the best a
researcher can do is document the reasons for dropout
so that these reasons can be investigated and possibly
mitigated in further studies.
o The threat of mortality is very hard to eliminate.
3.2 Demand effects:
Demand effects may occur when participants are aware of whether
they were assigned to the treatment or control group.
Solution:
o To reduce the likelihood of demand effects, a single-
blind experiment can be conducted.
o In a single-blind experiment, participants are unaware of
which group they have been assigned to until after the
experiment is completed.
3.3 Experimenter bias:
When the researchers administering the experimental treatment are
aware of each participant’s group assignment, they may inadvertently
treat those in the control group differently from those in the treatment
group.
Solution:
o Double-blind study, neither the participants nor the
experimenters know which participants are in the
treatment group and who is in the control group.
4. External validity in experimental research:
If we find evidence for a causal relationship in an experiment, can we
conclude that the same causal relationship generalizes to other people,
places, and times?
Not necessarily, the types of participants that participate in a lab
experiment may be very different from the population as a whole.
5. Balancing internal and external validity in experimental research:
It is possible to realize both internal validity and external validity in a
study, doing so is difficult.
If an experiment's internal validity is high because all possible
confounding variables can be controlled for, the experimenters know
that a cause-and-effect relationship is almost certainly true. However,
the artificial environment in a lab has low external validity.
Therefore, researchers may first run a lab experiment to show that a
causal relationship holds, followed by a field experiment to show that
the relationship generalizes to the real world.
6. Random sampling vs. random assignment
Random sampling is the random selection of subjects from a
population. It, therefore, pertains to the study's external validity.
Random assignment involves randomly assigning each subject in the
sample to the experimental conditions. Therefore, it pertains to the
study's internal validity.
10. Quality Research
A. The Basics of Quality Research
1. What is quality research?
Qualitative research is a research strategy that collects and analyzes non-
numerical data: words rather than numbers.
Qualitative data can be primary or secondary:
Primary qualitative data are qualitative data that are collected first-
hand by the researcher for a specific research purpose. Researchers can
collect primary qualitative data through interviews, focus groups, or
observations.
Secondary qualitative data are qualitative data that have already
been collected by someone else for a different purpose. The researcher
reanalyzes these data for a new purpose. Secondary qualitative data
sources include written texts, such as company and financial reports,
press releases, blogs, etc.
Qualitative research is helpful to generate insights into less mature
topics: in order to clarify key constructs and develop new theoretical
frameworks.
Qualitative research is typically inductive rather than deductive, because
the researcher develops a theoretical framework during empirical research.
Quantitative and qualitative research studies are not each other’s opposites
and they complement each other.
B. Collecting Primary Qualitative Data
Three well-known data collection options for primary qualitative data:
1. Interviews
2. Focus Groups
3. Observations
1. Interviews
Interview is a conversation where the researcher asks questions and listens
while the respondent answers.
Types of interviews:
Structured interview has a carefully worded set of interview
questions. The interviewee can typically be brief in his/her
responses.
Unstructured interview - the interviewer does not have a
planned sequence of questions to be asked to the interviewee.
o The interviewer usually begins the interview with a broad
question.
o Next questions are very much dependent on the answers
given by the interviewee.
Semi-structured interview is a hybrid form of the structured
and unstructured interview approach.
o semi-structured interview is based on a set of
predetermined questions, but leaves room for the
interviewee to elaborate on his responses and for the
interviewer to introduce additional questions based on the
interviewee’s answer.
o Most popular choice when collecting interview data.
Designing a semi-structured interview (4 stages):
1. Setting the scene:
Introduce yourself
Briefly inform the interviewee about the purpose of the
interview and why s/he was chosen to be among those
interviewed.
Ask for permission to audiotape the interview and assure
confidentiality: explain that the interviewee's anonymity will
be preserved.
2. Warm-up Questions:
A few warm-up questions: easy-to-answer, non-sensitive
questions.
3. Interview:
Ask the main questions of interest, organized per subtopic.
Start with an open-ended question.
Based on answers continue with probing questions.
(Probing questions are follow-up questions that help the
interviewee to think through issues.)
4. Summarizing:
Summarize/rephrase important information given by the
interviewee, to make sure you interpreted his/her answers
correctly.
Interview data:
After the interview takes place, you must transcribe it.
Reproduce exactly what you and the interviewee have said
in the language in which the interview was conducted.
Advisable to transcribe interviews immediately as then it is
still possible to recall the interviewee's non-verbal cues
during the interview.
On average, one hour of interviewing takes four hours of
transcribing.
2. Focus groups
Focus group is an unstructured interview conducted by
a moderator with a small group of participants, varying from
approximately 8 to 14 participants.
Moderator asks questions in an interactive setting where
participants are free to talk to each other.
Listening to others expressing their experiences and ideas
stimulates participants to express their own opinions.
2.1 The Role of the Moderator in focus groups:
o Ensure that all members participate in the discussion and that no
member dominates the group.
Designing a focus group:
1. Setting the scene:
o Moderator introduces the topic and explains the purpose of the
focus group.
o Ask for consent to record the focus group. It is very important
that the participants taking part in the focus group understand
how recordings will be used.
2. Introduction:
o Ask everyone to introduce themselves, stating their first name
clearly (important for writing up the recording later).
o Moderator can also ask participants to give a short answer to
an introductory question to get everyone involved in the
discussion from the outset.
3. Discussion:
o A topic guide needs to be planned in advance (outlines the
areas for discussion during the focus group, with key ideas and
questions to be discussed)
o Useful to construct the topic guide with the thought of a
conversation in mind rather than interview questions.
o Therefore, include topic questions, possibly with areas for
prompting rather than exact questions.
o Moderator should also be prepared to tactfully steer the group
back to the topics under consideration if the conversation goes
too much off track.
4. Closing round:
o Good to end the discussion with a ‘closing round’, asking each
participant, in turn, to offer final reflections or answer a final
question.
o This is followed by informing the participants of the next steps
and how they can stay informed or involved with the research.
Focus-group data:
Tape Recording
Transcript of those recordings
Moderator’s notes from the discussion
Choosing between focus-group and interviews:
Focus group is preferred:
o When interaction helps, people can build on each others answers
o When respondents can say what is relevant in <10min
Interviews:
o When interaction hurts (for example, sensitive topics)
o When detailed answers are needed:
Complex topics
Expert respondents
3. Observational studies
Observational studies involve systematically recording the behaviors of
small groups of people in their natural surroundings.
Dimensions of observational research:
1. First dimension pertains to whether the researcher's identity is
revealed (overt observation) or concealed (covert observation)
during the study.
2. Second dimension pertains to the extent to which the researcher
participates in the activities of the organization that s/he is
observing.
Leads to four types of observation research that are labeled:
1. Complete participant
2. Complete observer
3. Observer as participant
4. Participant as observer
Complete participant
In the complete participant role, the researcher:
o Tries to become a member of the group which s/he is
researching
o Does not reveal his/her true purpose to those he is observing
Complete Observer
In the complete observer role, the researcher:
o Does not take part in the activities of those s/he is observing
o Does not reveal his/her purpose to those he is observing
Observer-as-participant
In the observer-as-participant role, the researcher:
o Does not take part as a member of the group which s/he is
observing
o Reveals his/her true purpose to those he is observing
Participant-as-observer
In the participant-as-observer role, the researcher:
o takes part as a member of the group which s/he is observing
o reveals his/her true purpose to those he is observing
Observational data:
Note-making is very important in observational studies.
Your notes must consist of:
Primary observations: notes about what happened or what was
said.
Experiential data: notes on your perceptions and feelings as you
experience the process you are researching.
Contextual data: notes on the research setting (e.g.,
organizational structure, communication patterns, ...) that may
help you interpret other data
Choosing between Observations and Interviews:
Observations are preferred:
o To provide direct information about subjects’ behavior
o When directly asking subjects would lead to distorted
information.
Interviews are preferred:
o To identify the reasons underlying subjects' behavior
o When observation would affect subjects' behavior.
C. Validity in Qualitative Research
1. Threat to internal validity in qualitative research
Two common threats to internal validity in qualitative research are:
Researcher bias is the influence of researchers' prior knowledge and
assumptions on their study.
Respondent bias refers to participants not providing honest
responses to the researcher.
Researchers bias:
Researcher bias occurs when the researcher skews the process
toward a specific research outcome by introducing a systematic error
in the sample data.
The results then deviate from the true outcomes.
An example is selective perception and interpretation, the
researcher perceives what he wants to hear in a message while
ignoring opposing viewpoints.
Respondent biases:
Two types of respondent biases are Authority bias and Conformity
bias.
o Authority bias - the tendency to blindly follow or believe the
instructions and views of a person in authority.
If the interviewer, moderator in a focus group, is
perceived as a person in authority by the interviewee,
the latter may adapt his answers or behavior in an
attempt to come across as obedient.
o Conformity bias occurs when individuals sway their opinion
to match the opinion of the majority.
In a focus group, the participants who express their
opinions first inherently influence the responses of the
others.
Can be controlled by systematically varying the order in
which participants speak and by encouraging debate and
controversial opinions.
The moderator should convince participants that no
stupid or wrong answers exist.
2. How to increase the internal validity of a qualitative research study?
The following strategies can be used:
1. Triangulation - the research will be conducted from different or
multiple perspectives.
a. can take the form of using several moderators or different
locations, or using multiple individuals to analyze the same
data.
2. Peer debriefing refers to receiving feedback from other people at
different stages of your research.
3. Member checking refers to testing the emerging findings with your
research participants. This can be done in multiple ways:
a. You may send your participants the interview transcripts and
ask them to read these transcripts and provide comments or
corrections.
b. You may send participants an e-mail and ask them to verify
your interpretations before you jump to conclusions.
c. You may schedule a 'validation' interview. This is a follow-up
interview.
4. Negative case analysis is analyzing those cases that do not match
the trends or patterns emerging from the rest of the data.
3. External validity
External validity represents a problem for qualitative researchers because
they tend to use small samples.
Generalizability can be enhanced by doing a thorough job of describing the
research context and the assumptions that were central to the research.
The person who wishes to “transfer” the results to a different context is
then responsible for judging how sensible the transfer is.
D. Coding qualitative data
Qualitative data analysis is the process of disaggregating qualitative
data into smaller parts, called information chunks, and
then reconnecting these smaller parts into concepts.
Qualitative data analysis involves deconstructing and
reconstructing the data.
The challenge of qualitative research is that there are no widely
accepted rules about how qualitative data must be analyzed.
The approach involves coding data.
There are different approaches to coding one’s qualitative data.
We focus on the approach based on grounded theory.
Grounded theory is an approach where one builds one's theory
from the ground up, starting from the data.
How to proceed to de- and reconstruct qualitative data?
1. Step, Re-read your interview transcripts
Re-read your interview transcripts not just literally but also spend
some time dwelling on the underlying meanings.
2. Step, Code
Coding means assigning labels to qualitative data.
When coding, we try to identify patterns in qualitative data.
Occurs in three cycles:
1) Open coding
o Aim: Reducing data
o Making extensive transcripts shorter by converting them into
concise codes
2) Axial coding:
o Aim: Identify concepts
o Reconstructing data
3) Selective coding:
o Aim: Creating theory
o From concepts from the second cycle, we develop a theory.
Cycle 1: Open coding:
Starts with preparing the transcripts for coding
o Create a new word or excel document
Start each (sub)sentence on a separate line
Analyze each line for notable information fragments
Label the information:
o Assign a provisional code to each information fragment
o Rule of thumb:
Summarize each information fragment into 1-3 words,
including one word from the text
Reduce the long list of codes to a shorter one:
o Look for similar codes
o Rule of thumb: 20-25
Cycle 2: Axial coding:
Identify concepts/themes:
o Concepts/themes = codes that share similarities
o Move away from the terms used by the respondents
o Rule of thumb:
5-7 concepts/themes
Be exhaustive
o Concepts/themes need to fully reflect everything present in the
data
Cycle 3: Selective coding:
Identify the core concepts
Formulate a theory about the relationships between these concepts
Selective coding thus means identifying the cause of the phenomena
and its consequences
Leads to writing a research report
Open coding and axial coding are done repeatedly.
Remember that after conducting x number of interviews do not begin open
coding all of them.
Better to start by open coding a small number of interviews which will yield a
set of codes that gives you a reasonably good idea of the concepts and
themes to look for in the rest of the data.
These initial codes can help refine your questions for the next batch of
interviews.
Software for qualitative data analysis
There's software out there to help you crunch through qualitative data:
Kwalita
Nvivo
ATLAS
This software will not do the analysis for you, it may help you to stay
organized.
E. Content analysis
Qualitative data can sometimes be analyzed quantitatively.
Content analysis is a research tool for analyzing the presence, meanings,
and relationships of certain words, themes, or concepts in qualitative data
(e.g., newspaper articles, annual reports, ...)
Quantifying the qualitative data involves:
Counting the frequency with which words/themes occur
o Also referred to as Content Analysis.
Requirement: Large samples
o Samples need to be large enough to allow for statistical analysis:
Minutes of meetings between companies and trade unions
Announcements to shareholders
Newspaper articles
News bulletins on TV
Post on social media
Content analysis consist of two steps:
1. Determine the unit of coding
a. Word
o Example: “conflict”, “tasty” & “not tasty”
b. Theme
o Example: resolution of a conflict, SMEs
c. Time, space
o Example: time on TV news, space in newspapers
2. Count the frequency of occurrence:
How often does the word “conflict” occur in the minutes of union
meetings?
How many newspaper articles are about SMEs?
It is not always necessary to quantify qualitative data.
For instance, when conducting in-depth interviews to identify potential
explanatory variables, it is not meaningful to quantify the data from these
interviews.