PLS-SEM TUTORIAL
By Uliyatun Nikmah
Structural Equation Modeling
(SEM)
Multivariate analysis involves the application of
statistical methods that simultaneously analyze
multiple variables.
SEM is a class of multivariate techniques that
combine aspects of factor analysis and regression,
enabling the researcher to simultaneously examine
relationships among measured variables and latent
variables as well as between latent variables.
The most widely applied method is certainly
covariance-based SEM (CB-SEM). But there is a
variance-based partial least squares SEM (PLS-
SEM) approach, an alternative technique for
SEM.
PLS-SEM should not simply be viewed as a less
stringent alternative to CB-SEM but rather as a
complementary modeling approach to SEM.
Multivariate Methods
Organization of Multivariate Methods
Primarily Primarily
Exploratory Confirmatory
First-generation • Cluster analysis • Analysis of
techniques • Exploratory factor variance
analysis • Logistic
• Multidimensional regression
scaling • Multiple
regression
• Confirmatory
factor
analysis (CFA)
Second- Partial least squares Covariance-based
generation structural equation structural equation
techniques modeling (PLS- modeling (CB-SEM)
SEM)
Confirmatory vs Exploratory
test
Confirmatory when testing the
hypotheses of existing theories and
concepts
Exploratory when searching for latent
patterns in the data in case there is no
or only little prior knowledge on how
the variables are related.
Terms in PLS
Path models are diagrams used to visually display
the hypotheses and variable relationships that are
examined when SEM is applied
Constructs are variables that are not directly
measured
Indicators, also called items or manifest variables,
are the directly measured proxy variables that
contain the raw data.
Outer model measurement model (there are two
models: Outer model of Exogenous and Endogenous
latent variable)
Inner model structural model
Theory: measurement and structural theory
Errors
Path Model
Reflective vs Formative
Measures
Reflective vs Formative
Measures
Reflective vs Formative
Measures
Key Characteristics of PLS-SEM
Data Characteristics
Sample Size • Neglectable identification issues with small sample
sizes
• Achieves high levels of statistical power with small
sample sizes
• Larger sample sizes increase the precision (i.e.,
consistency) of PLS-SEM estimations
Distribution • No distributional assumptions; PLS-SEM is a
nonparametric method
• Influential outliers and collinearity may influence the
results
Missing Value Highly robust as long as missing values are below a
reasonable level (less than 5%)
Scale of • Works with metric data and quasi-metric (ordinal)
Measurement scaled variables
• The standard PLS-SEM algorithm also accommodates
binary coded variables, but additional considerations
are required when they are used as control variables,
moderators, and in the analysis of data from discrete
choice experiments
Model Characteristics
Number of items Handles constructs measured with single- and multi-
in each item measures
construct’s
Key Characteristics of PLS-SEM
Model Characteristics
Model Handles complex models with many structural model
complexit relationships
y
Model No causal loops (no circular relationships) are allowed in the
setup structural model
Model Estimation
Objective Aims at maximizing the amount of unexplained variance in the
dependent measures (i.e., maximizes the R² values)
Efficiency Converges after a few iterations (even in situations with complex
models and/or large sets of data) to the optimum solution (i.e.,
the algorithm is very efficient)
Nature of Viewed as proxies of the latent concept under
constructs investigation, represented by composites
Construct • Estimated as linear combinations of their indicators (i.e., they
scores are determinate)
• Used for predictive purposes
• Can be used as input for subsequent analyses
• Not affected by data limitations and Inadequacies
Parameter • Structural model relationships are generally
estimates underestimated, and measurement model relationships
are generally overestimated when
solutions are obtained using data from common factor
Key Characteristics of PLS-SEM
Model Evaluation
Evaluation of the The concept of fit—as defined in CB-SEM—does not
overall model apply to PLS-SEM. Efforts to introduce model fit
measures have generally proven unsuccessful
Evaluation of the • Reflective measurement models are assessed on the
measurement grounds of indicator reliability, internal consistency
models reliability, convergent validity, and discriminant validity
• Formative measurement models are assessed on the
grounds of convergent validity, indicator collinearity,
and the significance and relevance of indicator weights
Evaluation of the • Collinearity among sets of predictor constructs
structural model • Significance and relevance of path coefficients
• Criteria to assess the model’s in-sample (i.e.,
explanatory) power and out-of-sample predictive power
(PLSpredict)
Additional Methodological research has substantially
analyses extended the original PLS-SEM method by
introducing advanced modeling, assessment, and
analysis procedures.
Which one to use: PLS-SEM / CB-
SEM???
Use PLS-SEM when:
The goal is predicting key target constructs or identifying key
"driver" constructs.
Formatively measured constructs are part of the structural
model. Note that formative measures can also be used with
CB-SEM, but doing so requires construct specification
modifications (e.g., the construct must include both formative
and reflective indicators to meet identification requirements).
The structural model is complex (many constructs and many
indicators).
The sample size is small and/or the data are non-normally
distributed.
The plan is to use latent variable scores in subsequent
analyses.
Use CB-SEM when:
The goal is theory testing, theory confirmation, or the
comparison of alternative theories.
PLS-SEM vs CB-SEM
Distinctions
Determining Sample Size
Researchers commonly use this method
in determining sample size in PLS-SEM
where sample size should be equal to the
larger of:
1. 10 times the largest number of
formative indicators used to measure a
single construct, or
2. 10 times the largest number of
structural paths directed at a particular
construct in the structural model.
Cohens’s Statistical Power
Analysis
See also: G*Power
Higher Order Variable
Bootstrapping
PLS-SEM relies on a nonparametric bootstrap
procedure to test coefficients for their
significance. In Bootstrapping, a large number of
subsamples (i.e., bootstrap samples) are drawn
from the original sample with replacement.
Replacement means that each time an
observation is drawn at random from the
sampling population, it is returned to the
sampling population before the next observation
is drawn.
As a rule, 5,000 bootstrap samples are
recommended
Bootstrapping Procedure
Evaluating Structural
Output
VIF value > 0.20 (less than 5)
Number of bootstrap samples must be at least as large as
the number of valid observations but should be 5,000
Critical values for a two-tailed test are 1.65 (significance
level = 10%), 1.96 (sig. level = 5%), and 2.57 (sig. level = 1
%). In applications, you should usually consider path
coefficients with a 5% or less probability of error as
significant.
R2 values of 0.75, 0.50, or 0.25 for the endogenous
constructs can be described as respectively substantial,
moderate, and weak
The f2 values of 0.02, 0.15, and 0.35 indicate an exogenous
construct's small, medium, or large effect, respectively, on
an endogenous construct
Q2 values > 0 indicate that the exogenous constructs have
predictive relevance for the endogenous construct under
consideration. As a relative measure of predictive relevance
Referensi
Hair, J. F., Hult, G. T. M., Ringle, C. M., & Sarstedt, M. (2022). A
Primer on Partial Least Squares Structural Equation Modeling
(PLS-SEM), 3rd ed. Thousand Oaks, CA: Sage.
[Link]
[Link]
pls/
[Link]