Understanding Structural Equation Modeling
Understanding Structural Equation Modeling
(Introduction)
Psych 203
Quantitative Research in Psychology
Topic #10
X1 X2
Work Turnover
Engagement
Intent
Org’l
Identity
Sex
Job Engagement
Satisfaction
Linear Regression
Path Analysis
X1 X15
Job
X2 Satisfaction
Work X16
Engagement
X3
Org’l X17
X4 Identity
X5
X1 X15
Job
X2 Satisfaction
Work X16
Engagement
X3
Org’l X17
X4 Identity
X5
Climate
ENDOGENOUS constructs
• act as dependent variables in the model
• Determined by factors within of the model
• Path (one-headed arrow) towards them are present
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Advantages
• SEM can estimate multiple interrelated dependence
relationships
• SEM can Incorporate latent variables into the analysis
• Theoretical concepts (e.g. satisfaction, quality, organizational
climate) are better represented using multiple measures
• Statistical estimation of the relationships between concepts
are improved by accounting for measurement errors (SEM
deals with this in the analysis of the measurement model)
• Measurement Model
2
• Structural Model
3
• Convergent Validity
2
• Construct Reliability
3
• Test hypotheses on
individual path/regression
2 weights
• Describe coefficient of
determination (R2)
3
Evaluate Modify/
Specify Estimate (how well the
the model the parameters data fit the Respecify the
model) model
Specify
Specify the Estimate the Evaluate (how
well the data fit
Modify/
Respecify the
model parameters the model) X1 X2
the model model
e3 X3
X6 X7 X8 X9 X10 X11
e6 e7 e8 e9 e10 e11
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Evaluate (how Modify/
Specify the
model
Estimate
Estimate the
parameters
well the data fit the Respecify the
model) model
the parameters
Job
Satisfaction
.28*
.30*
.09 .32*
.35* Org’l Work Turnover
Identity Engagement
Intent
.42*
.45*
Climate
Job
Satisfaction
.28*
.30*
.09 .32*
.35* Org’l Work Turnover
Identity Engagement
Intent
.42*
.45*
Climate Χ2 =207.31, df=84 p<.001
NFI=.90, NNFI =.89
CFI=.92, RMSEA=.053, CI=.050, .056
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Specify the Estimate the Evaluate (how
well the data fit the
Modify
Modify/
Respecify the
model parameters Respecify
model) model the
model
Job
Satisfaction
.28*
.30*
.42*
.45* .21*
Climate
Χ2 =207.31, df=84 p<.001
NFI=.90, NNFI =.89
CFI=.92, RMSEA=.053, CI=.050, .056
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Kinds of Research Questions
• Adequacy of the model
• Testing theory
• Amount of variance in the variables accounted for by
the factors
• Reliability of the indicators
• Parameter estimates
• Mediation
• Group differences
• Longitudinal differences
• Multilevel Modeling
X1 X2
TURN1 TURN2
Lavaan Basics: JobSat
• Spacing does not matter
• New lines should be included for each new
“equation” Engagement Turnover
• =~ creates latent variables (is measured by)
• Latent =~ variable1 + variable2…
• You name it =~ names from the dataset OrgID
• Prediction is that latent variable predicts
Eng_Vig Eng_Ded Eng_Abs
the manifest variables
• <~ creates latent variables (is predicted by)
• Latent <~variable1 + variable2 IDORG2 IDORG3 IDORG4 IDORG5
• Prediction is that latent variable is
predicted by the manifest variables # measurement model
• ~indicates regression
• Y~X Jobsat =~ JOBS1 + JOBS2 + JOBS3
• Names need to be defined in dataset or OrgID =~ IDORG2 + IDORG3 + IDORG4 + IDORG5
previously defined latents
• ~~indicates covariance/correlation Engagement =~ Eng_Vig + Eng_Ded + Eng_Abs
• Does not matter which order you write it in Turnover =~ TURN1 + TURN2
• Names need to be dfined in dataset or
previously defined latents
• You can set specific variances by using: # regressions
• variable ~~ NUMBER*variable Engagement ~ Jobsat + OrgID
• ~1 indicates intercept
• Variable ~ 1 Turnover ~ Engagement
Jobsat ~~ OrgID
Buchanan, E. M. (2020, March 15). Structural
EquationModeling. Psych 203 (Quantitative Research in Psychology)
[Link] Marshall N. Valencia (mnvalencia@[Link])
cross-loading of error
Specify
Specify the
model
Estimate the
parameters
Evaluate (how
well the data fit
Modify/
Respecify the covariance
the model)
the model model
VIG1 E1
Issues / Considerations VIG2 E2
1. Unidimensionality of measures Vigor
VIG3 E3
• Indicators (measured variables
VIG4 E4
should relate/load only to a
single construct
• Cross loadings are indications of Dedi-
DED2 E5
Issues / Considerations
2. Items per construct
– Good practice: minimum of 3 items per factor, preferably 4 (Hair,
et al 2010)
– Single item constructs are problematic (construct validity issues
and “identification” issues)
– Some concepts may be very simple that they can be captured by
a single item (e.g. directly observable behaviours such as job
hopping, purchase, and consumption); when in doubt, use
multiple items
– If you have several items (for example you have 9 items), you can
reduce them to 3 “parcels” (get the average of 3 items to
represent one indicator)
X1 X2 X3 X1 X2 X3 X4
X1 X2
E1 E2 E3 E1 E2 E3 E4
E1 E2
6 parameters to estimate 8 parameters to estimate
4 parameters to estimate X1 X2 X3 X4
X1 X2 X1 X2 X3 X1 Var1
X1 Var1 X1 Var1 X2 Cov1,2 Var2
X2 Cov 1,2 Var2 X2 Cov1,2 Var2 X3 Cov1,3 Cov2,3 Var 3
X3 Cov1,2 Cov2,3 Var 3 X4 Cov1,4 Cov2,4 Cov3,4 Var4
3 unique terms
6 unique terms 10 unique terms
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Dealing with Identification
problems
To compute the number of
• 3-indicator rule • Underidentified models will not unique variances/ covariances:
• “set the scale”of each construct generate a unique solution ½ [p(p+1)]
(in AMOS, one indicator is
always set to 1.0) • Just-identified models will have a 0 where p is the number of indicators
degrees of freedom thus it will have a (or measured items)
perfect fit; it’s a “saturated” model
with a chi square of 0.
• Overidentified models have more
unique covariance and variance terms df (degrees of freedom)
than parameters to be estimated; a …amount of mathematical
unique solution can be generated information available to
with positive df and a corresponding estimate model parameters
χ2 value df=½ [p(p+1)] – k
where k is the no. of estimated
• “error” messages in the EQS may be parameters
due to identification problems
X1 X2 X3 X1 X2 X3
X1 Var1 X1 Var1
X2 2 Var2 Observed minus Expected X2 1.5 Var2
X3 3 1 Var 3 X3 3.5 1 Var 3
ERROR (residual)
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Evaluate (how Modify/
Specify the
model
Estimate
Estimate the
parameters
well the data fit the Respecify the
model) model
the parameters
Issues / Considerations
1. Type of input data
• Metric (interval or ordinal) -- covariances can be computed;
some recent versions of softwares allow nonmetric data
(binary, nominal)
• Covariance or correlation matrix?
• When raw data is used, the software takes care of
generating the covariance and correlation matrix
• Correlational matrix can be derived from covariance matrix;
standardized solutions (from correlational matrix) are
easier to interpret
• Suggestion: use covariance matrix whenever possible
Issues / Considerations
2. Missing data
• More than 10% of data items are missing – serious
considerations…
• Remedies:
• Complete case (listwise)
• All-available (pairwise)
• Model-based (ML/EM)
• Full information maximum likelihood (FIML)
• If missing data are random, less than 10%, high factor loadings
(.7 or above), any approach is appropriate; otherwise see Hair et
al (2010) or Enders & Bandalaos (2001)
Issues / Considerations
3. Multivariate normality of data
• Deviations from multivariate normality of data impacts on the
estimates of the parameters
• Screen variables for outliers (univariate and multivariate),
skewness and kurtosis
Remedies:
• Apply data transformations
• exclude cases that contribute to non-normality
• Use estimation procedures that handle non-normal data (e.g. “ADF”
method )
Issues / Considerations
4. Linearity
• SEM techniques examine only linear relationships
• Remedy: raise the measured variables to powers (e.g. square the
scores)
5. Multicolinearity and Singularity
• Extremely high correlations (>.90) between measured variables
• Inspect the determinant of the covariance matrix – extremely small
determinant may indicate multicollinearity or singularity (usually
EQS will abort -- error)
• Remedy: delete the variable causing the singularity or create
composite variables
6. Residuals
• Residuals should be small and centered around zero. The frequency
distribution of the residual covariances should be symmetrical
• Non symmetrical distributions may signal poor fit
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Evaluate (how Modify/
Specify the
model
Estimate
Estimate the
parameters
well the data fit the Respecify the
model) model
the parameters
Issues / Considerations
7. Sample size Minimum Model Complexity Measurement Model Characteristics
• Provides a basis for the Sample Size
estimation of sampling
error 100 5 or fewer constructs More than 3 items (or observed
• How large a sample is variables) per construct
needed to produce High item communalities (.6 or
trustworthy results? higher)
Considerations: 150 7 or fewer constructs Modest communalities (.5) and no
• Multivariate normality underidentified constructs
of data
300 7 or fewer constructs Lower communalities (below .45),
• Estimation technique
and/or multiple underidentified
• Model complexity (fewer than 3 items) constructs
• Amount of missing
data 500 Large number of Some with lower communalities,
constructs and/or having fewer than 3 items
• Average error variance
among the reflective
indicators
Psych 203 (Quantitative Research in Psychology)
Marshall N. Valencia (mnvalencia@[Link])
Step 3: Evaluate (how well the data fit
the model)
X1 X2 X3 X1 X2 X3
X1 Var1 X1 Var1
X2 2 Var2 Observed minus Expected X2 1.5 Var2
X3 3 1 Var 3 X3 3.5 1 Var 3
ERROR (residual)
Adjusted goodness-of-fit At least .90 is adequate (Bollen, 1990), values > 0.95
index (AGFI) is excellent
Root Mean Square Residual Values less than 0.08, values < 0.05 is excellent
(RMR)
Standardized root mean Values less than 0.08, values < 0.05 is excellent
square residual (SRMR)
Root Mean Square Error of Values .08 to .10 is mediocre fit (Mac Callum et al,
Approximation (RMSEA) 1996) adequate, less than .06 is good (Hu & Bentler,
1999), less than .05 is excellent (Steiger, 2007)
m ≤ 12 12< m <30 M ≥ 30
Chi-square χ2 n.s. p-values even with Sig p-values expected Sig. p-values
good fit expected
For N>250
CFI or TLI .95 or better Above .92 Above .90
RNI .95 or better, not used Above .92, not used Above .90, not used
with N>1,000 with N>1,000 with N>1,000
SRMR Biased upward, use other .08 or less (with CFI .08 or less (with CFI
indices or .92) above .92)
RMSEA Values < .07 (with CFI=.97 Values < .07 (with Values < .07 (with
or higher) CFI=.92 or higher) CFI=.90 or higher)
Psych 203 (Quantitative Research in Psychology)
(Hair, et al, 2010) Marshall N. Valencia (mnvalencia@[Link])
JOBS1 JOBS2 JOBS3
TURN1 TURN2
JobSat
Engagement Turnover
Chi-square χ2
Relative χ2 (χ2/df) OrgID
CFI (>=.90)
IFI (>=.90)
NFI (>=.90)
TLI (>=.90)
RMSEA (<=.08)
TURN1 TURN2
JobSat
Engagement Turnover
OrgID
Eng_Vig Eng_Ded Eng_Abs