100% found this document useful (1 vote)
37 views20 pages

Guide To Instrument Development

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
100% found this document useful (1 vote)
37 views20 pages

Guide To Instrument Development

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Practical Assessment, Research, and Evaluation

Volume 26 Volume 26, 2021 Article 1

January 2021

A Practical Guide to Instrument Development and Score


Validation in the Social Sciences: The MEASURE Approach
Michael T. Kalkbrenner
New Mexico State University - Main Campus

Follow this and additional works at: [Link]

Recommended Citation
Kalkbrenner, Michael T. (2021) "A Practical Guide to Instrument Development and Score Validation in the
Social Sciences: The MEASURE Approach," Practical Assessment, Research, and Evaluation: Vol. 26,
Article 1.
DOI: [Link]
Available at: [Link]

This Article is brought to you for free and open access by ScholarWorks@UMass Amherst. It has been accepted for
inclusion in Practical Assessment, Research, and Evaluation by an authorized editor of ScholarWorks@UMass
Amherst. For more information, please contact scholarworks@[Link].
A Practical Guide to Instrument Development and Score Validation in the Social
Sciences: The MEASURE Approach

Cover Page Footnote


Special thank you to my colleague and friend, Ryan Flinn, for his consultation and assistance with
copyediting.

This article is available in Practical Assessment, Research, and Evaluation: [Link]


vol26/iss1/1
Kalkbrenner: The MEASURE Approach

A peer-reviewed electronic journal.


Copyright is retained by the first or sole author, who grants right of first publication to Practical Assessment, Research & Evaluation. Permission
is granted to distribute this article for nonprofit, educational purposes if it is copied in its entirety and the journal is credited. PARE has the
right to authorize third party reproduction of this article in print, electronic and database forms.
Volume 26 Number 1, January 2021 ISSN 1531-7714

A Practical Guide to Instrument Development and Score


Validation in the Social Sciences: The MEASURE Approach
Michael T. Kalkbrenner, New Mexico State University

The research and practice of social scientists who work in a myriad of different specialty areas involve
developing and validating scores on instruments as well as evaluating the psychometric properties of
existing instrumentation for use with research participants. In this article, the author introduces The
MEASURE Approach to instrument development, an acronym of seven empirically supported steps
for instrument development, and initial score validation that he developed based on the
recommendations of leading psychometric researchers and based on his own extensive background
in instrument development. Implications for how The MEASURE Approach has utility for enhancing
the assessment literacy of social scientists who work in a variety of different specialty areas are
discussed.

Introduction are already financially burdened by the cost of required


textbooks. Similarly, applied social scientists (e.g.,
Assessment literacy is a pertinent issue in social counselors and teachers) who are not affiliated with a
sciences research, as researchers tend to assess latent university that provides access to electronic data bases
variables (e.g., personality, morale, and other attitudinal might also have limited access to these resources.
variables) that are abstract in nature, which are
generally appraised by inventories (Gregory, 2016). To The literature is lacking a single refereed journal
this end, social science researchers and practitioners article (a one-stop-shop) that includes a practical
are responsible for understanding the basic outline of the instrument development and validation
foundations, operations, and applications of testing, of scores process based on a number of synthesized
including instrument development (Standards for recommendations of prominent expert, contemporary
Educational and Psychological Testing, 2014). psychometric researchers. Such an article has potential
Instrumentation with strong psychometric support, to provide social scientists with a single and accessible
however, tends to be underutilized by social scientists resource for developing their own measures as well as
when conducting program evaluation and other types for evaluating the rigor of existing measures for use
of research (Tate et al., 2014). The extant literature with research participants. The primary aim of the
includes a series of peer-reviewed journal articles (e.g., present author was to introduce The MEASURE
Benson, 1998; Kane, 1992; Mvududu & Sink, 2013), as Approach to instrument development. MEASURE
well as textbooks and book chapters (e.g., Bandalos & (Figure 1) is an acronym comprised of the first letter of
Finney, 2019; DeVellis, 2016; Dimitrov, 2012; Fowler, the following seven empirically supported steps for
2014; Gregory, 2016; Kane & Bridgeman, 2017), which developing and validating scores on measures: (a)
collectively outline the instrument development Make the purpose and rationale clear, (b) Establish
process as well as guidelines for testing the validity of empirical framework, (c) Articulate theoretical
inferences from the scores. However purchasing these blueprint, (d) Synthesize content and scale
resources is infeasible for many graduate students, who development, (e) Use expert reviewers, (f) Recruit
participants, and (g) Evaluate validity and reliability.
Published by ScholarWorks@UMass Amherst, 2021 1
Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 2
Kalkbrenner, The MEASURE Approach

The MEASURE Approach was developed based on An instrument development study might also be
the guidelines of leading psychometricians, primarily necessary if a researcher determines that an existing
Benson (1998), DeVellis (2016), Dimitrov (2012), and instrument is inappropriate for use with their target
Mvududu and Sink (2013), as well as the author’s population (e.g., cross-cultural fairness issues). In some
extensive background and experience with instrument instances, developing an original measure with a
development and score validation. Finally, an exemplar diverse population can be more appropriate than
description of an instrument development study confirming scores on an established measure (step 7)
conducted by Kalkbrenner and Gormley (2020) is that was developed and normed with a different
presented to provide an example of applying each step population. Suppose for example, a researcher is
in The MEASURE Approach. seeking a screening tool for appraising mental health
distress among Spanish speaking clients. There might
Figure 1. The MEASURE Approach to Instrument
be utility in creating a new screening tool (based on the
Development
culture) rather than trying to validate scores on an
Make the purpose and rationale clear existing measure with Spanish speaking clients, as the
Establish empirical framework nature and breadth of the construct of measurement
(content validity, step 2) can vary substantially between
Articulate theoretical blueprint different cultures. Thus, even if an existing measure of
Synthesize content and scale development mental health distress is found to be statistically sound
with Spanish speaking clients, it might fail to capture
Use expert reviewers unique elements of mental health distress in the culture
Recruit participants (see Kane, 2010, for an overview of fairness-related
considerations in testing and assessment). After
Evaluate validity and reliability making the purpose clear, a researcher should provide
a rationale to justify why creating a new instrument is
necessary.
Step 1: Make the Purpose and
When providing a rationale for developing a new
Rationale Clear instrument, researchers should (a) present a summary
Researchers should first define the purpose of of their review of the extant measurement literature,
conducting an instrument development study by telling (b) cite any similar instruments that already exist, and
the reader what they are seeking to measure and why (c) articulate the construct(s) that existing measures fail
measuring the proposed construct is important to capture in order to highlight a gap in the existing
(DeVellis, 2016; Dimitrov, 2012). As part of this step, measurement literature (DeVellis, 2016; Dimitrov,
researchers should review the existing literature on the 2012). Finally, test developers should discuss how their
proposed construct of measurement to determine if proposed instrument has potential to fill the
they can use/adapt an existing measure or if an aforementioned gap in the measurement literature and
instrument development study is necessary (Mvududu articulate how filling this gap has significant potential
& Sink, 2013). If a measure exists in the literature, to advance future research and practice (see Fu and
researchers should carefully evaluate the rigor of the Zhang, 2019, as well as Kalkbrenner and Gormley,
instrument development study by comparing the 2020, for examples of providing a rationale for
procedures that the test developers employed to instrument development based on these steps).
established empirical standards (e.g., The MEASURE
Approach). An instrument development study is
necessary if the literature is lacking a measure to
appraise the researcher’s desired construct of
measurement. An instrument development study
might also be necessary if a researcher determines that
an existing instrument is potentially psychometrically
flawed (e.g., lacking reliability or validity evidence, step
7).
[Link]
DOI: [Link] 2
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 3
Kalkbrenner, The MEASURE Approach

Step 2: Establish Empirical sources (e.g., peer-reviewed journal articles) that


collectively provide a rationale for the intended
Framework construct of measurement. At this stage of
Benson (1998) suggested instrument development development, the empirical framework can be general
undergoes a Substantive Stage in which test developers in nature. The idea is to refer to at least one empirical
situate the study within the context of a theoretical source that will set the framework for developing items
framework. Similarly, in the Establish Empirical that capture the proposed construct of measurement.
Framework stage, researchers are tasked with
identifying a theory(ies) and/or synthesized findings
from the extant literature to set an empirical Step 3: Articulate Theoretical
framework for the item development process. In this Blueprint
context, an empirical framework refers to at least one
theory or scholarly source (e.g., peer-reviewed) that Researchers can begin to refine and organize their
provides a series of principles or assumptions that empirical framework by creating a theoretical
underlie the proposed construct of measurement. For blueprint. A theoretical blueprint (Figure 2) is a tool for
example, a test developer might refer to Maslow’s enhancing the content validity of a measure by offering
Hierarchy of Needs (Maslow, 1943) as the empirical researchers two primary advantages, including (a)
framework for developing a measure to appraise the creating the content and domain areas for the construct
extent to which one’s various needs are satisfied. The of measurement and (b) determining the approximate
goal in step 2 is to provide an overview of the proportion of items that should be developed across
theoretical underpinnings for the proposed construct each content and domain area (Menold et al., 2015;
of measurement, which is an important step for Summers & Summers, 1992). Content areas in a
ensuring content validity or the extent to which test blueprint refer to the specific subject aspects for the
items adequately represent the scope of a construct of construct of measurement. Domain areas in a blueprint
measurement (Lambie et al., 2017). Four primary refer to the various application-based dimensions of
methods for demonstrating content validity in social the construct of measurement. The content and
sciences research include (a) empirical framework, (b) domain areas on a blueprint should be derivatives of
theoretical blue print (step 3), (c) expert review (step the extant literature and, in most cases, multiple
5), and (d) pilot testing (step 6). plausible/logical content and domain areas can be
generated for a construct of measurement; thus, there
In some instances, the literature might be lacking is usually not one “right” or “correct” content or
an established theory that a researcher can use to set an domain area for any given measure. Researchers are
empirical framework for the item development tasked with providing a rationale from the extant
process. In these instances, a researcher can build their
literature to justify the utility of their content and/or
own theoretical framework for item development by
domain areas.
synthesizing the findings from a number of empirical

Figure 2. Example Theoretical Blueprint: Mental and Physical Health (Kalkbrenner & Gormley, 2020)
Domain Areas
Frequency Intensity Duration
Diet 7 6 6
Exercise 7 6 6
Content Areas
Stress Management 12 9 9
Avoiding Toxins 4 4 4

Published by ScholarWorks@UMass Amherst, 2021 3


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 4
Kalkbrenner, The MEASURE Approach

Researchers should refer to the extant literature testing (step 7). Items should be brief, clear, and
(step 1) and the empirical framework (step 2) to written at approximately a sixth-grade reading level
determine the breadth of their proposed construct of (see DeVellis, 2016, for a comprehensive overview of
measurement. Researchers can also seek assistance strategies for developing sound items).
from content experts (step 5) to help with item
Using a Research Team in the Item Development Process
development and determining the breadth of their
proposed construct of measurement. Seeking input The initial process of creating items that are
from a panel of content experts who are representative intended to measure a latent construct is qualitative in
of the field of study might be especially helpful in cases nature, thus, there is utility in incorporating tenants of
where there is a gap in literature on the proposed triangulation of multiple researchers (i.e., a research
construct of measurement. To enhance content team, see Carter et al., 2014) from qualitative inquiry
validity, test developers should adjust for the relative into the item development process. Researchers should
importance of the items across each content and first individually create a pool of items based on the
domain area for the construct of measurement. In empirical framework (step 2) and blueprint (step 3).
other words, more items should be developed for the Researchers should seek to develop an exhaustive list
intersecting content and domain areas that represent a of items (i.e., as many as possible) during the first
greater scope of the construct of measurement. The round of item development. The researcher can then
purpose of numbering the intersecting cells on a edit/reduce their list by looking for redundancy. Once
blueprint (Figure 2) is to denote the approximate each research team member has created their own list
proportion of items that will be developed to represent of potential items, they can come together for a series
each cell. Not every instrument, however, will be based of meetings in which they review and discuss each team
on a theoretical framework that lends itself to the member’s list of items and eventually come to a
blueprint matrix that is depicted in Figure 2. As such, consensus about the initial pool of items that will be
a test developer might include only content area(s) or sent to the expert reviewers (step 5). Conducting a
only domain area(s) on their blueprint. Blueprint qualitative pilot study with the targeted population is
construction is a flexible procedure, which allows another way that researchers can enhance the rigor in
researchers to customize this tool to enhance content the item development process. Specifically, researchers
validity in the subsequent item development process. can conduct individual interviews and/or focus groups
with participants that meet the inclusion criteria of the
target population. Emergent codes and themes from
Step 4: Synthesize Content and Scale the qualitative interviews might have utility for guiding
the item development process.
Development
Assembling the Instrument
Synthesize Content
Self-administered questionnaires should be
Before developing an initial pool of items, transparent and relatively easy to follow (Fowler,
researchers should be clear about the parameters of 2014). Researchers are encouraged to implement a
their proposed construct of measurement and reflect standard convention for each element in the measure;
on how their construct differs from other latent for example, using all uppercase letters for the
variables in order to avoid redundancy (DeVellis, 2016; instructions, italicized text for scale points (see Scaling
Fowler, 2014). The empirical framework (step 2) and section below), and regular text for test items.
blueprint (step 3) can be instrumental tools to Different font styles (e.g., Times New Roman, Calibri)
synthesize content for the purpose of refining the can also be utilized to clearly denote different elements
parameters of the construct of measurement during the of the survey. Instructions should be as short as
item development process. Researchers should possible and include visual cues (e.g., arrows, bolded
develop approximately three to four times as many text) when appropriate, as participants tend to read
items that will comprise the final version of the instructions briefly, if at all (Fowler, 2014). Moreover,
measure (DeVellis, 2016) as multiple (potentially definitions should be presented for any vague or
problematic) items are usually deleted during the abstract terms. For example, The Revised Fit, Stigma,
expert review (step 5) and during reliability/validity and Value (FSV) Scale, a screening tool for measuring
[Link]
DOI: [Link] 4
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 5
Kalkbrenner, The MEASURE Approach

barriers to seeking personal mental health counseling has usefulness for appraising cumulative/hierarchical
services, provides respondents with a definition of constructs in which test takers who endorse a strong
counseling from the American Counseling Association statement also endorse milder statements by default
(Kalkbrenner et al., 2019). The response options or (DeVellis, 2016). For example, asking respondents to
scale points (see Scaling section below) on an endorse (select agree or disagree) with each of the
instrument tend to appear above the items and can be following statements, I feel happy occasionally, I feel happy
repeated after every 10 to 15 questions depending on most of the time, and I feel happy all of the time. Someone
the length of the measure. who selects agree for I feel happy all of the time will almost
certainly also select agree for I feel happy most of the time.
Test questions should be as brief and concise as
Additionally, a binary scale in which respondents are
possible and do not necessarily have to be complete
asked to select one of two possible options has
sentences. Item stems can have utility for increasing
particular utility for appraising observed variables with
brevity and decreasing respondent fatigue. On The
dichotomous response options. For example, asking
Revised FSV Scale, for example, participants are asked
respondents to indicate (e.g., yes or no) if they have a
to reply to the following stem: “I am less likely to
high school diploma. Moreover, checklists allow test
attend counseling because....” to a number of items
takers to select multiple response options and are
(e.g., item “... it would suggest I am unstable,”
useful when more than one answer might apply to a
Kalkbrenner et al., 2019, p. 26). Researchers can refer
survey item. For example, a researcher might provide
to the theoretical blueprint (step 3) as an aid for
participants with a list of every state in the U.S. and ask
ordering the items. When ordering the items, the
them to select all of the states that they have visited.
subject area clusters (intersecting content and domain
areas on the blue print) should be interspersed to A semantic differential scale allows one to capture
reduce the likelihood of a response set. Instruments are the connotative meaning of stimuli or objects by
typically revised and sometimes reassembled including unipolar or bipolar adjectives as scale points
throughout steps 5 to 7 as items on the test are usually (DeVellis, 2016). Similarly, on a visual analogue scale
revised/removed following expert review (step 5) respondents are asked to place a mark at a specific
and/or during validity testing (step 7). point on a line between scale points that represent the
opposite ends of a continuum. See DeVellis (2016, pp.
Scaling
129-130) for examples of semantic differential and
Researchers should work together to determine the visual analogue scales. Finally, a Rasch scale is based on
format of measurement or scale for their instrument. item response theory (Amarnani, 2009) and is centered
Likert scaling is one of the most commonly used on the notion that test takers are more likely to respond
scaling formats in the social sciences (DeVellis, 2016). correctly to items that measure easier degrees of a trait
When creating a Likert scale, items are presented in (Boone et al., 2017). Test takers are provided with a
declarative statements with anchor definitions (i.e., more difficult or easier subsequent item based on
response options) that designate fluctuating amounts whether they answered the previous question correctly.
of agreement or approval of the statement, for Rasch scales have utility for high-stakes testing (e.g.,
example, 1= strongly disagree, 2= disagree, 3= neutral, 4= intelligence tests, tests of cognitive ability). Reviewing
agree, 5= strongly agree. It is important to label each the intricacies of item response theory and Rasch
anchor definition on the scale. The number and format scaling are beyond the scope of this manuscript,
of anchor definitions should be determined by the however, refer to Boone et al. (2017) for a primer on
construct of measurement (DeVellis, 2016; also see Rasch analysis and scaling. Ultimately, researchers
Vagias, 2006, for a variety of Likert scale response should choose their scaling option based on the nature
anchors). Likert scaling is particularly appropriate for of their construct of measurement (e.g., Likert scaling
measuring attitudinal constructs (e.g., personality, for attitudinal measures, Rasch scaling for high stakes
beliefs, values, or emotions). testing). See DeVellis (2016) for a detailed overview of
Despite the popularity of Likert scales in social selecting a scaling option that is consistent with one’s
science research, a myriad of additional scaling construct of measurement.
methods are available. Guttman scaling, for example,

Published by ScholarWorks@UMass Amherst, 2021 5


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 6
Kalkbrenner, The MEASURE Approach

Step 5: Use Expert Reviewers asked to rate on a Likert scale (step 4) the extent to
which each survey item represents a content area of the
Once the raw version of the instrument (initial proposed construct of measurement. Researchers also
pool of items and scaling format) is assembled, the tend to attach an open-response option to Likert scale
measure should be sent to a group of external expert questions so that reviewers can discuss the reasons
reviewers who are knowledgeable in the content area behind their ratings. Ikart (2019) provides a
(Ikart, 2019; Lambie et al., 2017). Experts are comprehensive overview of using expert reviewers in
sometimes consulted for assistance with item the instrument development process, including but not
development (step 2), however, different expert limited to creating these forms.
reviewers (i.e., people who did not contribute to
developing the original pool of items) should be
included in this phase to provide a fresh/non-biased Step 6: Recruit Participants
perspective. The number of expert reviewers tends to
range between three and five, however, upwards of 20 Pilot Testing
expert reviewers have been noted in the literature Before collecting data from human subjects,
(Ikart, 2019). The primary purpose of the expert review researchers should review and obtain proper
process is to maximize the measure’s content validity institutional review board (IRB) approval. Pilot testing
by obtaining feedback from a panel of experts (also referred to as preliminary testing) involves
regarding “how relevant they think each item is to what administering the instrument to a small developmental
you intend to measure” (DeVellis, 2016, p. 135). sample that is similar to the target population. Pilot
Test developers are responsible for justifying what testing allows researchers to test their procedures and
constitutes an “expert” in a given content area. Expert check for errors in data imputation (e.g., a survey
reviewers (approximately 10+ years’ experience) are question that asks for a written response, however, the
typically classified into two possible groups for question format is set to only allow a single numeric
ensuring the rigor and content validity of items, entry) or technology errors (e.g., particular web
including (a) survey and questionnaire experts, and/or browsers that do not support the survey platform or
(b) substantive or subject matter experts (Ikart, 2019). issues with broken or inconsistent hyperlinks). Pilot
Reviewers with survey and questionnaire expertise are testing also provides an opportunity to solicit feedback
well versed in best practices and mechanics of from participants about the content and readability of
questionnaire design and item development. Subject the items. There are a number of guidelines for what
matter experts have a wealth of knowledge/experience constitutes a small pilot sample, however, pilot samples
with the construct of measurement and ensure that the tend to range between 25 and 150 participants
collective pool of items sufficiently captures the (Browne, 1995; Hertzog, 2008). Pilot study data should
extensiveness of the construct. be reviewed for information about item content,
including clarity and readability as well as for any errors
Expert reviewers can be solicited via email list in the administration procedures. Researchers can
serves associated with professional organizations. Test
tentatively compute initial item analyses, for example,
developers can also use their personal contacts (e.g., inter-item correlations and descriptive statistics. Ideally
current/former professors, employers, co-workers) for the pilot sample is 100+ for computing initial item
suggestions about potential expert reviewers. Expert analyses (Field, 2018), however, researchers can
reviewers can be hired (depending on funding compute these analyses with smaller samples as long as
accessibility). Expert reviewers are sometimes added as they consider the limitations of a small sample size
co-authors of the manuscript if their input significantly when interpreting the results. Researchers might
influences the measure. In most cases, there is utility in conduct a factor analysis (step 7) with the pilot data as
giving the expert reviewers an opportunity to make long as their sample size is sufficient (next section). If
direct comments on the instrument itself (i.e., track pilot study participants highlight issues related to item
changes in MS Word) as well as soliciting their content and readability, researchers should revise and
feedback on a brief survey or form to solicit additional repeat the pilot process.
feedback. For example, expert reviewers might be
[Link]
DOI: [Link] 6
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 7
Kalkbrenner, The MEASURE Approach

Sample Size for the Main Study are > 0.70 (Knekta et al., 2019). If, however,
communalities are < 0.50, a sample size of 300+ would
Researchers should determine their minimum
be required to obtain accurate estimates. Moreover, as
sample size for the main study prior (a priori) to
the number of factors (subscales) increases the sample
launching data collection (Mvududu & Sink, 2013).
size must also increase. For example, a model with
Factor analysis (step 7) is one of the most common
seven or more factors would require a sample size of
statistical tests for validating scores on newly
developed measures (Bandalos & Finney, 2019; 500+.
Benson, 1998; Mvududu & Sink, 2013). In general, Based on the synthesized recommendations of the
larger samples are desirable for factor analysis due to leading psychometric researchers cited in this section,
increases in statistical power, however, there is not a this writer recommends that test developers determine
clear consensus in the literature for determining the their a priori minimum sample size by following one of
minimal sample size for factor analysis (Knekta et al., the two following criteria, whichever yields a larger
2019). Originally, sample size guidelines for factor sample: (a) an STV ratio of 10:1 or (b) a sample size of
analysis were based on general benchmarks. For 200 participants. Suppose, for example, a measure is
example, Comrey and Lee (1992) offered the following comprised of 50 items. The minimum sample based on
guidelines for sample size in psychometric research: 50 an STV ratio of 10:1 would be 10*50 or 500. Before
= very poor, 100 = poor, 200 = fair, 300 = good, 500 = very the cessation of data collection, however, researchers
good, and > 1,000 = excellent. In more recent years, should check their sample size with the guidelines
many psychometric researchers determine their provided by Bandalos and Finney (2019) and Knekta
minimum a priori sample size by calculating the ratio et al. (2019, see the previous paragraph) as the unique
between the number of participants and the number of properties of the data (e.g., communalities, number of
estimated parameters or variables being analyzed, items/factors) should be considered when making final
sometimes referred to as the subjects-to-variables ratio decisions about when one has achieved a sufficient
(STV, Beavers et al., 2013; Mvududu & Sink, 2013). sample size.
The recommended size of this ratio varies substantially
Obtaining a Sufficient Sample Size: Accessing Participants
between different psychometricians, from as low as 3:1
to as high as 20:1 (Mvududu & Sink, 2013), however, There are a variety of strategies for recruiting
10:1 is typically considered acceptable. However, this participants to obtain a sufficient sample size for
ratio might be insufficient for estimating the minimum psychometric analyses (Sharon, 2018). Convenience
necessary sample size for brief measures sampling in public locations (with the proper
(approximately 19 or less items) as the sample size for approvals) can be a cost-effective strategy for accessing
psychometric studies should include at least 200 participants. When conducting survey research with
participants (Comrey & Lee, 1992). college students, for example, a researcher might
recruit prospective participants as they enter the library
A number of contemporary psychometricians (e.g., or student union. Researchers can also consider using
Bandalos & Finney, 2019; Knekta et al., 2019) reject a their personal contacts (e.g., current/former
one size fits all approach for determining sample size.
professors, employers, co-workers) to distribute
Sample size in psychometric research varies as a recruitment messages for participation in research. For
function of communality: “amount of variance in the example, a researcher might ask one of their
variables that is accounted for by the factor solution, current/former professors to send a recruitment email
the number of variables per factor, and the interactions to all of the students in their department. Researchers
of these two conditions” (Bandalos & Finney, 2019, p. who are affiliated with an organization (e.g., university)
102). Generally, more simplistic models (i.e., fewer might have institutional support available to aid in data
items and factors/subscales) require smaller samples; collection for IRB approved research. As just one
Wolf et al. (2013) demonstrated that a sample size of example, many universities make their entire student
30 was sufficient for confirming a unidimensional registry (i.e., email addresses of all enrolled students)
factor solution with factor loadings > 0.80. Similarly, a publicly available. Researchers also sometimes offer
sample size as low as 100 can be sufficient for factor small incentives (e.g., small electronic gift cards, bag of
analysis with three factors and item communalities that candy) to all participants or give participants the option

Published by ScholarWorks@UMass Amherst, 2021 7


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 8
Kalkbrenner, The MEASURE Approach

of entering a raffle to win a prize. When offering survey platforms offer user-friendly item construction
incentives for participation in survey research, options (e.g., matrices for building Likert scales, slider
however, there exist a number of ethical considerations options for visual analogue scaling, written response,
(Singer & Bossarte, 2006). For example, incentives multiple choice, and more). Most electronic survey
cannot exert undue influence or be coercive, including platforms generate anonymous electronic links, which
but not limited to offering excessive monetary can be sent to prospective participants via mass email
compensation. What constitutes an excessive or distribution or posted on websites. In addition, the
inappropriate incentive varies by context, thus majority of these platforms also allow users to upload
researchers should work with their research teams, a contact list of prospective participants and use a
institutional review board, and consult the extant piped text option to personalize each individual
literature to determine an appropriate incentive for a message. Suppose for example, a researcher has a
particular study. See Singer and Bossarte (2006) for a registry spreadsheet of 20,000 prospective participants
detailed overview of practical and ethical with their information organized into columns (e.g.,
considerations when offering incentives in research. first name, last name, email address…). They can
personalize the greeting in each message by using a
While convenience sampling methods tend to be
piped text option, which will automatically insert each
cost effective, its use comes with a cost to the
participant’s name in the greeting field (e.g., Dear
representativeness of the sample. Data collected via
${m://FirstName}). Electronic survey platforms also
convenience sampling, for example, tends to represent
eliminate the need for raw data entry as data are
scores from participants who have opportune and a
downloaded directly into SPSS or Excel data
proclivity to participate in survey research (i.e., people
who like to take surveys). To this end, more rigorous spreadsheets.
sampling techniques (e.g., random sampling) tend to
enhance the generalizability of results. Alvi (2016)
offers a free and comprehensive manual on various Step 7: Evaluate Validity and
sampling techniques in social science research. Reliability
Depending on funding accessibility, researchers can
also hire data collection contracting companies (e.g., The final step in initially validating scores on a new
IMPAQ, 2020; Qualtrics Sample Services, 2020) for measure involves testing for validity (the scale is
data collection. Qualtrics Sample Services (2020), for measuring what it is intended to measure) and
example, is a data collection contracting company with reliability (consistency of scores) evidence of the
a national sampling pool of over 96 million participants measure and its subscales (Gregory, 2016). In a
and they can recruit random samples, stratified by landmark article, Kane (1992) introduced an argument-
variables of interest (i.e., adults in U.S. stratified by the based approach to validity based on the notion that
most recent census data). Qualtrics Sample Services making an interpretive argument is “the framework for
can also access specific samples, for example, Latinx collecting and presenting validity evidence and seeks to
females in a certain age range, first-generation college provide convincing evidence for its inferences and
students, high school students, and a number of other assumptions” (p. 527). According to Kane
specific populations. Similarly, Amazon Mechanical interpretative arguments can never be proven with
Turk (2020) is crowdsourcing marketplace where absolute certainty. To this end, test developers are
researchers can recruit prospective participates and tasked with presenting multiple forms of evidence to
offer them monetary compensation to incentivize their demonstrate the plausibility of their interpretative
voluntary participation. Finally, there are a number of argument for validity evidence (Kane, 1992; Kane &
companies (e.g., Redi Data, 2020) that sell randomly Bridgeman, 2017). Validity is a unitary construct,
generated email lists of a target population. however, there exist a number of sources of validity
evidence, including content validity (steps 3 to 5),
Electronic Survey Research. Online survey criterion-related validity, and construct validity (Kane
platforms, for example, Qualtrics (2020), REDCap & Bridgeman, 2017; Lenz & Wester, 2017).
(2020), eSurveysPro (2020), and SurveyMonkey (2020),
are becoming increasingly popular. These electronic
[Link]
DOI: [Link] 8
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 9
Kalkbrenner, The MEASURE Approach

Criterion-Related Validity types of factor analysis: exploratory factor analysis


(EFA) and confirmatory factor analysis (CFA;
Demonstrating criterion-related evidence involves
examining associations between test scores and a non- Bandalos & Finney, 2019; Mvududu & Sink, 2013).
test criterion (Kane, 1992). Criterion-related validity Exploratory Factor Analysis. The primary
evidence includes concurrent validity or the extent to purpose of EFA is to uncover the underlying
which test scores relate to a non-test criterion in the dimensionality within groups of test items by detecting
present. For example, a test developer who compares how the items cluster together into subscales
high school students’ scores on an anti-bullying (subscales are also known as dimensions or factors),
questionnaire to their teachers’ ratings of bullying in each of which constitute an aspect of the larger
the classroom is testing criterion-related validity. A construct that the researcher is seeking to measure
high association between the teacher’s ratings and anti- (Beavers et al., 2013). The EFA is exploratory in nature
bullying scores would yield concurrent validity and the analysis will isolate latent factors that explain
evidence for the test, as scores on the measure are the covariance (correlations) among a group of items
consistent (concur) with a non-test criterion reference (Mvududu & Sink, 2013). Prior to computing EFA, the
(the teacher’s rating). Criterion-related validity following three preliminary tests should be conducted
evidence can also include predictive validity or the to test the factorability of the data (i.e., determine if the
degree to which test scores predict a non-test criterion data set is appropriate for factor analysis): inter-item
in the future or past. For example, a test developer correlation matrix, Bartlett’s test of Sphericity, and the
might evaluate the predictive validity of a career Kaiser–Meyer–Olkin (KMO) Test for Sampling
readiness instrument by testing the extent to which Adequacy (Beavers et al., 2013; Mvududu & Sink,
readiness scores predict respondents’ future 2013). There are a number of additional important
employers’ ratings of their job performance. Criterion- considerations in EFA, including factor extraction,
related evidence has utility for supporting one’s factor rotation, factor retention (Beavers et al., 2013,
interpretative validity argument, however, pp. 4-11), and naming the rotated factors (Mvududu &
demonstrating construct validity evidence is widely Sink, 2013, pp. 90 - 91).
considered a cornerstone of validating scores on newly Confirmatory Factor Analysis. CFA is a “theory
developed tests (Bandalos & Finney, 2019; Benson, testing strategy” based on structural equation modeling
1998). for determining the extent to which the factor solution
Construct Validity of an existing measure maintains internal structure with
a new sample of participants (Mvududu & Sink, 2013,
Construct validity refers to the extent to which an
p. 91). In an instrument development study (or when
instrument accurately appraises a theoretical or
testing the psychometric properties of an established
hypothetical construct and is the most rigorous form
measure with a new population), researchers should
of validity evidence for validating scores on newly
collect data from a new sample and compute a CFA to
developed tests (Benson, 1998; Kane & Bridgeman,
test the fit between the dimensionality of the
2017). Specifically, tests of internal structure and relations
hypothesized factor solution with a new sample
with other established theoretical constructs are two of the
(Bandalos & Finney, 2019). Model fit is determined by
most extensively used methods for demonstrating
investigating a combination of goodness-of-fit indices
construct validity in social science research (Gregory,
such as: incremental, absolute, and parsimonious
2016; Kane & Bridgeman, 2017; Swank & Mullen,
(Bandalos & Finney, 2019, p. 115). Determining model
2017).
fit is a complex task and “it is naïve to believe that
Internal Structure and Factor Analysis model fit can be properly assessed by a single index”
Factor analysis, a series of psychometric analysis (Bandalos & Finney, 2019, p. 115). Psychometric
for testing the dimensionality (internal structure) of the researchers offer general cutoff values for particular fix
construct of measurement, is probably the most widely indexes (Hu & Bentler, 1999; Schreiber et al., 2006);
used procedure for testing construct validity in social however, these values should be used as general
sciences research (Bandalos & Finney, 2019; Benson, guidelines rather than absolute standards. When
1998; Mvududu & Sink, 2013). There are two primary evaluating model fit, researchers should assess fit

Published by ScholarWorks@UMass Amherst, 2021 9


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 10
Kalkbrenner, The MEASURE Approach

holistically by considering the implications of multiple literature scores suggest that depression and anxiety are
fit indexes. In addition to evaluating fit indexes, separate theoretical constructs (i.e., separate constructs
researchers should also consider correlation residuals, in the same domain).
parameter estimates, and convergence problems (see Employing a Multi-Faceted Approach to Construct Validation
Bandalos & Finney, 2019, p. 115) when evaluating
model fit. On one level, evaluating construct validity by
testing a new measure’s relation with other
Relations with Other Established Theoretical Constructs theoretically-related measures presents a potential
Examining the relationship between scores on temporal-validity issue, as one is using an old test to
newly developed tests with established theoretical validate scores on a new test (Gregory, 2016). Factor
constructs is also a popular method of demonstrating analysis yields information about the internal
construct validity in social sciences research (Benson, dimensionality of instrumentation, however, it does
1998; Strauss & Smith, 2009; Swank & Mullen, 2017). not yield evidence about precisely what is being
In fact, Benson (1998) refers to testing the relationship measured (Benson, 1998). To this end, correlating
between scores on a new test with other theoretically- scores on a new test with an established test has greater
related measures as “the most crucial” stage in utility for isolating the precise construct of
conducting a strong program of construct validation measurement. It is important to note that no test is
(p. 14). One approach is to test convergent validity or inherently valid (i.e., one can only validate scores on a
“the relationship among different measures of the test rather than validate the test itself). Thus, tests are
same construct” (Strauss & Smith, 2009, p. 1). For only valid for certain purposes, with particular
example, the developer of a new Depression Severity populations, at specific points of time. Psychometric
inventory might test the correlation between scores on support for a test is strongest when researchers
their new measure with scores on an established conduct a series of psychometric studies in which they
screening tool for depression (e.g., the Beck demonstrate different forms of validity evidence for
Depression Inventory). Higher correlations (e.g., r > scores on the test among various populations. To this
0.5, see Swank & Mullen, 2017, p. 272) would provide end, there is utility in employing a multi-faceted
stronger convergent validity evidence, as scores on the method of construct validation. For example,
new screening tool are similar (converge) with scores researchers can employ factor analysis to uncover the
on an established measure for appraising the intended dimensionality (internal structure) of an instrument as
construct of measurement (e.g., Depression Severity). well as testing the convergence/divergence of the
measure with other well-established tests. Moreover,
Assessing discriminant validity (also known as
initial validity testing can reveal insights for improving
divergent validity) is another method of establishing
the construct and content validity of instrumentation.
construct validity by demonstrating “that a measure of
In such instances, test developers can make revisions
a construct is unrelated to indicators of theoretically
to the items and repeat steps 5 to 7. The decision about
irrelevant constructs in the same domain” (Strauss &
whether to revise and retest items should be made via
Smith, 2009, p. 1). Referring to the example in the
research team consensus, which can include
previous paragraph, the test developer might correlate
scores on their new Depression Severity index with an consultation with content experts (step 5).
established measure of Anxiety Severity (e.g., The Beck Reliability Evidence
Anxiety Inventory [BAI]). Based on the extant Once a researcher has established validity evidence
literature (e.g., Nguyen et al., 2019), one should expect for scores on their instrument, they should compute a
only a minimal-to-moderate relationship between test of the measure’s reliability or consistency of scores.
symptoms of anxiety and depression (i.e., divergence There are numerous forms of reliability evidence; test-
between theoretically different contracts in the same retest, alternative forms, inter-rater, and internal
domain). Thus, a minimal-to-moderate correlation consistency, (see Bardhoshi & Erford [2017] for a
(e.g., r < 0.4, see Swank & Mullen, 2017, p. 272) detailed overview of each form of reliability evidence).
between scores on the new Depression Severity index This author will focus on internal consistency reliability
and the BAI would support the new scale’s in this manuscript since psychometric researchers tend
discriminant validity as consistent with the extant
[Link]
DOI: [Link] 10
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 11
Kalkbrenner, The MEASURE Approach

to employ cross-sectional research designs, in which relation to the total composite score. McDonald's
data are collected at only one specific point in time. Omega is advantageous when tau-equivalence is not
met as it allows the associations between each item and
Cronbach's Coefficient Alpha
the total scale to vary. Nájera Catalán (2018) provide a
Cronbach's coefficient alpha (α) is widely cited series of recommendations for interpreting ω.
(Bardhoshi & Erford, 2017; Cho, 2016; Dunn et al., McDonald's Omega is the most popular alternative to
2014; McNeish, 2018) as the most commonly used Cronbach's coefficient alpha (Dunn et al., 2014;
measure of internal consistency reliability across the McNeish, 2018), however, other options exist
social sciences and represents the mean value of all including the greatest lower bound (GLB) method and
possible split-half combinations of the items on a Coefficient H. Discussing the intricacies of these
measure or subscale (Cronbach, 1951). Cronbach's estimates is beyond the scope of this manuscript,
coefficient alpha ranges from 0 to 1, with values closer however, see Bendermacher (2017) for more on the
to 1 denoting stronger reliability evidence. There is GLB method and McNeish (2018) for more on
much debate in the literature regarding the lowest Coefficient H. The overall take-away message is that
acceptable cutoff value for α. George and Mallery there is no single, supreme reliability estimate for all
(2003) offer the following guidelines: “ α > .9 – tests, as the derivative of each estimate is based on
Excellent, α > .8 – Good, α > .7 – Acceptable” (p. 231). different measurement models. Thus, researchers are
However, the threshold for an “acceptable” coefficient tasked with carefully selecting and explaining the most
alpha value should depend on the construct of appropriate reliability estimate for their particular study
measurement (Taber, 2018) and the stakes or (Cho, 2016; Dunn et al., 2014; McNeish, 2018).
consequences for test takers that are attached to the
test. For example, reliability evidence should be
stronger for high-stakes testing (e.g., tests of The Measure APPROACH: An
intelligence or college readiness tests) than for Example
attitudinal screening tools (e.g., interest inventories or
non-diagnostic personality tests). Thus, it is the test Kalkbrenner and Gormley (2020) employed the
developer’s responsibility to provide a rationale for steps in The MEASURE Approach to develop the
acceptable internal consistency reliability estimates Lifestyle Practices and Health Consciousness
based on the nature of the test. Inventory (LPHCI). Kalkbrenner and Gormley made
their purpose and rationale clear (step 1) by (a)
Alternatives to Cronbach's Coefficient Alpha describing their intention to create a measure for
Despite the popularity of Cronbach's coefficient appraising lifestyle practices of holistic wellness or
alpha in social sciences research, its use is sometimes integrated dimensions of physical and mental health,
called into question (e.g., McNeish, 2018; Taber, 2018). (b) exposing a gap in measurement literature for
Specifically, coefficient alpha is notoriously misused in appraising integrated aspects of mental and physical
instances when the data do not meet certain key health with a single, relatively brief composite scale,
assumptions (Dunn et al., 2014) including, the and (c) highlighting the need for such a measure in the
assumption of tau-equivalence or the notion that each integrated primary health care climate in the U.S.
scale item equally contributes to the total composite Specifically, the LPHCI had great potential for
scale score. This is problematic since Cronbach's measuring a new latent variable (Global Wellness) for
coefficient alpha tends to underestimate (sometimes enhancing the future research and practice of
substantially) the internal consistency reliability practitioners, especially those who work in integrated
estimate of scores on a scale in the absence of tau- behavioral health settings.
equivalence (McNeish, 2018). The empirical framework for the LPHCI (step 2)
Composite reliability estimates (e.g., McDonald's was developed based on two well-established
Omega [ω]) are a viable alternative to Cronbach's theoretical models of healthy lifestyle practices for
coefficient alpha as both estimates produce an internal preventing disease and optimizing physical and mental
consistency reliability coefficient based on the ratio health, including Servan-Schreiber’s life-style practices-
between the variance accounted for by each item in based anti-cancer method (diet, exercise, stress

Published by ScholarWorks@UMass Amherst, 2021 11


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 12
Kalkbrenner, The MEASURE Approach

management, and avoiding toxins) and Chopra and (DeVellis, 2016). For example, LPHCI item 14,
Fisher’s Big Five (mixed nuts, coffee, exercise, vitamin “skipped a meal despite feeling hungry” is brief, clear,
D, and meditation). According to Chopra, Fisher, and and written at a fifth-grade Flesh-Kincaid level.
Servan-Schreiber lifestyle practices that are only Kalkbrenner and Gormley (2020) utilized a research
implemented in a single facet of one’s life are seldom team during the item development process. Each
sufficient for preventing disease; rather, one’s research team member individually developed separate
engagement in a number of integrated lifestyle lists of possible LPHCI items based on the empirical
practices geared towards enhancing both physical and framework and blueprint. The team then engaged in a
mental health are essential for promoting their optimal series of meetings until a consensus was reached about
health and wellness (Chopra & Fisher, 2016; Servan- the items that became the initial pool of LPHCI items
Schreiber, 2009). The theoretical premise of both (Kalkbrenner & Gormley 2020).
Servan-Schreiber and Chopra and Fisher’s models of
The initial pool of LPHCI items were sent to three
holistic wellness was consistent with Kalkbrenner and expert reviewers (step 5). Collectively, the reviewers
Gormley (2020)’s aim to develop a screening tool of had over 65 years’ experience working in medical,
mental and physical wellness, thus they used these academic, and clinical mental health settings. Two of
models to set the major theoretical framework for the reviewers were subject matter experts and one was
developing a theoretical blueprint (Figure 2) and the a survey and questionnaire expert. Kalkbrenner and
initial pool of LPHCI items. Gormley (2020) then pilot tested the LPHCI with a
Kalkbrenner and Gormley (2020) created a small sample (N = 125) of the target population; no
theoretical blueprint (step 3) for the LPHCI and the technology issues emerged and participants did not
content areas (diet, exercise, stress management, and suggest any revisions to the items, thus researchers
avoiding toxins) were comprised of the four major proceeded to launch data collection for the main study.
tenants of Servan-Schreiber’s model of mental and Sample size for the main study was based on an STV
physical wellness (Figure 2). The domain areas on the ratio of 10:1. Specifically, there were a total of 42
LPHCI blueprint (Figure 2) included frequency, LPHCI items to enter into the EFA, thus, the
intensity, and duration, which Kalkbrenner and minimum sample based on a STV ratio of 10:1 was
Gormley (2020) adapted from the application-based 10*42 or 420. Participants were recruited via a data
dimensions of Servan-Schreiber’s as well as Chopra collection service (Qualtrics Sample Services, 2020) to
and Fisher’s models of holistic wellness. On the obtain a random national sample (stratified by the U.S.
LPHCI blueprint (Figure 2), for example, the diet, census data) of adults living in the U.S.
exercise, and avoiding toxins content areas of Servan- Upon the completion of data collection,
Schreiber’s model are all related to physical health. Kalkbrenner and Gormley (2020) evaluated initial
Thus, Kalkbrenner and Gormley (2020) included more reliability and validity evidence (step 7) for scores on
items in the stress management cells of the blueprint the LPHCI by conducting EFA, CFA, higher-order
to adjust for the relative importance of holistic wellness CFA, and tests of internal consistency reliability. The
(i.e., create a more equal focus on their aim to appraise EFA revealed four latent factors or subscales that
both mental and physical wellness). The number of comprised the LPHCI. In other words, the EFA
items in each intersecting content and domain area on identified four groups of observed variables (test
the blueprint (Figure 2) are only approximations of the items) that clustered together to form four LPHCI
total number of items that comprised the initial pool subscales. Data from a second sample of participants
of items. were entered into a CFA, which revealed acceptable
Kalkbrenner and Gormley (2020) began model fit. Finally, tests of internal consistency
synthesizing content and developing their scale (step 4) reliability (Cronbach’s coefficient alpha) produced
by referring to their theoretical framework consisting acceptable reliability evidence for the LPHCI scales.
of Chopra’s and Servan-Schreiber’s theories as well as Kalkbrenner and Gormley (2020) argued that α > .70
a blueprint (Figure 2) to guide the item development was acceptable reliability evidence for the LPHCI
process. They made sure the items were brief, clear, because (a) the LPHCI is an attitudinal screening tool,
and written at approximately a sixth-grade reading level (b) there were no diagnostic or high-stakes implications
[Link]
DOI: [Link] 12
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 13
Kalkbrenner, The MEASURE Approach

for test takers, (c) the construct of measurement was Educators who teach classes in testing, research
exploratory, and (d) three of the four LPHCI subscales methods, assessment, or psychometrics can refer to
were comprised of relatively few items (shorter scales The MEASURE Approach for lesson planning and
tend to produce lower reliability estimates). potentially, include this article as a required or
Collectively, the EFA/CFA results and tests of internal supplemental course reading. Practitioners who work
consistency reliability produced adequate validity and in applied social science fields (e.g., counseling,
reliability estimates for the LPHCI. However, a psychology, or social work) can refer to The
number of poorly worded items were removed during MEASURE Approach to instrument development to
the expert review and validity evidence phases, which review the rigor of existing instrumentation before use
made Kalkbrenner and Gormley (2020) concerned with their clients. The MEASURE Approach was
about the content validity of the final factor solution. designed to help social scientists gain a greater
To this end, Kalkbrenner and Gormley (2020) revised understanding of the instrument development process
the content of the LPHCI items to reflect a more and validating scores on tests, which is consistent with
comprehensive scope of the construct of measurement the Standards for Educational and Psychological
and repeated steps 5 to 7. The results of the second Testing (2014) and has potential to increase assessment
round of item development, data collection, and literacy and promote methodological rigor in social
psychometric testing yielded adequate content validity, sciences research. The overview of the MEASURE
construct validity, and internal consistency reliability Approach presented in this manuscript can serve as a
estimates for scores on the LPHCI. one-stop-shop or a single resource that students and
professionals can refer to for outlining empirically
supported steps in the instrument development and
Conclusions score validation process.
The MEASURE Approach was designed to
provide social scientists with a single resource (one-
stop-shop) for outlining seven practical steps in the
instrument development process (Table 1). The
MEASURE Approach is rooted in classical test theory,
with an emphasis on supporting the creation of
measures that demonstrate construct validity evidence
and evidence of test content (Lenz & Wester, 2017).
Such research will require large sample sizes (step 6)
and the emergent evidence will be sample-specific until
future researchers demonstrate reliability and validity
generalizations. A number of further test development
considerations can be relevant to social science
researchers, for example, cognitive interviews
(Peterson et al., 2017), invariance testing (Dimitrov,
2010), higher-order CFA (Credé et al., 2015), using
tests outside the normative sample (Hays & Wood,
2017), high-stakes testing (Boone et al., 2017), and
cultural/language adaptations (Lenz et al., 2017).
The MEASURE Approach to instrument
development has a number of implications for
informing the research and the practice of social
science professionals. Researchers can refer to The
MEASURE Approach to instrument development
when evaluating the rigor of existing instrumentation
or when creating new measures for use in research.

Published by ScholarWorks@UMass Amherst, 2021 13


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 14
Kalkbrenner, The MEASURE Approach

Table 1. The MEASURE Approach to Instrument Development


Step Summary Statement
Make the purpose & State the purpose of the instrument development study and provide a rationale for
rationale clear creating a new instrument by (a) reviewing the extant literature and citing any similar
instruments that already exist and articulate the construct(s) that the existing measures
fail to capture in order to highlight a gap in the existing measurement literature, (b)
discuss how the proposed instrument has significant potential to fill the aforementioned
gap in the measurement literature, and (c) articulate how filling this gap has potential to
advance future research and practice.
Establish empirical Identify a theory (or combination of theories) to set an empirical framework for the item
framework development process. If the literature is lacking an operationalized theory, a researcher
can build their own empirical framework by citing a series of empirical sources to provide
a rationale for the intended construct of measurement and define the scope of the
proposed construct of measurement.
Articulate theoretical A theoretical blueprint (Figure 2) is a tool for enhancing the content validity of a measure
blueprint by organizing the content and domain areas for the construct of measurement and
determining the approximate proportion of items that should be developed across each
content and domain area.
Synthesize content and Referring to the empirical framework (step 2) and theoretical blueprint (step 3),
scale development researchers should first create a large list of potential items individually. Then, researchers
can come together for a meeting(s) to review and compare their separate lists of possible
items and negotiate until a consensus is reached about the pool of items that will be sent
to the expert reviewers. The initial pool of items should include approximately three to
four times as many items that will comprise the final version of the measure.
Use expert reviewers The initial pool of items is sent to approximately three to five external, expert reviewers.
Typically, reviewers are either (a) survey/questionnaire experts, who are well versed in
psychometrics and the mechanics of item development or (b) substantive/subject matter
experts who are knowledgeable in the content area.
Recruit participants Administer the instrument to a small pilot sample that is similar to the target population
and review the pilot data for data imputation and technology issues as well as participant
feedback about item content and readability. Then launch data collection for the main
study by following one of the following criteria, whichever yields a larger sample: (a)
subjects-to-variables ratio of approximately 10:1 or (b) 200 participants.
Evaluate validity and Test the validity (the scale is measuring what it is intended to measure) and reliability
reliability (consistency of scores) evidence of scores on the measure and its subscales. The
MEASURE Approach is centered on demonstrating evidence of construct validity and
internal consistency reliability.

[Link]
DOI: [Link] 14
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 15
Kalkbrenner, The MEASURE Approach

References Browne, R.H. (1995). On the use of a pilot sample for


sample size determination. Statistics in Medicine, 14,
Alvi, M.H. (2016). A manual for selecting sampling 1933 – 1940.
techniques in research. Munich Personal RePEc [Link]
Archive. [Link] Carter, N., Bryant-Lukosius, D., DiCenso, A., Blythe,
[Link]/70218/1/MPRA_paper_70218.pdf J., & Neville, A. (2014). The use of triangulation in
American Educational Research Association, qualitative research. Oncology Nursing Forum, 41(5),
American Psychological Association, National 545–547. [Link]
Council on Measurement in Education, & Joint 547
Committee on Standards for Educational and Cho, E. (2016). Making reliability reliable: A
Psychological Testing (U.S.). (2014). Standards for systematic approach to reliability
educational and psychological testing. coefficients. Organizational Research Methods, 19(4),
[Link] 651–682.
tandards [Link]
Amarnani, R. (2009). Two theories, One theta: A Comrey, A. L., & Lee, H. B. (1992). A first course in
gentle introduction to item response theory as an factor analysis. Lawrence Erlbaum
alternative to classical test theory. The International Credé, M., & Harms, P. (2015). 25 years of higher‐
Journal of Educational and Psychological Assessment, 3, order confirmatory factor analysis in the
104–109. organizational sciences: A critical review and
Bandalos, D.L., & Finney, S.J. (2019). Factor analysis: development of reporting
Exploratory and confirmatory. In G.R. Hancock, recommendations. Journal of Organizational
L. M. Stapleton, & R.O. Mueller (Eds.), The Behavior, 36(6), 845–872.
reviewer’s guide to quantitative methods in the [Link]
social sciences (pp. 98-122). Routledge. Cronbach, L. J. (1951). Coefficient alpha and the
Bardhoshi, G., & Erford, B. (2017). Processes and internal structure of tests. Psychometrika, 16(3), 297-
procedures for estimating score reliability and 334. [Link]
precision. Measurement and Evaluation in Counseling DeVellis, R. F. (2016). Scale development (4th ed.). Sage.
and Development. 50(4), 256–263. Dimitrov, D. (2012). Statistical methods for validation of
[Link] assessment scale data in counseling and related fields.
Beavers, A. A., Lounsbury, J. W., Richards, J. K., American Counseling Association.
Huck, S. W., Skolits, G. J., & Esquivel, S. L. Dimitrov, D. (2010). Testing for factorial invariance
(2013). Practical considerations for using in the context of construct validation. Measurement
exploratory factor analysis in educational and Evaluation in Counseling and Development, 43(2),
research. Practical Assessment, Research & 121–149.
Evaluation, 18(5/6), 1-13. [Link]
Bendermacher, N. (2017). An unbiased estimator of Dunn, T., Baguley, T., & Brunsden, V. (2014). From
the greatest lower bound. Journal of Modern Applied alpha to omega: A practical solution to the
Statistical Methods, 16(1), 674–688. pervasive problem of internal consistency
[Link] estimation. The British Journal of Psychology, 105(3),
Benson, J. (1998). Developing a strong program of 399–412. [Link]
construct validation: A test anxiety eSurveysPro. (2020). [Online survey platform
example. Educational Measurement, Issues and software]. Outside Software Inc.
Practice, 17(1), 10–17. [Link]
[Link] Fowler, F. J. (2014). Survey research methods (5th ed.).
3992.1998.tb00616.x Sage
Boone, W., & Noltemeyer, A. (2017). Rasch analysis: Fu, M., & Zhang, L. (2019). Developing and
A primer for school psychology researchers and validating the Career Personality Styles
practitioners. Cogent Education, 4(1). Inventory. Measurement and Evaluation in Counseling
[Link]

Published by ScholarWorks@UMass Amherst, 2021 15


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 16
Kalkbrenner, The MEASURE Approach

and Development, 52(1), 38–51. Kane, M., & Bridgeman, B. (2017). Research on validity
[Link] theory and practice at ETS. In R. E. Bennett & M.
George, D., & Mallery, P. (2003). SPSS for windows step von Davier (Eds.), Methodology of educational
by step: A simple guide and reference. 11.0 update (4th measurement and assessment. Advancing human
ed.). Allyn & Bacon. assessment: The methodological, psychological and policy
Gregory, R. J. (2016). Psychological testing: History, contributions of ETS (p. 489–552). Springer Science
principles and applications (Updated 7th edition). + Business Media. [Link]
Pearson. 3-319-58689-2_16
Hays, D., & Wood, C. (2017). Stepping outside the Knekta, E., Runyon, C., & Eddy, S. (2019). One size
normed sample: Implications for doesn’t fit all: Using factor analysis to gather
validity. Measurement and Evaluation in Counseling and validity evidence when using surveys in your
Development. 50(4), 282–288. research. CBE Life Sciences Education, 18(1).
[Link] [Link]
Hertzog, M. (2008). Considerations in determining Kuder, G. F., & Richardson, M. W. (1937). The
sample size for pilot studies. Research in Nursing & theory of the estimation of test reliability.
Health, 31(2), 180–191. Psychometrika, 2(3), 151-160.
[Link] [Link]
Hu, L., & Bentler, P. (1999). Cutoff criteria for fit Lambie, G., Blount, A., & Mullen, P. (2017).
indexes in covariance structure analysis: Establishing content-oriented evidence for
Conventional criteria versus new psychological assessments. Measurement and
alternatives. Structural Equation Modeling: A Evaluation in Counseling and Development. 50(4), 210–
Multidisciplinary Journal, 6(1), 1–55. 216.
[Link] [Link]
Ikart, E. (2019). Survey questionnaire survey Lenz, A., Gómez Soler, I., Dell’Aquilla, J., & Uribe, P.
pretesting method: An evaluation of survey (2017). Translation and cross-cultural adaptation of
questionnaire via expert reviews technique. Asian assessments for use in counseling
Journal of Social Science Studies, 4(2). research. Measurement and Evaluation in Counseling
[Link] and Development. 50(4), 224–231.
IMPAQ. (2020). [Online survey platform software]. [Link]
IMPAQ International, LLC. Lenz, A., & Wester, K. (2017). Development and
[Link] evaluation of assessments for counseling
Kalkbrenner, M.T., & Gormley, B. (2020). professionals. Measurement and Evaluation in
Development and initial validation of scores on Counseling and Development. 50(4), 201–209.
the Lifestyle Practices and Health Consciousness [Link]
Inventory (LPHCI). Measurement and Evaluation in Maslow, A. (1943). A theory of human motivation.
Counseling and Development. 53(4), 219-237. Psychological Review, 50(4), 370–396.
[Link] [Link]
Kalkbrenner, M.T., Neukrug, E.S., & Griffith, S.A., McNeish, D. (2018). Thanks coefficient alpha, we’ll
(2019). Appraising counselors’ attendance in take it from here. Psychological Methods, 23(3), 412–
counseling: The validation and application of the 433. [Link]
Revised Fit, Stigma, and Value Scale. Journal of Menold, J., Jablokow, K., Purzer, S., Ferguson, D., &
Mental Health Counseling. 4, 21-35. Ohland, M. (2015). Using an instrument blueprint
[Link] to support the rigorous development of new
Kane, M. (1992). An argument-based approach to surveys and assessments in engineering
validity. Psychological Bulletin, 112(3), 527–535. education. ASEE Annual Conference and Exposition,
[Link] Conference Proceedings, 122. American Society for
Kane, M. (2010). Validity and fairness. Language Engineering Education.
Testing, 27(2), 177–182. Mvududu, N. H., & Sink, C. A. (2013). Factor analysis
[Link] in counseling research and practice. Counseling
[Link]
DOI: [Link] 16
Kalkbrenner: The MEASURE Approach
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 17
Kalkbrenner, The MEASURE Approach

Outcome Research and Evaluation, 4(2), 75-98. Summers, S., & Summers, S. (1992). Instrument
[Link] development: Writing the items. Journal of Post
Nájera Catalán, H. (2018). Reliability, population Anesthesia Nursing, 7(6), 407–410.
classification and weighting in multidimensional [Link]
poverty measurement: A Monte Carlo Study. Social SurveyMonkey. (2020). [Online survey platform
Indicators Research, 142(3), 887–910. software]. Bain & Company, Inc., Fred Reichheld
[Link] and Satmetrix Systems, Inc.
Nguyen, D., Wright, E., Dedding, C., Pham, T., & [Link]
Bunders, J. (2019). Low self-esteem and its _lp&ut_source2=sem&ut_source3=header
association with anxiety, depression, and suicidal Swank, J., & Mullen, P. (2017). Evaluating evidence
ideation in Vietnamese secondary school students: for conceptually related constructs using bivariate
A cross-sectional study. Frontiers in correlations. Measurement and Evaluation in Counseling
Psychiatry, 10(SEP). and Development. 50(4), 270–274.
[Link] [Link]
Peterson, C., Peterson, N., & Powell, K. (2017). Taber, K. S. (2018). The use of Cronbach’s alpha
Cognitive interviewing for item development: when developing and reporting research
Validity evidence based on content and response instruments in science education. Research in Science
Processes. Measurement and Evaluation in Counseling Education, 48(6), 1273–1296.
and Development. 50(4), 217–223. [Link]
[Link] Tate, K., Bloom, M., Tassara, M., & Caperton, W.
Qualtrics. (2020). [Online survey platform software]. (2014). Counselor competence, performance
SAP. [Link] assessment, and program evaluation: Using
Qualtrics Sample Services [Online sampling service psychometric instruments. Measurement and
service]. (2020). Evaluation in Counseling and Development, 47(4), 291–
[Link] 306. [Link]
services/online-sample/ Urdan, T. C. (2010). Statistics in Plain English (3rd ed.).
REDCap. (2020). [Online survey platform software]. New York, NY: Routledge.
Vanderbilt. [Link] Vagias, W. M. (2006). Likert-type scale response anchors.
Redi Data. (2020). [data services and direct marketing [Link]
solutions] [Link] Assessment/likert-
Schreiber, J. B., Nora, A., Stage, F. K., Barlow, E. A., type%20response%[Link]
& King, J. (2006). Reporting structural equation Wolf, E. J., Harrington, K. M., Clark, S. L., & Miller,
modeling and confirmatory factor analysis results: M. W. (2013). Sample size requirements for
A review. Journal of Educational Research, 99(6), 323- structural equation models: An evaluation of
337. [Link] power, bias, and solution propriety. Educational and
Sharon, T. (2018). 43 ways to find participants for research. Psychological Measurement, 73(6), 913–934.
Medium. [Link] [Link]
ways-to-find-participants-for-research-
ba4ddcc2255b
Singer, E., & Bossarte, R. M. (2006). Incentives for
Survey Participation: When are they “Coercive”?
American Journal of Preventive Medicine, 31(5), 411–
418.
[Link]
Strauss, M., & Smith, G. (2009). Construct validity:
Advances in theory and methodology. Annual
Review of Clinical Psychology, 5(1), 1–25.
[Link]
53639

Published by ScholarWorks@UMass Amherst, 2021 17


Practical Assessment, Research, and Evaluation, Vol. 26 [2021], Art. 1
Practical Assessment, Research & Evaluation, Vol 26 No 1 Page 18
Kalkbrenner, The MEASURE Approach

Citation:
Kalkbrenner, M. (2021). A Practical Guide to Instrument Development and Score Validation in Social Sciences
Research: The MEASURE Approach. Practical Assessment, Research & Evaluation, 26(1). Available online:
[Link]

Corresponding Author
Michael Kalkbrenner
Office: O’Donnell 202 J
New Mexico State University
Las Cruces, NM, 88003-8001

Email: mkalk001 [at] [Link]

[Link]
DOI: [Link] 18

You might also like