Mixed Methods for Instrument Development
Mixed Methods for Instrument Development
For the most optimal reading experience we recommend using our website.
A free-to-view version of this content is available by clicking on this link, which
includes an easy-to-navigate-and-search-entry, and may also include videos,
embedded datasets, downloadable datasets, interactive questions, audio
content, and downloadable tables and resources.
Our study’s weighting between quantitative and qualitative phases enabled a more thorough under-
standing of a complex phenomenon like IPV [interpersonal violence] disclosure in emergency de-
partments that would not have occurred using a RCT alone. (Catallo et al., 2013, p. 9)
The use of qualitative research methods prior to designing and administering a survey will help en-
sure a more valid survey, open the researcher up to new hypotheses to test, provide a means of
exploring mixed analysis techniques, and further review and revise the survey. (Crede & Borrego,
2013, p. 75)
In This Chapter
Instrument Development
Data collection is an important part of any evaluation. The development of data collection instruments can
occur as part of a larger evaluation. However, instrument development can also be a standalone endeavor.
Either way, evaluators are typically involved in the development or adaptation of instruments for data collec-
tion purposes. Hence, this chapter is dedicated to the explanation of different approaches to development of
instruments for data collection as viewed through the various philosophical stances and evaluation branches.
Data sources are commonly divided into primary and secondary (Mertens, 2015b). Primary data sources
include such sources as people interviewed, surveyed, or tested; documents produced as part of the evalu-
ation study (e.g., student portfolios or art projects); and geographic information system mapping that occurs
during the study. Secondary sources include data existing before the study is designed and created for pur-
poses other than the evaluation, such as administrative records, websites, published or unpublished studies,
and extant data bases. The focus of this chapter is on the development of primary data collection instruments
such as questionnaires, surveys, checklists, interview guides, and observation guides.
Guidance for the development of instruments can be found in most major research and evaluation textbooks
(see Mertens, 2015b; Mertens & Wilson, 2012). However, the guidance rarely extends to how to use a mixed
methods design for instrument development. The type of mixed methods designs used for instrument de-
velopment differ based on the branch of the evaluation theory tree in which evaluators situate themselves.
As in Chapter 2 on evaluations of interventions, this chapter looks at designs of mixed methods studies for
instrument development for the four branches of evaluation—methods, use, values, and social justice—and
dialectical pluralism (DP). Study examples of the use of mixed methods to develop quantitative instruments
for all four branches and DP are presented. An additional example of mixed methods to develop a qualitative
instrument is presented under the Values branch. The use of mixed methods for this purpose is less common,
but the example provides insight into how mixed methods can be used for this purpose. Perhaps it will stimu-
late others to consider this approach.
As mentioned, one of the most common uses of mixed methods is in the process of developing data collection
instruments. This can take the form of using a qualitative stage to get ideas for what should be included in an
instrument, followed by a quantitative stage to establish the instrument’s psychometric qualities. It could also
take the form of a quantitative stage of development first, followed by a qualitative stage to explore the applic-
ability of the instrument in diverse cultural contexts. Or it could take a multistage form with several iterations
of qualitative and quantitative cycles to inform refinement of instruments.
Clarke et al. (2011) used a concurrent mixed methods design, collecting quantitative and qualitative data si-
multaneously to establish the psychometric properties of a scale to measure mental well-being in adolescents
and to get insights into the adolescents’ perceptions of the scale items. See Sample Study 3.1 for details of
the Clarke et al. (2011) study.
Concurrent mixed methods design for instrument development: Measuring mental health in adolescents
(Clarke et al., 2011).
Problem: Mental health issues in adolescents are traumatic and can have far-reaching consequences in terms
of quality and longevity of life and social costs. In the United Kingdom, the National Institute for Health and
Clinical Excellence established the prevention of emotional and behavioral problems and promotion of men-
tal and emotional health as a priority. The evaluators in this study identified a problem because of a lack of
validated measures of mental illness and health in childhood and adolescence.
Evaluand: The evaluators validated the Warwick-Edinburgh Mental Well-Being Scale (WEMWBS) with ado-
lescents in England and Scotland.
Sample: The sample for the quantitative portion of the study consisted of 1,650 students ages 9 to 11 in Eng-
land and ages 13 to 16 in Scotland. The sampling strategy for this portion of the study was as follows: “Six
schools (three each from two cities, one in Scotland and one in England) were selected to reflect variation
by geographical location, socioeconomic deprivation (based on proportion of children in receipt of free school
meals) and educational attainment (proportion of children achieving 5+ GCSE grades A–C (England)/5+
awards at SCQF Level 4 (Scotland)” (p. 3). A random sample of 12% of these students were asked to take the
measure twice to establish reliability. The evaluators also conducted 12 focus groups with 40 students ages
13–14 and 40 who were ages 15–16. They were members of the same cohort (who were not included in the
1,650 students used to establish the measure’s psychometric properties). The focus groups were conducted
in same-sex, same-age groups, two in each of the schools in the larger sample.
Data collection: The WEMWBS was administered to the quantitative sample during lesson times. The focus
group participants completed the measurement instrument and then discussed their reactions to the instru-
ment in groups of six to eight students for approximately an hour.
Data analysis: Descriptive statistics and multiple linear regression analysis were conducted on the quantita-
tive data to compare responses by sociodemographic characteristics. Psychometric statistics were calculat-
ed to obtain the validity and reliability information on the instrument. Confirmatory factor analysis was done
to test the factor structure of the instrument. Qualitative data were coded based on the protocol used in the
focus groups. The codes were combined into themes and the themes were compared to see if there were
differences by gender, location, or age.
Results: The Cronbach alpha was high (.87) with a 95% confidence interval of .85–.88, and the test-retest
reliability was moderately high (.66), with a 95% confidence interval of .59–.72, indicating an acceptable level
of reliability. Validity was judged to be high based on correlations with other measures of similar constructs.
The factor analysis demonstrated one underlying factor. The qualitative data indicated that the participants
generally found the scale “simple, short, and easy to complete” (p. 6). However, they also indicated that some
of the items were redundant and some were embarrassing, and they also suggested new items that could be
added to the scale.
Benefits of using MM: The instrument developers concluded that they had produced a valid and reliable in-
strument for this age group. Because of the concurrent mixed methods design, the qualitative findings were
not used to make changes to the instrument. The authors stated that they did not see the benefit of making
changes based on the qualitative findings because they did not want to lose the continuity with the adult ver-
sion of this instrument that had been previously validated.
1. The first step to develop an instrument, whether using mixed methods or not, is to conduct a compre-
hensive literature review to determine the current status of instruments in existence related to the
concept of interest and/or to identify the full scope of literature about the concept of interest in order
to develop items for an instrument. In the Clarke et al. (2011) study, the investigators examined the
literature, not to identify the full scope of literature related to the concept but to identify measures al-
ready in use to measure mental health in the United Kingdom. Their review of literature revealed that
there was no scale validated “to monitor teenage population mental wellbeing and to evaluate inter-
ventions and programmes targeted to this age group” (p. 488).
2. A common mixed methods design used to develop instruments in the Methods branch is the sequen-
tial mixed methods design where the scope of the concept is determined using qualitative methods,
including literature review and interviews with relevant stakeholders. This is followed by the use of
quantitative methods to test the quality of the instrument using standard psychometric analysis strate-
gies. This quantitative stage of the study can be combined with the concurrent use of qualitative
methods to determine problems with administration or interpretation of meaning in the items. Use of
a concurrent mixed methods design is a bit unusual for the purpose of instrument development be-
cause it does not allow for the modification of the instrument based on findings from one portion of
the study or the other. However, evaluators can use this design to establish psychometric properties
of instruments and to understand the perceptions of the people who complete the measure about the
measure itself. This is the design Clarke et al. (2011) chose to establish the validity and reliability of
the instrument in their study.
3. In the quantitative portion of a mixed methods study to develop an instrument, the procedures mirror
practice in instrument development without the mixed methods component in that the concern is with
calculating the reliability and validity coefficients from a defined population.
4. The addition of a qualitative portion in mixed methods allows for the investigation of problems with
administration and problematic content either because of difficulties in understanding the meaning or
perceptions of relevance or comprehensiveness of the items.
5. If evaluators are examining an instrument validated on one population for use with another popula-
tion, they will have access to the established psychometric properties established for the original pop-
ulation. However, this may contribute to a reluctance to change items in the scale because it would
have implications for the established reliability and validity estimates.
1. Race relations can be harmonious or they can be contentious. A community service agency wants an
instrument to measure the quality of race relations that is reliable and valid and perceived by the in-
tended respondents as containing appropriate content. They also want to be sure that it is easy for
people to complete. Design a concurrent MM study to develop such an instrument. What are the ad-
vantages and limitations of using this MM design for this purpose?
2. Instead of using a concurrent MM design for designing an instrument, what would the design look like
if it was sequential or multistage MM design? How would that change what the evaluator does?
In both the Methods and the Values branch, development of an instrument to collect quantitative data includes
mixed methods in the form of literature review, stakeholder input, and establishment of psychometric proper-
ties such as reliability and validity. When the instrument development mixed methods study comes from the
Values branch, the literature review and stakeholder input take on increased presence in the study with an
aim to obtain improved understanding of complex cultural and linguistic issues present in the diverse stake-
holder community, as well as to increase the breadth of items to reflect the experiences of the stakeholders
in all its diversity and complexity. Literature review is treated as a qualitative data collection strategy, that is,
document review and the results of the document review are used to enrich understanding of the construct
to be measured by the instrument when it is developed. If an ethnographic methodology is used as the
qualitative part of the study, then the evaluator will spend time observing in naturalistic settings, interviewing
stakeholders formally and informally, and interacting with stakeholders multiple times over the course of the
instrument development to ensure he or she is capturing the complexity of the concept in culturally respon-
sive ways. The interaction with stakeholders (intended targeted population and experts) continues through
the piloting of the instrument; this is followed by larger-scale testing of the instrument to establish reliability
and validity.
Mixed methods designs can also be used to develop qualitative data collection instruments (e.g., interview
guides or observation guides). This is less commonly done by evaluators, but the approach represents a po-
tential area of growth for mixed methods design. Hence, I include Sample Study 3.3 (Phillips, Dwan, Hep-
worth, Pearce, & Hall, 2014) that demonstrates the use of a mixed methods design to develop a qualitative
evaluation strategy.
Crede and Borrego (2013) used an ethnographic methodology as the primary approach in the design of a
quantitative survey instrument to determine factors that influence retention in graduate engineering programs.
They began with a literature review and then collected ethnographic data over a 9-month period as the begin-
ning point for their instrument development. The design and process of their study are presented in Sample
Study 3.2.
Quantitative instrument development with qualitatively dominant approach: Retention of graduate engineering
students (Crede & Borrego, 2013).
Problem: In the United States, the preparation of graduate-level science and engineering students is viewed
as necessary to maintain a technological edge to foster innovation and a workforce able to contribute to the
national economy. The graduation rate from doctoral-level engineering programs is lower for U.S. students
than for international students.
Design: Sequential exploratory mixed methods with extensive ethnography. The evaluators conducted an
ethnographic study for 9 months before they developed an instrument to measure factors that influence re-
tention in engineering graduate school.
Sample: The qualitative sample came from three research groups from two engineering departments at a
large public university. Electrical engineers were the most frequent (n = 20 from six countries) and aerospace
engineering was the smallest group (n = 4 from China and the United States); 12 students came from a mul-
tidisciplinary program in aerospace and electrical engineering. These students were from the United States
and India. The quantitative sample came from four universities: a large public university in the Midwest, a
historically black university in the East, a large private university in the West, and a large public university in
the East.
Data collection: The qualitative data collection consisted of an extensive literature review followed by 9
months of ethnographic observations and interviews that focused on the student experience in terms of lan-
guage and culture. Ethnographic data collection included formal (20 semistructured interviews) and informal
interviews, lengthy periods of observation, and participation in most research group activities. The quantita-
tive data collection involved the use of a survey based on the results of the qualitative data collection. The
students and experts in graduate engineering education were involved with the evaluation team members by
reviewing drafts and commenting on clarity of language and survey items. The draft instrument was piloted
with 50 graduate engineering students. The evaluators revised the survey based on the results of the draft
administration results and an additional review by an expert in engineering education and a small group of
students who had participated in the ethnographic part of the study. The administration of the final survey was
distributed online to students at four universities across the United States and was completed by 837 students
(18.5% response rate).
Data analysis: In the qualitative analysis, themes that arose in the literature on graduate student retention
were combined with themes that emerged from qualitative analysis of the ethnographic data. The quantitative
analysis consisted of calculation of an internal consistency index (Cronbach alpha), an indicator of validity.
Results: The qualitative themes that emerged from the analysis included international diversity; expectations
of students and faculty; climate in terms of acceptance, belonging, and competitiveness; organizational sup-
port from experienced students, resources, and advisors; individual preferences such as belief about the im-
portance of the work and values of diversity and teamwork; and feeling valued in terms of the research group.
These themes formed the basis of the survey the evaluators developed; the survey consisted of items related
to the graduate school experience that were measured using a five point Likert-type scale. The final survey
had seven subscales, with Cronbach alpha values that ranged from .86 to .63.
Benefits of using MM: The survey was developed based on the ethnographic data. At several points during
the quantitative testing of the survey, the evaluators included qualitative data collection in the form of review
with students and experts. The internal consistency was improved after the expert review resulted in reword-
ing of several questions and the addition of several items related to student development (learning). The
ethnographic data helped the evaluators create an instrument reflective of the concept of international diver-
sity in engineering graduate research groups.
Phillips et al. (2014) used a qualitatively dominant mixed methods design to develop qualitative data-gathering
instruments and strategies that could be used to evaluate small, community-based organizations. They were
interested in the development of qualitative instruments and strategies as well as in maximizing the trustwor-
thiness and authenticity of them. These are criteria from the constructivist paradigm that are often used in
qualitative studies that parallel the concept of validity in the postpositivist paradigm.
Values branch mixed methods design for development of qualitative instruments and strategies for data col-
Problem: Limited evaluations have been conducted to assess the quality of health care delivered by small,
community-based organizations. The staff providing the services have limited time to engage in evaluation
activities because of time constraints and other pressures.
Evaluand/instrument and strategies: The Qualitative Rapid Appraisal Rigorous Analysis (Q-RARA) is a set
of qualitative instruments and strategies designed specifically for evaluating small, community-based health
care organizations. The Q-RARA is predominantly qualitative with a small component of quantitative data be-
ing collected. It was developed to study nurses in general practice in Australia.
Design: Qualitatively dominant Values branch concurrent mixed methods design. The design combined the
use of a constructivist version of rapid appraisal (Nyanzi, Manneh, & Walraven, 2007) and qualitative mixed
methods designs. The evaluators “chose to work within an interpretive theoretical perspective involving ex-
tensive analytic input from researchers by choosing to locate the study within the nurses’ work environments,
and to interpret their actions, positions and the meanings they gave to their roles” (p. 561). Data were collect-
ed concurrently and combined at the stage of analysis.
Sample: Data were collected from “25 practices that varied in size, organizational structure and geographic
location across two Australian states” (p. 561). The sample included doctors, nurses, and practice managers,
although the authors do not tell us the specific number of people in the sample.
Data collection: The qualitative data collection strategies included using in-depth interviews with nurses, doc-
tors, and practice managers; structured observations and unstructured observations of nurses’ activities; pho-
tographs of nurse-identified worksites; floor plans; and field notes. Quantitative data collection included a
questionnaire that summarized staff numbers and working hours and a social scan. Social scanning “included
the collection of publicly available census and health service provision data to describe the socio-geographic
setting of each practice” (p. 561). The data were collected by a single evaluator at each site for a period of
one day in order to minimize disruption of the provision of services.
Data analysis: The evaluation team analyzed the data to determine the quality and rigor of the data produced
and to determine the capacity of the Q-RARA to adequately reflect the stakeholders’ views of the trustwor-
thiness and authenticity of the data produced by this set of instruments and strategies. Lincoln and Guba
(1985) proposed that “trustworthy data are credible, transferable, dependable, and confirmable. Data should
also be assessed for authenticity, a domain that includes fairness, and educative, ontological, tactical, and
catalytic authenticity” (p. 567). The evaluators analyzed the data related to trustworthiness by comparing their
processes and findings against the criteria defined by Lincoln and Guba, especially regarding the credibility
of the data. Credibility was determined by ascertaining the extent to which prolonged engagement and per-
sistent observation were possible in the form of determining the access to the backstage of general practice
that the evaluators were permitted in each setting. In addition, peer debriefing was conducted by means of
the senior member of the team reviewing the field notes and reflections of each site visit to provide feedback
as a peer to the evaluators who collected the data. The evaluation team met fortnightly to offer opportunities
for critical reflection (i.e., progressive subjectivity: a process designed to allow for the evaluators to become
aware of any biases they hold, become aware of identify shifts in their understandings throughout the study,
and examine ways to avoid the influence of their own biases in the data collection). Member checks were
conducted to share and check the findings and interpretations with the participants by providing summaries
of the findings for verification. The authors also posted deidentified findings on a website for practicing nurses
to obtain feedback. Transferability was determined by analyzing the extent to which background information
was provided in each case to provide sufficient information for readers to make a judgment about the transfer-
able nature of the findings to their own setting. Dependability, that is, the stability of the study procedures and
methods over time, was assessed through multiple meetings over the course of the study with multiple stake-
holders to review the process and document any changes. The final trustworthiness criterion is confirmability,
or the documentation that the findings are grounded in data. A confirmability audit was conducted that linked
all case summaries, analysis documents, and nonpublic documents to deidentified codes for each participant/
site.
Authenticity is also multifaceted. Fairness, the extent to which diverse stakeholder perspectives are included
and used in the formulation of recommendations, was determined by analyzing the range of participants and
documenting their involvement in the formulation of recommendations. Ontological authenticity, that is, the
extent to which stakeholders’ understanding of reality is improved or expanded, was assessed by document-
ing the inclusion of diverse perspectives that might not have been broadly known. Educative authenticity, that
is, the extent to which stakeholders develop a better understanding of the perspectives of other stakeholders,
was determined by analyzing the process used to distribute the findings to diverse audiences. Catalytic au-
thenticity, that is, the extent to which participation in the evaluation and exposure to the findings elicits action
and change, was determined by analyzing changes in policies and practices. Tactical authenticity, that is, the
extent to which stakeholders feel empowered by having participated in the evaluation, was determined by
documenting future uptake of the study recommendations by participants.
Results: The Q-RARA was judged to be trustworthy because, even though prolonged engagement was pre-
cluded by the one-day timeframe of data collection, the evaluators were granted access to observe the par-
ticipants in meetings, on their rounds, and as they moved across the spaces in the practice in each of the
settings. Persistent observation was also not possible; however, the evaluators did repeat their observations
of specific phenomenon throughout the day at each site. The peer debriefing between the evaluators in the
field and the senior member of the team resulted in clarification of the interview process and questions. The
process of progressive subjectivity resulted in challenges to assumptions evaluators had made during the
study. Member checks resulted in feedback on specific topics. Transferability was viewed as sufficient be-
cause of the extensive background information included for each case site. Dependability was also viewed as
sufficient because of the documentation of changes that occurred in procedures and protocols. Evidence of
confirmability revealed that links could be made between all the data collection documents and their data with
the sites in which the data were collected. The instruments and procedures were viewed as fair because many
diverse stakeholders were included and their advice was used to develop the recommendations. Ontological
authenticity was deemed to be acceptable because of the differences in perspectives revealed, particularly
with regard to the less recognized role of nurses as educators. Educative authenticity was supported by the
increased understanding by doctors of the multiple roles played by nurses. Catalytic authenticity was judged
to be strong because “the national body responsible for preparing doctors for general practice has now fund-
ed trials of practice nurses training general practice registrars” (p. 569). Tactical authenticity was also judged
to be strong because the findings were adopted by the chief nurse in the Department of Health and Ageing,
and additional funding has been provided to support the role of nurses as educators.
Benefits of using MM: The authors noted that prior evaluations of this type of organization have relied on col-
lection of data by means of quantitative measures or single qualitative methods such as interviews or focus
groups. “Single method studies do not capture the richness and variety of organizational functioning, the ways
that staff members interact and use their time, and the impact of the spatial environment of the organization
on their work” (p. 559). The use of a mixed methods design resulted in an approach that included primarily
qualitative instruments and strategies, with some quantitative data collection, that was well received by the
participants, particularly by the practice nurses who felt that their work was validly portrayed and evaluated.
They appreciated the nondisruptive nature of the data collection. The use of this mixed methods approach
permitted the collection of a range of data, inclusion of a multidisciplinary team, and use of an iterative ap-
proach to data analysis that “allowed us to engage in a dialogue with the data and produce authentic and
trustworthy findings” (p. 572).
Guidance for Designing a Mixed Methods Instrument Development Study in the Values Branch
1. Conduct a comprehensive literature review to determine what is already known about the phenome-
non under study.
2. Plan a qualitative study such as an ethnographic approach or a multimethod qualitative strategy that
includes purposeful sampling criteria based on established criteria for relevant stakeholder groups,
being inclusive of diversity in the population.
3. Use qualitative data collection methods that include observation, formal and informal interviews, and
document reviews that focus on the cultural considerations relevant to the phenomenon you want to
measure.
4. If the primary focus is on the development of a quantitative instrument, involve stakeholders in the
interpretation of the qualitative data and its use to formulate an instrument that is responsive to the
data already collected. Disaggregate both quantitative and qualitative data by dimensions of diversity
relevant in that context (e.g., international and domestic students). If the primary focus of the study is
to test the quality of qualitative instruments and procedures, consider use of the methods used in the
Phillips et al. (2014) study that used Lincoln and Guba’s (1985) framework for assessing trustworthi-
ness and authenticity.
5. Pilot test the instrument with a representative quantitative sample for a quantitative instrument and
with a qualitative sample for a qualitative instrument.
6. Involve experts and other stakeholders to review the results of the pilot test; make revisions of the
instrument.
7. Conduct a larger-scale quantitative study with the instrument to establish reliability and validity for a
quantitative instrument. Conduct a follow-up study for a qualitative instrument to determine its sus-
tained trustworthiness and authenticity.
1. The Centers for Disease Control and Prevention has a program to prevent sexually based violence
that focuses on working with men to address cultural perceptions of masculinity and healthy relation-
ships between men and women (as well as between men and men and women and women). Design
an instrument development evaluation using qualitative methods as the primary method combined
with a quantitative strategy to establish reliability and validity. How does the sensitivity of the concept
influence your design? How does the challenge to cultural norms that support oppression of women
or vulnerable partners influence the design? How would you change the design if you were develop-
ing a qualitative instrument?
2. What measurements are you using in your evaluation work? How could these measurements be im-
proved by using a qualitative phase as a basis for revision? Design a Values branch MM evaluation
to improve the measurement instrument you use.
Instrument development has largely been the purview of quantitative evaluators who use statistical calcula-
tions to demonstrate the reliability and validity of the instruments. However, evaluators in the Use branch point
out that many questions cannot be answered by quantitative data alone that are pertinent to establishing the
quality of measurement instruments. Hence, they recommend the use of mixed methods to answer a wider
range of questions related to the validation of inferences that can be drawn from the results based on quan-
titative instruments. Combining strategies such as factor analysis and cognitive interviewing can provide the
evaluator with greater confidence about construct validity.
Multiphase instrument development: Measuring participation by stakeholders in evaluations (Daigneault & Ja-
cob, 2014).
Problem: Evaluations sometimes claim to be participatory but do not proffer evidence that supports the extent
or nature of the participation by stakeholders.
Evaluand: Participatory Evaluation Measurement Instrument (PEMI). This purpose of this instrument is to
measure stakeholder participation in evaluations.
Design: Multistage MM instrument development in the Use branch. The authors noted that their initial plan
was to conduct a quantitative study, but they obtained qualitative data unexpectedly in the form of open-end-
ed comments and informal e-mail correspondence that turned the study into a mixed methods study. Thus,
the first stage of the study is a concurrent MM design with quantitative and qualitative data being collected
simultaneously. Then the authors extended the study to include a second stage that included revision of the
instrument and collection of additional quantitative data for validation purposes.
Sample: For the quantitative portion of the study, they used a purposive sampling strategy to select 40 evalu-
ation studies that displayed various levels of stakeholder participation from peer-reviewed journals published
between 1985 and 2010.
Data collection: The authors developed an instrument, the Participatory Evaluation Measurement Instrument,
based on a review of literature on stakeholder participation and a framework on participation developed by
Cousins and Whitmore (1998). Their instrument contained items about the extent and nature of participation
using a 5-point scale. As a quantitative validation of the PEMI, the evaluators used the PEMI to assess the
degree of participation exhibited in 40 articles that were screened to represent participatory approaches to
evaluation. The intention of the first stage was to collect quantitative data that could be used to establish the
interrater reliability using Cohen’s kappa and the intraclass correlation coefficient. The evaluators shared the
results of their application of the PEMI with the authors of the 40 studies by means of an online survey.
The online survey had two sections. The first section focused on questions relative to the PEMI. Re-
spondents had to check boxes about types of participants, steps in which they were involved, and
their level of control on the evaluation process. On the next page of the online survey, 5-point indices
were automatically generated from respondents’ answers for each dimension and for the overall lev-
el of participation. These scores were presented to the authors of the evaluation cases for reactions
in the first part of the survey. Respondents’ opinions were measured on an ordinal scale of agree-
ment and an open-ended question asked respondents to justify their choice (i.e., ‘‘Why?’’). The sec-
ond part of the survey contained 11 four-point Likert-type questions from which the calculation of the
Evaluation Involvement Scale (EIS) was derived (Toal, 2009).
Respondents generated qualitative data in their responses to the survey’s open-ended questions that con-
tained richer data than the evaluators had anticipated, and so they decided to conduct a thematic analysis of
the qualitative comments. In addition, two of the authors of the original 40 published studies supplemented
their survey responses with substantive comments made in e-mails; these were also included in the qualita-
tive data analysis. Following the analysis from this round of data collection, the evaluators decided to revise
the instrument to respond to the concerns raised by the authors of the articles that were part of the literature
review. In the final stage, the evaluators used the revised instrument to again quantitatively rate the level of
participation in the published articles from the first stage. They then surveyed the authors again, sharing with
them the change in scores that resulted when the revised instrument was used to rate their work.
Data analysis: In the first stage, the authors’ scores were compared with the raters’ scores for agreement. The
evaluators noted that the article authors only somewhat agreed that the PEMI score accurately represented
the level of participation from their point of view. The evaluators decided to integrate this quantitative finding
with a qualitative analysis of the comments respondents wrote for the open-ended questions and e-mails. The
evaluators then decided that the qualitative data included rich data, so they conducted a thematic analysis
of the survey and e-mail comments. They used these data to revise the instrument and then conducted a
quantitative stage of data collection to obtain psychometric data about the quality of the instrument. Analysis
consisted of comparison of the scores from the first use of the instrument to those obtained with the revised
instrument. The Mann-Whitney U test was used to compare the responses of article authors whose scores
stayed the same in the first and second administration of the scale with those whose scores increased.
Results: The initial quantitative stage indicated an acceptable level of reliability and weak to strong convergent
validation. However, when the evaluators completed the qualitative thematic analysis, they found that people
who felt the instrument underrepresented the extent of participation had expressed disagreement with the
way the concept of participation was rated, the relevance of participation by different types of stakeholders,
and inability of the scale to capture the multifaceted nature of the participation. The revised instrument re-
sulted in 20% of the scores being unchanged and increased the scores for the other 80%, thus supporting
increased validity of the instrument. The level of disagreement about the accuracy of the scores as measures
of stakeholder participation yielded by the PEMI decreased, again supporting increased validity.
Benefits of using MM: The use of mixed methods led the evaluators to revise the PEMI: “The results generat-
ed in the mixed methods phase provide additional evidence of the necessity of revising the PEMI. Quantitative
data indicated that the alignment between the scores derived from the PEMI and the respondents’ opinions
with respect to the level of participation was only partial. Qualitative data revealed that an overarching theme
derived from the respondents’ answers was the underrepresentation of stakeholder participation by the PE-
MI’s scores, thus suggesting the need to find a less conservative concept structure for the revised version
of the instrument” (p. 16). The evaluators used the integrated results of the first stage to revise and test the
validity of the PEMI during a second quantitative stage.
The instrument development and construct validation developed by Onwuegbuzie et al. (2010) contains 10
phases:
1. Conceptualize the construct of interest. Conduct a literature review and consult with a diverse set of
experts on the construct to be measured. Conducting a pragmatic mixed methods systematic litera-
ture review is discussed further in Chapter 5 and can provide guidance for the evaluator at this stage.
Input from experts can be obtained by a mixed methods strategy involving focus groups and surveys.
2. Identify and describe behaviors that underlie the construct. This phase involves the analysis of the
data collected in Step 1. Data from the experts can be analyzed using qualitative strategies such as
thematic analysis, grounded theory, or ethnographic analysis. The preliminary construct elements can
be shared with experts through a quantitative strategy such as the Delphi technique. This process
can be repeated as many times as necessary to reach a clear concept of the construct.
3. Develop initial instrument. A team can contribute to the item writing using a table of specifications to
ensure coverage of the construct. The items should be formatted into a quantitative form, such as a
Likert scale, with added room for qualitative comments about each item.
4. Pilot test initial instrument. The draft instrument should be administered to a pilot group to obtain input
into the quality of the items by asking about such characteristics as clarity, relevance, and cultural
responsiveness. It is possible to calculate reliability and validity coefficients at this stage, but these
should be treated with caution because the pilot group is typically quite small. It is also possible that
the evaluator could conduct focus groups or interviews following the pilot administration to collect ad-
ditional qualitative data. The instrument should be revised based on this process.
5. Design and field test revised instrument. A larger sample is needed for field testing the instrument in
order to permit the use of factor analysis of the responses. In a mixed methods design, the evaluator
should include qualitative data collection through open-ended questions in addition to the quantitative
responses to the items.
6. Validate revised instrument. Quantitative analysis phase: In this phase, the evaluator employs the
traditional quantitative analytic strategies to establish the validity and reliability of the instrument (e.g.,
factor analysis, correlational analysis, discriminant and convergent validity).
7. Validate revised instrument. Qualitative analysis phase: The qualitative data from Stage 5 can be an-
alyzed using any number of qualitative analytic strategies (e.g., narrative analysis, content analysis,
and semiotics).
8. Validate revised instrument. Mixed analysis phase: Qualitative-dominant crossover analyses. This
involves either qualitizing quantitative data or conducting a quantitative analysis of the qualitative da-
ta. This could take the form of disaggregation of groups based on high, medium, or low scores and
comparing their qualitative data.
9. Validate revised instrument. Mixed analysis phase: Quantitative-dominant crossover analyses. This
involves quantitizing qualitative data perhaps by noting themes that emerge and transforming them
into factors that could be analyzed quantitatively.
10. Evaluate the instrument/construct evaluation process and product. The evaluator should review all
the data and findings that emerged from the nine previous steps and engage in a debriefing with
members of the team and experts. If concerns arise about the items or the process, then the instru-
ment should be revised and the steps repeated until a satisfactory version of the instrument emerges
(adapted from Onwuegbuzie et al. 2010).
Note: Additional examples of the development of quantitative instruments that used the Onweugbuzie et al.
(2010) framework include (1) Koskey, Sondergeld, Stewart, and Pugh (2016) in their development of a quan-
titative instrument to measure the transformative experiences that students in school have in the form of
extending their classroom learning into everyday experience and (2) David, Hitchcock, Ragan, Brooks, and
Starkey (2016) in their development of a quantitative instrument to measure trust between college students
and their athletic trainers.
1. The teacher accreditation organization requires that universities provide data on the change in
teacher candidates’ dispositions as they progress through their university program and in the first 5
years of their teaching careers. The measurement is to be used several times during the program and
needs to align with the conceptual framework that guides the programmatic decisions. Design a
mixed methods instrument development study using the Use branch characteristics for an instrument
that could be used to address this need.
2. Examine the development of instruments used in evaluation studies in a domain of your interest. To
what extent does the development process include the use of mixed methods and reflect the Use
branch characteristics?
Development of instruments in the Social Justice branch reflects the assumption of the transformative par-
adigm (Mertens, 2015b; Mertens & Wilson, 2012). Participation of stakeholders is a priority throughout the
entire process of development in order to incorporate diverse values and meanings based on people’s ex-
periences and understandings. Specific efforts are made to ensure that inclusion of traditionally underrepre-
sented groups is accomplished in culturally respectful ways. The transformative ontological assumption holds
that different versions of reality emerge from different social positions and that different methods are needed
to accurately understand the realities of those who are marginalized and oppressed. Mixed methods within
the Social Justice branch can “capture the diversity of people’s viewpoints with regard to their social loca-
tions” (Ungar & Liebenberg, 2011, p. 128). Cultural differences are of the utmost importance, and evaluators
in this branch ask themselves questions such as “Does a measure in one culture relate to the measure of the
same factor in another?. . . [And] how do we balance assumptions of homogeneity across Minority and Ma-
jority World contexts with the need for sensitivity to within group and between group heterogeneity?” (Ungar
& Liebenberg, 2011, p. 129).
Ungar and Liebenberg (2011) worked with a team of instrument developers across the world to develop a
measure of youth resilience (see Sample Study 3.5). They stated their goal as follows: “Our goal when con-
structing the CYRM-28 [Child Youth and Resilience Measure] was to build a more culturally sensitive measure
with face and item validity (we wanted a measure that was perceived as relevant by all our global partners
and showed the potential for discriminant validity in multiple contexts). Achieving this goal meant a more rec-
iprocal research design congruent with mixed methods as used within a transformative paradigm” (p. 129).
Problem: Many studies of youth in low- and middle-income countries focus on what is wrong with young peo-
ple. However, youth who survive in challenging contexts have strengths that, if recognized, can be seen as
part of the solution to the problems they face. A good measure of youth resilience grounded in cultures across
the world could assist in the identification of resilience.
Sample: The instrument was pilot tested with 1,451 youth, and 89 youth participated in interviews in 11 dif-
ferent countries. Sites were purposefully chosen to maximize variation. The qualitative sample of youth was
chosen to meet these criteria: “(a) cultural differences, (b) differences in the nature of the risks facing individ-
ual youth (all participants were sampled from one population of youth-at-risk identified locally, such as youth
living in poverty, exposed to violence, or racially marginalized), and (c) the ability of the principal investigator
to locate an academic partner with the capacity to supervise the research locally” (p. 131). For the quantitative
testing of the instrument, 60 or more youth were purposefully chosen by the local advisory committee who
knew youth who were placed at risk by virtue of living in a conflict zone, poverty, family breakdown, marginal-
ization, or addiction in the family.
Data collection: A team of instrument developers from 14 communities in 11 countries collaborated on the
creation of CYRM-28. Local advisory committees (LAC) were established at each site; these were composed
of five individuals who were knowledgeable about youth in their communities. The evaluators conducted an
audit of strengths that “were the most relevant to populations under stress by conducting focus groups (and
later qualitative interviews) with youth and those responsible for their well-being from Minority and Majority
World contexts where their youth are exposed to extreme adversity” (p. 128).
Data analysis: Initial listing of elements that reflected resilience were created by the members of the interna-
tional team in a three-day conference. The group analyzed the elements using qualitative strategies that re-
sulted in 32 domains (e.g., assertiveness, problem-solving ability) organized into four clusters (individual, rela-
tional, community, and cultural). Following the focus groups and comments by LAC members, the evaluators
analyzed the items to identify common and unique elements and constructed the draft instrument. The data
from the piloting of the instrument were analyzed using factor analysis. The factor analysis revealed common-
ality and redundancy across some items, allowing the evaluators to reduce the common core of questions on
the measure to 28.
Results: “As such, by having begun with exploratory qualitative data, the questions contained in the quan-
titative measure are rooted in the experiences of individuals from multiple cultures and contexts. Findings
from the analysis of additional qualitative data also informed the quantitative analysis and findings, affecting
the structure of the CYRM-28. In this way, the CYRM-28 is designed to demonstrate good content validity
within each research site in which it was piloted while still sharing enough homogeneity to make it useful for
cross-national comparisons” (Ungar & Liebenberg, 2011, p. 128). The CYRM-28 consists of 28 items that are
standard and also allows local areas to add questions that do not appear in the common list of elements in
order to reflect specific local aspects of resilience. The questions are answered using a 5-point scale ranging
from 1, Not at all, to 5, A lot.
Benefits of using MM: “Mixed methods were necessary to identify emic factors (including community values
related to resilience) relevant to young people in cultures and contexts that are underrepresented in the Mi-
nority World literature. The use of mixed methods also allowed us to compare the results of our quantitative
findings with young people’s descriptions of their experiences of complex interactions to nurture and maintain
well-being within their challenging social ecologies” (Ungar & Liebenberg, 2011, p. 128). Mixed methods also
contributed to stronger internal validity and generalizability across contexts.
Development of an instrument in the Social Justice branch follows a logic similar to that used in the develop-
ment of any instrument. However, the influence of the transformative paradigm’s assumptions leads to con-
siderations of inclusion and representation that might not surface in other evaluation branches. The following
guidance is offered:
1. Defining the problem. The meaning of the construct being measured needs to be negotiated across
cultures and in a way that is inclusive of marginalized populations within the cultures. This means that
local knowledge is needed to ascertain how individuals can be invited and supported in respectful
ways. The use of literature review is also an appropriate part of this first step; however, it needs to be
conducted with a critical eye toward whose voices are being represented in the literature.
2. Identifying the study design. Mixed methods designs allow for inclusion of qualitative methods that
encourage discussions of variability and opportunities for tolerance and celebration of differences and
commonalities. Quantitative methods of instrument construction can be used to establish psychome-
tric properties and can be paired with qualitative data collection to ascertain the meanings and ac-
ceptability of the items as interpreted by the targeted population.
3. Identifying participants. The selection of participants for the various stages of the mixed methods
study needs to be done with an awareness of the diversity within the communities. The use of local
advisory committees can be helpful in this regard if the members are engaged with the communities
and understand the power structures and bases of discrimination.
4. Construction of the measure. Evaluators can rely on the scope of the concept to be measured as it is
represented in the literature with an important caveat. They need to ensure that the items for the in-
strument are grounded in the culture and complexity of the community. A good way to accomplish this
is to have an intensive qualitative process that precedes and operates continuously throughout the
development of the items. If comparisons are to be made across groups, then attention needs to be
given to the language and translation issues.
5. Analysis and interpretation. The use of mixed methods encourages the co-construction of meaning of
the constructs and helps to refine the selection of items. Face-to-face meetings within sites and be-
tween sites help ensure the measure demonstrates high face validity across cultures (adapted from
Ungar & Liebenberg, 2011, p. 142).
1. You have been hired by a firm to develop an instrument to determine if people in a community under-
stand changes that need to be made in order to address environmental pollution and their willingness
to engage in the recommended behaviors. How would you begin this process? Design a social justice
mixed methods study to develop an instrument for this purpose.
2. Attitudes toward law enforcement officials is noted as a problem in trying to establish positive rela-
tionships between police officers and community members. What advice would you provide a com-
munity organization that wants to explore how to improve relationships with law enforcement in their
community based on a measurement instrument that you develop? Design a social justice mixed
methods study to develop an instrument for this context.
Dialectical pluralism applied to instrument development is similar to the description provided in Chapter 2
on evaluation of interventions in that two or more paradigms are used to frame the study (Shannon-Baker,
2015b). The processes used to develop the instrument stay true to the assumptions of the paradigms used in
the framing of the study and then are brought into conversation with each other to identify new understandings
that emerge from the different perspectives (Greene & Hall, 2010). Johnson and Stefurak (2013) recommend
that the qualitative and quantitative parts of the study be conducted by cross-perspective research teams that
represent expertise in each of the methods used. Through respectful dialogue among the team members, it
is possible to incorporate understandings into the instrument development that would not be available if one
method of data collection was used alone.
De-la-Cueva-Ariza et al. (2014) provide a description of a method for applying DP in the development of an
instrument to study patient satisfaction with nursing care in intensive care contexts. Their publication consists
of a description of their proposed methodology; therefore no results of the study can be presented.
Problem: Patient experience with nursing care is different in intensive care settings than in other health set-
tings. No instrument exists that is specific to the intensive care context.
Evaluand: Survey for intensive care patients to measure their satisfaction with nursing care in Spain.
Sample: The qualitative sample size is estimated to be from 27 to 36 participants who will be purposefully
chosen for maximum variation with individuals who meet the following criteria: be at least 18 years old, have
spent at least 48 hours in the intensive care unit, and are able to describe their experience in the intensive
care unit. The quantitative sample will be about 200 patients who were admitted to the intensive care unit
using the same criteria used for the qualitative stage.
Data collection: The study will begin with a qualitative stage that will include in-depth interviews with patients
who were admitted to intensive care and have spent at least 48 hours in that unit. Qualitative data collection
will include an in-depth interview, field diaries, and a discussion group with experts. Documents will also be
analyzed; the evaluators do not specify what the documents will be. The questionnaire itself will be the data
collection instrument for the quantitative part of the study. It will be developed based on the findings from the
qualitative part of the study.
Data analysis: Grounded theory will be used to analyze the qualitative data. The evaluators described their
data analytic process for the qualitative portion of the study as follows:
First, an open encoding will allow for conceptualization (identification of concepts) and finding the
properties and dimensions in the data. Therefore, we have the foundation and the initial structure
to build a theory. Second, an axial codification consisting of matching categories to subcategories
and linking them based on their properties and dimensions: accommodating the properties of a cat-
egory and its dimensions (this task begins during the open codification); identifying the variety of
conditions, actions/interactions and consequences associated with the phenomenon of satisfaction;
matching category with its subcategories by sentences indicating the relationship among them; and
searching for clues in the data denoting how they can be related to the main categories.
Finally, the selective codification in the analysis will lead to the integration of concepts around a cen-
tral category and completing the categories that need to be further perfected and developed, show-
ing the depth and complexity of thought of the theory being developed. (p. 206)
Quantitative analysis will include descriptive and inferential statistics. The psychometric properties of the in-
strument will be determined using standard Methods branch strategies. “The scale’s metric properties will be
determined by: content validity with establishment of the Pearson correlation for quantitative variables and
the sensitivity and specificity for the qualitative ones. The construct’s validity will be defined through factorial
analysis of the instrument’s items. And reliability decided by the internal coherence established by Cronbach
alpha and temporal stability through test-retest reliability calculated using the intra-class correlation coefficient
will be cross-checked” (p. 207).
Results: The qualitative data will be used to categorize items that can be included in the questionnaire. The
quantitative data will be used to establish the instrument’s psychometric properties.
Benefits of using MM: The dialectical pluralism approach to instrument development allows the evaluators
to adhere to the criteria for rigor that are defined for qualitative data collection in the Values branch, as well
as those for quantitative data collection as defined in the Methods branch. The instrument will be developed
based on the real experiences of patients in the intensive care unit and will have rigorous psychometric prop-
erties.
Instrument development using a DP approach borrows ideas from the paradigms used to frame the DP ap-
proach in each particular study.
1. If the plan for the study begins with a qualitative portion that reflects the assumptions of the construc-
tivist paradigm and the Values branch, then the evaluators need to adhere to the methodological as-
sumptions associated with that stance. This can include a comprehensive literature review. However,
in the de-la-Cueva-Ariza et al. (2014) study, the evaluators propose using a grounded theory ap-
proach in the qualitative portion of the study that does not rely on preestablished categories that might
emerge from the literature. The basic logic of grounded theory is that the inquiry begins without
preestablished categories of concepts; rather, relevant ideas are allowed to emerge from interactions
with participants in the field (Charmatz, 2014; Corbin & Strauss, 2008). Data are collected and ana-
lyzed simultaneously, and then data from each instance of interaction are compared with the analysis
of others. This provides for a refinement of tentative categories that emerge from earlier data collec-
tion and analysis through a process of constant comparison. The intent is to develop an understand-
ing of the concept to be measured that is grounded in the experiences of the target population. For
example, de-la-Cueva-Ariza et al. (2014) plan to develop a grounded theory that explains the con-
cept of patient satisfaction in relation to nursing care in intensive care contexts. The specifics of their
grounded theory analysis are found in the summary of Sample Study 3.6.
2. The quantitative phase of the instrument development is informed by the results of the qualitative
phase, thus bringing the qualitative and quantitative perspectives into dialogue with each other. Items
for the questionnaire are developed that reflect the concepts that emerged from the qualitative data
analysis. It is common practice to have these items reviewed by a group of experts in the field prior
to administering them in a draft form to a sample from the target population. A revised draft is usually
administered to a sample of the target population in a format that allows for determining the appropri-
ateness, comprehension, and interpretation of the items. The draft instrument is then revised again
based on the analysis of this administration.
3. At this stage, the instrument is generally tested quantitatively to determine its reliability and validity by
administering it to another sample and conducting factor analysis and comparison to other known in-
struments for validity and calculating reliability coefficients such as Cronbach alpha and the test-retest
approach.
1. You have been hired by a community outreach group to help them measure the impact of their pro-
grams in terms of attitudes of community members toward members of a variety of religious groups.
How would you use DP to frame an approach to the development of an instrument for this purpose?
2. If an organization says it has heard of this concept of DP but doesn’t understand exactly what that
would mean for its development of instruments, how would you explain the concept to that organiza-
tion? Think of an example you can use as part of your explanation and integrate the example into the
information you would share with the organization.
Part of the strategies for instrument development share common territory across the philosophical frames
and evaluation branches in that a literature review is generally used as a starting point and engagement with
stakeholders is necessary to collect quantitative and qualitative data about their experience with the instru-
ment. However, mixed methods evaluation designs for instrument development manifest different emphases
depending on the philosophical framing and evaluation branch used by the evaluators. For example, in the
Methods branch, the emphasis is on quantitative data to establish the reliability and validity of an instrument,
often in relation to a norm group. Qualitative data can be used to determine coverage and understanding of
items. In the Values branch, the focus is much more on the collection of qualitative data to inform the devel-
opment of the instrument. This can take the form of an ethnographic study to determine nuances of cultural
understandings. The quantitative part of the study is also used to determine reliability and validity. In the Use
branch, the goal is to develop an instrument that is psychometrically sound, and qualitative data are used to
serve that goal. As in the sample study in this chapter, the qualitative data were “accidentally” collected by
comments shared with the instrument developers. For the Social Justice branch, evaluators are more likely to
spend more time upfront collecting qualitative and quantitative data to understand the context and complexity
of the stakeholders more fully. They will also be consciously inclusive of members of marginalized populations
in the sample that participates in the instrument development. In DP, separate strategies are used for qualita-
tive data collection (which may include literature reviews and focus groups), and these are used to inform the
development, implementation, and interpretation of quantitative data collection.
In the next chapter, we look at another important venue in which evaluators work: policy evaluation. This area
is replete with complexity because of the explicitly political nature of the context and the power relations in-
herent in such contexts.
[Link]
Trustworthiness in mixed methods evaluations is assessed by criteria such as credibility, transferability, dependability, and confirmability, as proposed by Lincoln and Guba. These criteria ensure that the findings are credible, applicable in other contexts, consistent over time, and grounded in data. Trustworthiness is important because it confirms that the evaluation's findings are robust and can be trusted by diverse stakeholders for making informed decisions .
Dialectical pluralism in mixed methods instrument development involves using two or more paradigms to frame the study, allowing for diverse perspectives to inform the development process. Through respectful dialogue among cross-perspective research teams, the approach facilitates the incorporation of varied understandings into the instrument that would not be possible through a single method, thus enriching the instrument's development and applicability .
Pre-testing with a sample from the target population allows evaluators to test the appropriateness, comprehension, and interpretation of the instrument items. Feedback from this stage informs revisions that improve item clarity and relevance, thus refining the instrument to better meet the needs of its intended audience and improving its validity and reliability .
Ensuring face validity across cultures involves significant challenges such as accommodating cultural nuances in item wording, addressing language and translation issues, and guaranteeing the relevance and appropriateness of items for different cultural contexts. These challenges require a thorough qualitative process to ground the items in cultural complexity and continuous engagement with diverse cultural insights during development .
Progressive subjectivity involves evaluators engaging in critical reflection to identify and mitigate their biases throughout the study. It allows evaluators to become aware of any preconceived notions, document shifts in understanding, and develop strategies to minimize bias influence on data collection and interpretation, thus enhancing the credibility and neutrality of the findings .
In the Values branch of mixed methods design, stakeholder input is crucial as it enhances understanding of complex cultural and linguistic issues and helps increase the breadth of items to reflect the stakeholders' diverse experiences. This engagement ensures that the instrument is relevant, comprehensive, and culturally sensitive, thus improving its utility and acceptance .
A concurrent mixed methods (MM) study allows for the simultaneous collection of quantitative and qualitative data, which can be beneficial in providing a comprehensive view when developing an instrument to measure race relations. The advantages include obtaining immediate qualitative insights into how participants perceive the quantitative items, which helps in refining these items for cultural relevance and clarity. A limitation is the potential complexity in integrating both sets of data effectively to inform instrument development, as well as the substantial resources needed to manage concurrent data collection processes .
Incorporating both qualitative and quantitative perspectives in the final phase of instrument development is critical because it ensures that the instrument not only measures what it is intended to but also resonates with the target population's experiences and insights. This dual perspective helps refine the instrument, enhances its reliability and validity, and ensures it is culturally and contextually appropriate, thus improving its overall effectiveness and acceptance .
The qualitative component enhances instrument development by allowing for the investigation of issues related to the administration of the instrument, such as understanding, relevance, and comprehensiveness of the items. It helps identify problematic content and provides deeper cultural and contextual insights that can guide the refinement of the instrument, ensuring its validity and reliability across diverse populations .
Catalytic authenticity influences stakeholder engagement by stimulating action and change through exposure to evaluation findings, which can result in policy and practice transformations. Tactical authenticity empowers stakeholders, increasing their engagement by making them feel that their participation in the evaluation results in tangible outcomes, such as adoption of recommendations, thereby fostering a sense of ownership and commitment .