0% found this document useful (0 votes)
13 views10 pages

Guidelines for Effective Questionnaire Development

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views10 pages

Guidelines for Effective Questionnaire Development

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Reflections on Research

Practical Guidelines to Develop and Evaluate a Questionnaire

Abstract Kamal Kishore,


Life expectancy is gradually increasing due to continuously improving medical and nonmedical Vidushi Jaswal1,
interventions. The increasing life expectancy is desirable but brings in issues such as impairment of Vinay Kulkarni2,
quality of life, disease perception, cognitive health, and mental health. Thus, questionnaire building
and data collection through the questionnaires have become an active area of research. However, Dipankar De3
questionnaire development can be challenging and suboptimal in the absence of careful planning Departments of Biostatistics and
and user‑friendly literature guide. Keeping in mind the intricacies of constructing a questionnaire,
3
Dermatology, Post Graduate
Institute of Medical Education
researchers need to carefully plan, document, and follow systematic steps to build a reliable and and Research (PGIMER),
valid questionnaire. Additionally, questionnaire development is technical, jargon‑filled, and is not a 1
Department of Psychology,
part of most of the graduate and postgraduate training. Therefore, this article is an attempt to initiate MCM DAV College for Women,
an understanding of the complexities of the questionnaire fundamentals, technical challenges, and Chandigarh, 2Department of
sequential flow of steps to build a reliable and valid questionnaire. Dermatology, PRAYAS Health
Group, Amrita Clinic, Karve
Keywords: Instrument, psychometrics, questionnaire development, reliability, scale construction, Road, Pune, Maharashtra, India
validity

Introduction are measured with a group of items framed


from the “easiest” to the “most difficult.”
There is an increase in the usage of the
For example, for a stem, a participant may
questionnaires to understand and measure
have to choose from options (a) stand,
patients’ perception of medical and (b) walk, (c) jog, and (d) run. It requires
nonmedical care. Recently, with increased a strict ordering of items. The Rasch
interest in quality of life associated with method adds the stochastic component
chronic diseases, there is a surge in to the Guttman method which lay the
the usage and types of questionnaires. foundation of modern and powerful
The questionnaires are also known as technique item response theory for scale
scales and instruments. Their significant construction. All the approaches have their
advantage is that they capture information fair share of advantages and disadvantages.
about unobservable characteristics such However, Likert scales based on classical
as attitude, belief, intention, or behavior. testing theory are widely established and
The multiple items measuring specific preferred by researchers to capture intrinsic
domains of interest are required to characteristics. Therefore, in this article, we
obtain hidden (latent) information from will discuss only psychometric properties
participants. However, the importance of required to build a Likert scale. Address for correspondence:
questions or items needs to be validated Dr. Dipankar De,
A hallmark of scientific research is that it Additional Professor,
and evaluated individually and holistically.
needs to meet rigorous scientific standards. Department of Dermatology,
The item formulation is an integral part A questionnaire evaluates characteristics Post Graduate Institute
of Medical Education
of the scale construction. The literature whose value can significantly change and Research (PGIMER),
consists of many approaches, such as with time, place, and person. The error Chandigarh, India.
Thurstone, Rasch, Gutmann, or Likert variance, along with systematic variation, E‑mail: dr_dipankar_de@
methods for framing an item. The Thurstone plays a significant part in ascertaining [Link]
scale is labor intensive, time‑consuming, unobservable characteristics. Therefore, it
and is practically not better than the is critical to evaluate the instruments testing
Access this article online
Likert scale.[1] In the Guttman method, human traits rigorously. Such evaluations
cumulative attributes of the respondents are known as psychometric evaluations in Website: [Link]

context to questionnaire development and DOI: 10.4103/idoj.IDOJ_674_20


Quick Response Code:
This is an open access journal, and articles are
distributed under the terms of the Creative Commons How to cite this article: Kishore K, Jaswal V,
Attribution‑NonCommercial‑ShareAlike 4.0 License, which Kulkarni V, De D. Practical guidelines to develop and
allows others to remix, tweak, and build upon the work evaluate a questionnaire. Indian Dermatol Online J
non‑commercially, as long as appropriate credit is given and the 2021;12:266-75.
new creations are licensed under the identical terms.
Received: 21-Aug-2020. Revised: 11-Dec-2020.
For reprints contact: WKHLRPMedknow_reprints@[Link] Accepted: 25-Jan-2021. Published: 02-Mar-2021.

266 © 2021 Indian Dermatology Online Journal | Published by Wolters Kluwer - Medknow
Kishore, et al.: Developing a questionnaire

validation. The scientific standards are available to select specific essential requirements which are not directly
items, subscales, and entire scales. The researchers can a part of scale development and evaluation; however,
broadly segment scientific criteria for a questionnaire into
reliability and validity.
Questionnaire
Development  What to do?
Despite increasing usage, many academicians grossly Purpose

misunderstand the scales. The other complication is that Step-1 Item generation
 Literature review
 Expert and nonexpert
many authors in the past did not adhere to the rigorous  Existing questionnaire
standards. Thus, the questionnaire‑based research was Step-2 Item Formatting
 Unambiguous
criticized by many in the past for being a soft science.[2] Step-3 Preliminary
 Unloaded
 Unbarreled
The scale construction is also not a part of most of the Questionnaire
 Covering letter
graduate and postgraduate training. Given the previous Step-4 Validation  Order of items
discussion, the primary objective of this article is to  Questionnaire layout
Step-5
sensitize researchers about the various intricacies and Pilot Testing  Coefficient validity ratio
 Coefficient validity index
importance of each step for scale construction. The Step-6 Data Collection  Interrater agreement
emphasis is also to make researcher aware and motivate  Floor & ceiling effect
to use multiple metrics to assess psychometric properties. Step-7 Evaluation
 Rewording
 Deletion
Table 1 describes a glossary of essential terminologies used
 Data collection
in context to questionnaire.  Data Entry
 Data cleaning
The process of building a questionnaire starts with item
 Descriptive analysis
generation, followed by questionnaire development, and  Reliability
concludes with rigorous scientific evaluation. Figure 1  Validity

summarizes the systematic steps and respective tasks Figure 1: Flowchart demonstrating the various steps involved in the
at each stage to build a good questionnaire. There are development of a questionnaire

Table 1: Glossary of important terms used in context to psychometric scale


Term Definition
Psychometrics A science which deals with the quantitative assessment of abilities that are not directly observable, e.g.,
confidence, intelligence
Reliability Refer to the degree of consistency of instrument in measurements, e.g., is weighing machine giving similar results
under consistent conditions?
Validity Refer to the ability of an instrument to represent the intended measure correctly, e.g., is weighing machine giving
accurate results?
Likert scale A psychometric scale consists of multiple items that arrived through a systematic evaluation of reliability and
validity, e.g., quality‑of‑life score
Likert Item It is a statement with a fixed set of choices to express an opinion with the level of agreement or disagreement
Latent variable Represent a concept or underlying construct which cannot be measured directly. Latent variables are also known
as unobserved variables, e.g., health and socioeconomic status
Manifest variable A variable which can be measured directly. Manifest variables are also known as observed variables, e.g., blood
pressure and income
Double‑barrel item A question addressing two or more separate issues but provides an option for one answer, e.g., do you like the
house and locality?
Negative item It is an item which is in the opposite direction from most of the questions on a scale
Factor loadings Demonstrate the correlation coefficient between the observed variable and factor. It quantifies the strength of
the relationship between a latent variable (factor) and manifest variables. It is key to understand the relative
importance of items in the final questionnaire. An item with high factor loading is more important than others
Cross‑loading An observed variable with loading more than threshold value on two or more factors, e.g., education level with
value >0.35 for both teaching and research domains. The items with cross‑loadings are candidates for deletion
from a questionnaire
Reverse scoring The practice of reversing the score to cancel positive and negative loading on the same factor, e.g., changing the
maximum rating (such as strongly agree=5) to a minimum (such as strongly agree=1) or vice versa
Floor and ceiling The inability of a scale to discriminate between participants in a study as the high proportion of participants score
effect worst/minimum or best/maximum score, e.g., more than 80% responses are received by single option among the
five options for a Likert item. Item is poorly discriminating between participants and is a candidate for deletion
Eigenvalue An indicator of the amount of variance explained by a factor. The factor with the highest eigenvalue explains
the maximum amount of variance and practically makes a factor most important. The eigenvalue is obtained by
column sum of squares of factor loading

Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021 267


Kishore, et al.: Developing a questionnaire

these improve the utility of the instrument. The indirect advantageous to speak with experts telephonically, face
but necessary conditions are documented and discussed to face, or electronically, requesting their participation
under the miscellaneous category. We broadly segment before mailing the questionnaire. It is good to explain to
and discuss the questionnaire development process under them right in the beginning that this process unfolds over
three domains, known as questionnaire development, phases. The time allowed to respond can vary from hours
questionnaire evaluation, and miscellaneous properties. to weeks. It is recommended to give at least 7 days to
respond. However, a nonresponse needs to be followed
Questionnaire Development up by a reminder email or call. Usually, this stage takes
The development of the list of items is an essential and two to three rounds. Therefore, it is essential to engage
mandatory prerequisite for developing a good questionnaire. with experts regularly; else there is a risk of nonresponse
The researcher at this stage decides to utilize formats such from the study. Table 2 gives general advice to researchers
as Guttman, Rasch, or Likert to frame items.[2] Further, for making a cover letter. The researcher can modify the
the researcher carefully identifies the appropriate member cover letter appropriately for their studies. The authors can
of the expert panel group for face and content validity. consult Rubio and coauthors for more details regarding the
Broadly, there are six steps in the scale development. drafting of a cover letter.[4]

Step I Step IV
It is crucial to select appropriate questions (items) to The responses from each round will help in rewording,
capture the latent trait. An exhaustive list of items is the rephrasing, and reordering of the items in the scale. Few
most critical and primary requisite to lay the foundation questions may need deletion in the different rounds of
of a good questionnaire. It needs considerable work in previous steps. Therefore, it is better to evaluate content
terms of literature search, qualitative study, discussion with validity ratio (CVR), content validity index (CVI), and
colleagues, other experts, general and targeted responders, interrater agreement before deleting any question in the
and other questionnaires in and around the area of interest. instrument. Readers can consult formulae in Table 2 for
General and targeted participants can also advise on items, calculating CVR and CVI for the instrument. CVR is
wording, and smoothness of questionnaire as they will be calculated and reported for the overall scale, whereas
the potential responders. CVI is computed for each item. Researchers need to
consult Lawshe table to determine the cutoff value for
Step II CVR as the same depends on the number of experts in
It is crucial to arrange and reword the pool of questions the panel.[5] CVI >0.80 is recommended. Researchers
for eliminating ambiguity, technical jargon, and loading. interested in detail regarding CVR and CVI can read
Further, one should avoid using double‑barreled, excellent articles written by Zamanzadeh et al. and
long, and negatively worded questions. Arrange all Rubio et al.[4,6] It is crucial to compute CVR, CVI,
items systematically to form a preliminary draft of the and kappa agreement for each item from the rating of
questionnaire. After generating an initial draft, review importance, representativeness, and clarity by experts.
the instrument for the flow of items, face validity The CVR and CVI do not account for a chance factor.
and content validity before sending it to experts. The Since interrater agreement (IRA) incorporates chance
researcher needs to assess whether the items in the factor; it is better to report CVR, CVI, and IRA
score are comprehensive (content validity) and appear to measures.
measure what it is supposed to measure (face validity).
Step V
For example, does the scale measuring stress is measuring
stress or is it measuring depression instead? There is The scholars require to address subtle issues before
no uniformity on the selection of a panel of experts. administering a questionnaire to responders for pilot
However, a general agreement is to use anywhere from a testing. The introduction and format of the scale play a
minimum of 5–15 experts in a group.[3] These experts will crucial role in mitigating doubts and maximizing response.
ascertain the face and content validity of the questionnaire. The front page of the questionnaire provides an overview
These are subjective and objective measures of validity, of the research without using technical words. Further,
respectively. it includes roles and responsibilities of the participants,
contact details of researchers, list of research ethics
Step III (such as voluntary participation, confidentiality and
It is advisable to prepare an appealing, jargon‑free, and withdrawal, risks and benefits), and informed consent for
nontechnical cover letter explaining the purpose and participation in the study. It is also better to incorporate
description of the instrument. Further, it is better to include anchors (levels of Likert item) in each page at the top or
the reason/s for selecting the expert, scoring format, and bottom or both for ease and maximizing response. Readers
explanations of response categories for the scale. It is can refer to Table 3 for detail.

268 Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021


Kishore, et al.: Developing a questionnaire

Table 2: General overview and the instructions for rating in the cover letter to be accompanied by the questionnaire
Content Explanation
Construct Definition of characteristics of the measurement
Purpose To evaluate the content and face validity
How Please rate each item for its representativeness and clarity on a scale from 1 to 4
Evaluate the comprehensiveness of the entire instrument in measuring the domain
Please add, delete, or modify any item as per your understanding
Measure CVR CVI
Characteristics Importance Representative Clarity
Scoring 0-Not necessary 1-Not representative 1-Not clear
1-Useful 2-Need major revisions to be representative 2-Need major revisions to be clear
2-Essential 3-Need minor revisions to be representative 3-Need minor revisions to be clear
4-Representative 4-Clear
Formula CVR = (NE-N/2)/(N/2) CVIR=NR/N CVIC=NC/N
where NE=number of experts where CVIR=CVI for representativeness where CVIC=CVI for clarity
rated an item as essential NR=Number of experts rated an item as NC=Number of experts rated an
N=Total number of experts representative (3 or 4) item as clear (3 or 4)
N=Total number of experts N=Total number of experts

Table 3: A random set of questions with anchors at the top and bottom row
Items Strongly disagree Disagree Neutral Agree Strongly agree
(SD) (D) (N) (A) (SA)
Duration of disease (since onset) SD D N A SA
Number of relapse(s) of the disease SD D N A SA
Duration of oral erosions (present episode) SD D N A SA
Number of relapse(s) of oral lesions SD D N A SA
Persistence of oral lesions after subsidence of cutaneous lesions SD D N A SA
Change in size of existing lesion in last 1 week SD D N A SA
Development of new lesions in last 1 week SD D N A SA
Difficulty in eating normal food SD D N A SA
Difficulty in eating food according to their consistency SD D N A SA
Inability to eat spicy food SD D N A SA
Inability to drink fruit juices SD D N A SA
Excessive salivation/drooling SD D N A SA
Difficulty in speaking SD D N A SA
Difficulty in brushing teeth SD D N A SA
Difficulty in swallowing SD D N A SA
Restricted mouth opening SD D N A SA
Strongly disagree Disagree Neutral Agree Strongly agree

Step VI 0.7, respectively, for IQC and reliability, are suspicious


and candidate for elimination from the questionnaire.
Pilot testing of an instrument in the target population is an Cronbach’s α, a measure of internal consistency and IQC
important and essential requirement before testing on a large of a scale, indicates researcher about the quality of items in
sample of individuals. It helps in the elimination or revision measuring latent attribute at the initial stage. This process
of poorly worded items. At this stage, it is better to use is important to refine and finalize the questionnaire before
floor and ceiling effects to eliminate poorly discriminating starting the testing of a questionnaire in study participants.
items. Further, random interviews of 5–10 participants can
help to mitigate the problems such as difficulty, relevance, Questionnaire Evaluation
confusion, and order of the questions before testing it on The preliminary items and the questionnaire until this stage
the study population. The general recommendations are to have addressed issues of reliability, validity, and overall
recruit a sample size between 30 and 100 for pilot testing.[4] appeal in the target population. However, researchers need
Inter‑question (item) correlation (IQC) and Cronbach’s α to rigorously evaluate the psychometric properties of the
can be assessed at this stage. The values less than 0.3 and primary instrument before finally adopting. The first step

Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021 269


Kishore, et al.: Developing a questionnaire

in this process is to calculate the appropriate sample size Table 4: A sample of data entry format
for administering a preliminary questionnaire in the target (a) Illustration of master sheet
group. The evaluations of various measures do not follow a Participant Age Religion Family Height Weight Q1 Q2 Q3
sequential order like the previous stage. Nevertheless, these 1 25 1 1 185.0 85.0 1 5 2
measures are critical to evaluate the reliability and validity 2 26 3 1 155.0 63.0 2 5 1
of the questionnaire. 3 22 2 2 155.0 57.0 4 2 1
4 35 2 1 158.5 67.5 3 2 2
Data entry
5 49 1 2 175.0 64.0 2 4 3
Correct data entry is the first requirement to evaluate the 6 40 4 1 159.0 78.0 2 4 3
characteristics of a manually administered questionnaire. Qi→ith Question in the questionnaire, where i=1,2,3, … n
The primary need is to enter the data into an appropriate (b) Illustration of coding sheet
spreadsheet. Subsequently, clean the data for cosmetic and Variable Description Coding and Measurement
logical errors. Finally, prepare a master sheet, and data label valid range scale
dictionary for analysis and reference to coding, respectively. Participant A random None String
Authors interested in more detail can read “Biostatistics serial number
Series.”[7,8] The data entry process of the questionnaire is to participant
like other cross‑sectional study designs. The rows and Age Age in years None Interval
columns represent participants and variables, respectively. (30‑70 years)
It is better to enter the set of items with item numbers. Religion Religion of 1=Hindu Nominal
the participant 2=Sikh
First, it is tedious and time‑consuming to find suitable
variable names for many questions. Second, item numbers 3=Muslim
help in quick identification of significantly contributing and 4=Others
non‑contributing items of the scale during the assessment Q Level of 1=Strongly Ordinal
of psychometric properties. Readers can see Table 4 for agreement in disagree
more detail. the question 2=Disagree
3=Neutral
Descriptive statistics
4=Agree
Spreadsheets are easy and flexible for routine data entry and 5=Strongly
cleaning. However, the same lack the features of advanced agree
statistical analysis. Therefore, the master sheet needs to be
exported to appropriate software for advanced statistical
type of missingness. The typically recommended threshold
analysis. Descriptive analysis is the usual first step which
for the missingness is 5%.[10] There are broadly three types
helps in understanding the fundamental characteristics of
the data. Thus, report appropriate descriptive measures of missingness, such as missing completely at random,
such as mean and standard deviation, and median and missing at random, and not missing at random. After
interquartile/interdecile range for continuous symmetric identification of a missing mechanism, impute the data with
and asymmetric data, respectively.[9] Utilize exploratory single or multiple imputation approaches. Readers can refer
tabular and graphical display to inspect the distribution to an excellent article written by Graham for more details
of various items in the questionnaire. A stacked bar chart about missing data.[11]
is a handy tool to investigate the distribution of data Sample size
graphically. Further, ascertain linearity and lack of extreme
multicollinearity at this stage. Any value of IQC >0.7 The optimum sample size is a vital requisite to build a
warrants further inspection for deletion or modification. good questionnaire. There are many guidelines in the
Help from a good biostatistician is of great assistance for literature regarding recruiting an appropriate sample size.
data analysis and reporting. Literature broadly segments sample size approaches into
three domains known as subject to variables ratio (SVR),
Missing data analysis minimum sample size, and factor loadings (FL). The factor
Missing data is the rule, not the exception. Majority of analysis (FA) is a crucial component of questionnaire
the researchers face difficulties of finding missing values designing. Therefore, recent recommendations are to
in the data. There are usually three approaches to analyze use FLs to determine sample size. Readers can consult
incomplete data. The first approach is to “take all” which Table 5 for sample size recommendations under various
use all the available data for analysis. In the second domains. Interested readers can refer to Beavers and
method, the analyst deletes the participants and variables colleagues for more detail.[12] The stability of the factors
with gross missingness or both from the analysis process. is essential to determine sample size. Therefore, data
The third scenario consists of estimating the percentage and analysis from questionnaires validates the sample size

270 Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021


Kishore, et al.: Developing a questionnaire

Table 5: Sample size recommendations in the literature


Sample size criteria
Subject to variables ratio Minimum sample size Factor loading
Minimum 100 participants + SVR ≥5 At least 300 participants At least 4 items with FL >0.60 (minimum 100 participants)
51 participants + number of variables At least 200 participants At least 10 items with FL >0.40 (minimum 150 participants)
At least SVR >5 At least 150‑300 participants Items with 0.30 ≤ FL ≤0.40 (minimum 300 participants)
SVR→Subject to variable ratio, FL→Factor loading

after data collection. The Kaiser–Meyer–Olkin (KMO) sample size adequacy, IQC, and Bartlett’s test in step 7
criterion testing the adequacy of sample size is available in [Figure 1]. The value of EFA is used at the initial stages
the majority of the statistical software packages. A higher to extract factors while constructing a questionnaire. It is
value of KMO is an indicator of sufficient sample size for especially important to identify an adequate number of
stable factor solution. factors for building a decent scale. The factors represent
latent variables that explain variance in the observed data.
Correlation measures First and the last factor explain maximum and minimum
The strength of relationships between the items is an variance, respectively. There are multiple factor selection
imperative requisite for a stable factor solution. Therefore, criteria, each with its advantages and disadvantages. It
the correlation matrix is calculated and ascertained for is better to utilize more than one approach for retaining
same. There are various recommendations of correlation factors during the initial extraction phase. Readers can
coefficient; however, a value greater than 0.3 is a must.[13] consult Sindhuja et al. for the practical application of more
A lower value of the correlation coefficient will fail to form than one‑factor selection criteria.[14]
a stable factor due to lack of commonality. The determinant
and Bartlett’s test of sphericity can be used to ascertain the
Kaiser’s criterion
stability of the factors. The determinant is a single value Kaiser’s criterion is one of the most popular factor retention
which ranges from zero to one. A nonzero determinant criteria. The basis of the Kaiser criterion is to explain the
indicates that factors are possible. However, it is small in variance through the eigenvalue approach. A factor with
most of the studies and not easy to interpret. Therefore, more than one eigenvalue is the candidate for retention.[15]
Bartlett’s test of sphericity is routinely used to infer that An eigenvalue bigger than one simply means that a single
determinant is significantly different than zero. factor is explaining variance for more than one observed
variable. However, there is a dearth of scientifically
Validity rigorous studies to declare a cutoff value for Kaiser’s
Physical quantities such as height and weight are observable criterion. Many authors highlighted that the Kaiser criterion
and measurable with instruments. However, many tools over‑extract and under‑extract factors.[16,17] Therefore,
need regular calibration to be precise and accurate. The investigators need to calculate and consider other measures
standardization in context to the questionnaire development for extraction of factors.
is known as reliability and validity. The validity is the
Cattell’s scree plot
property which indicates that an instrument is measuring
what it is supposed to measure. Validation is a continuous Cattell’s scree plot is another widespread eigenvalue‑based
process which begins with the identification of domains and factor selection criterion used by researchers. It is popularly
goes on till generalization. There are various measures to known as scree plot. The scree plot assigns the eigenvalues
establish the validity of the instrument. Authors can consult on the y‑axis against the number of factors in the x‑axis.
Table 6 for different types of validity and their metrics. The factors with highest to lowest eigenvalues are plotted
from left to right on the x‑axis. Usually, the scree plots
Exploratory FA form an elbow which indicates the cutoff point for factor
FA assumes that there are underlying constructs (factors) extraction. The location or the bend at which the curve first
which cannot be measured directly. Therefore, the begins to straighten out indicates the maximum number of
investigator collects the exhaustive list of observed factors to retain. A significant disadvantage of the scree
variables or responses representing underlying constructs. plot is the subjectivity of the researcher’s perception of the
Researchers expect that variables or questions in the “elbow” in the plot. Researchers can see Figure 2 for detail.
questionnaire correlate among themselves and load on the
Percentage of variance
corresponding but a small number of factors. FA can be
broadly segmented in exploratory factor analysis (EFA) The variance extraction criterion is another criterion to
and confirmatory factor analysis. The EFA is applied on retain the number of factors. The literature recommendation
the master sheet after assessing descriptive statistics such varies from more than a minimum of 50–70%
as tabular and graphical display, missing mechanism, onward.[12] However, both the number of items and factors

Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021 271


Kishore, et al.: Developing a questionnaire

Table 6: Scientific standards to evaluate and report for constructing a good scale
Psychometric Component Definition Indices
properties
Validity Content validity The items are addressing all the relevant aspect of construct Content validity ratio
Content validity indices
Interrater agreement
Face validity The test appears to measure the intended measure Expert opinion (qualitative)
Construct validity The strong (rs) and weak (rw) correlation between same and Exploratory factor analysis
different construct, respectively Correlation coefficient
Criterion validity The correlation between a predictor measure (teamwork) Correlation coefficient
and criterion measures (actual performance in team)
Convergent The correlation between a scale and conceptually similar Correlation coefficient
validity scales or subscales of a scale Multitrait‑multimethod matrix
Reliability Internal The cohesiveness of items in measuring the same variable Coefficient α
consistency consistently Coefficient β
Coefficient Ω
Test‑retest Consistency of score for stable characteristics on separate Correlation coefficient
times Intra‑class correlation coefficient
Alternate forms Consistency of scores among the same sample for similar Correlation coefficient
tests
Descriptive Tabular display Display of essential data characteristics in rows and Mean (SD)
analysis columns Median (IQR)
Graphical display Visual display of large data to exhibit trends, patterns, and Box plot
relationships Bar graph
Missing MCAR Missing data is independent of observed or unobserved data Little’s MCAR
mechanism MAR Missing data is related to observed but not unobserved data Listing and Schlittgen (LS) test
NMAR Missing data is related to unobserved data No standard test (based on
assumptions)
Factorability Sample size
Minimum number of participants required to measure study KMO criteria
outcomes
Correlation A matrix displaying the inter‑correlations among the Determinant
matrix variables
Sphericity Refers to equality of correlations between different items Bartlett’s test
MCAR: Missing completely at random; MAR: Missing at random; NMAR: Not missing at random; KMO: Kaiser‑Meyer‑Olkin; SD: Standard
deviation; IQR: Interquartile range

8 preferred; however, there are recommendations to use a


value higher than 0.30.[3,15,18]
7

6 Very simple structure


Eigen value

5 Very simple structure (VSS) approach is a symbiosis of


4 theory, psychometrics, and statistical analysis. The VSS
3
criterion compares the fit of the simplified model to the
original correlations. It plots the goodness‑of‑fit value
2
as a function of several factors rather than statistical
1 significance. The number of factors that maximizes the
0 VSS criterion suggests the optimal number of factors to
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
Factor Number extract. VSS criterion facilitates comparison of a different
number of factors for varying complexity. VSS will be
Figure 2: A hypothetical example showing the researcher’s dilemma of
selecting 6, 10, or 15 factors through scree plot highest at the optimum number of factors.[19] However, it is
not efficient for factorially complex data.
will increase dramatically if there are a large number of
Parallel analysis
manifest (observed) variables. Practically, the percentage of
variance explained mechanism should be used judiciously Parallel analysis (PA) is a statistical theory‑based robust
along with FL. The FLs with greater than 0.4 value are technique to identify the appropriate number of factors. It

272 Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021


Kishore, et al.: Developing a questionnaire

is the only technique which accounts for the probability correlation >0.70 represents high reliability.[21] The change
that a factor is due to chance. PA simulates data to generate in study condition (recovery of patients after intervention)
95th percentile cutoff line on a scree plot restricted upon the over time can decrease test–retest reliability. Therefore, it
number of items and sample size in original data. The factors is important to report the time between repeated measures
above the cutoff line are not due to chance. PA is the most while reporting test–retest reliability.
robust empirical technique to retain the appropriate number
of factors.[16,20] However, it should be used cautiously for Parallel forms and split‑half reliability
the eigenvalue near the 95th percentile cutoff line. PA is Parallel form reliability is also known as an alternate form
also robust to distributional assumptions of the data. Since of consistency. There are two types of option to report
different techniques have their fair share of advantages and parallel form reliability. In the first method, different
disadvantages, researchers need to assess information on but similar items make alternative forms of the test. The
the basis of multiple criteria. assumptions of both the assessment are that they measure
the same phenomenon or underlying construct. It addresses
Reliability the twin issues of time and knowledge acquisition of test in
Reliability, an essential requisite of a scale, is also known as test–retest reliability. In the second approach, the researcher
reproducibility, repeatability, and consistency. It identifies randomly divides the total items of an instrument into two
that the instrument is consistently measuring the attribute halves. The calculation of parallel form from two halves is
under identical conditions. Reliability is a necessary known as split‑half reliability. However, randomly divided
characteristic of a tool. The trustworthiness of a scale can half may not be similar. The parallel from and split‑half
be increased by increasing and decreasing the systematic reliability are reported with the correlation coefficient. The
and random component, respectively. The reliability of an recommendations are to use a value higher than 0.80 to
instrument can be further segmented and measured with assess the alternate form of consistency.[24] It is challenging
various indices. Reliability is important but it is secondary to generate two types of tests in clinical studies. Therefore,
to validity. Therefore, it is ideal to calculate and report researchers rarely report reliability from two analogous but
reliability after validity. However, there are no hard and separate tests.
fast rules except that both are necessary and important
measures. Readers may consult Table 6 for multiple types General Questionnaire Properties
of indices for reliability. The major issues regarding the reliability and validity of
scale development have already been discussed. However,
Internal consistency there are many other subtle issues for developing a good
Cronbach’s alpha (α), also known as α‑coefficient, is one questionnaire. These delicate issues may vary from a choice
of the most used statistics to report internal consistency of Likert items, length of the instrument, cover letter, web
reliability. The internal consistency using the interitem or internet mode of data collection, and weighting of
correlations suggests the cohesiveness of items in a scale. The immediately preceding issues demand careful
questionnaire. However, the α‑coefficient is sample‑specific; deliberation and attention from the researcher. Therefore,
thus, the literature recommends the same to calculate and the researcher should carefully think through all these
report for all the studies. Ideally, a value of α >0.70 is issues to build a good questionnaire.
preferred; however, the value of α >0.60 is also accepted
for construction of new scale.[21,22] Researchers can increase Likert items
the α‑coefficient by adding items in the scale. However, a The Likert items are the fixed choice ordinal items which
value can either reduce with the addition of non‑correlated capture attitude, belief, and various other latent domains.
items or deletion of correlated items. Corrected interitem The subsequent step is to rank the questions of the Likert
correlation is another popular measure to report for internal scale for further analysis. The numerals for ranking can
consistency. A value of α <0.3 indicates the presence of either start from 0 or 1. It does not make a difference. The
nonrelated items. The studies claim that coefficient beta (β) Likert scale is primarily bipolar as opposite ends endorse
and omega (Ω) are better indices than coefficient‑α, but the contrary idea.[2] These are the type of items which
there is a scarcity of literature reporting these indices.[23] express opinions on a measure from strong disagreement
to strong agreement. The adjectival scales are unipolar
Test–retest scale that tends to measure variables like pain intensity
Test–retest reliability measures the stability of an instrument (no pain/mild pain/moderate pain/severe pain) in one
over time. In other words, it measures the consistency direction. However, the Likert scale (most likely–least
of scores over time. However, the appropriate time likely) can measure almost any attribute. The Likert scale
between repeated measures is a debatable issue. Pearson’s can either have odd or even categories; however, odd
product‑moment and intraclass correlation coefficient categories are more popular. The number of classifications
measure and report test–retest reliability. A high value of in the Likert scale can vary from anywhere between 3 and

Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021 273


Kishore, et al.: Developing a questionnaire

11,[2] although the scale with 5 and 7 classes have displayed lengthy questionnaire. There are different opinions about
better statistical properties for discriminating between either grouping or mixing the issues in an instrument.[24]
responses.[2,24] Grouping inflates intra‑scale correlation, whereas mixing
inflates inter‑scale correlation.[28] Both the approaches
Length of questionnaire have empirically shown to give similar results for at least
A good questionnaire needs to include many items to 20 or more items. The questions related to a particular
capture the construct of interest. Therefore, investigators domain can be assigned either equal or unequal weights.
need to collect as many questions as possible. However, the There are two mechanisms to assign unequal weights in
lengthier scale increases both time and cost. The response a questionnaire. In the first situation, researchers affix
rate also decreases with an increase in the length of the different importance to items. In the second method, the
questionnaire.[25] Although what is lengthy is debatable investigators frame more or fewer questions as per the
and varies from more than 4 pages to 12 pages in various importance of subscales in the scale.
studies,[26] the longer scales increase the false positivity
rate.[27] Conclusion
The fundamental triad of science is accuracy, precision,
Translating a questionnaire
and objectivity. The increasing usage of questionnaires in
Many a time, there are already existing reliable and valid medical sciences requires rigorous scientific evaluations
questionnaires for use. However, the expert needs to assess before finally adopting it for routine use. There are
two immediate and important criteria of cultural sensitivity no standard guidelines for questionnaire development,
and language of the scale. Many sensitive questions on evaluation, and reporting in contrast to guidelines such
sexual preferences, political orientations, societal structure, as CONSORT, PRISMA, and STROBE for treatment
and religion may be open for discussion in certain societies, development, evaluation, and reporting. In this article,
religions, and cultures, whereas the same may be taboo or we emphasize on the systematic and structured approach
receive misreporting in others. The sensitive questions need for building a good questionnaire. Failure to meet the
to be reframed considering regional sentiments and culture questionnaire development standards may lead to biased,
in mind. Further, a questionnaire in different language unreliable, and inaccurate study finding. Therefore, the
needs to be translated by a minimum of two independent general guidelines given in this article can be used to
bilingual translators. Similarly, the translated questionnaire develop and validate an instrument before routine use.
needs to be translated back into the original language by a
minimum of two independent and different bilingual experts Financial support and sponsorship
who converted the original questionnaire. The process Nil.
of converting the original questionnaire to the targeted
language and then back to the original language is known Conflicts of interest
as forward and backward translation. The subsequent steps There are no conflicts of interest.
such as expert panel group, pilot testing, reliability, and
validity for translating a questionnaire remain the same as References
in constructing a new scale. 1. Streiner DL, Norman GR, Cairney J. Health Measurement
Scales: A Practical Guide to their Development and Use. USA:
Web‑based or paper‑based Oxford University Press; 2015.
Broadly, paper and electronic format are the two modes 2. Chapple ILC. Questionnaire research: An easy option? Br Dent
of administering a questionnaire to the participants. Both J 2003;195:359.
techniques have advantages and disadvantages. The 3. Boateng GO, Neilands TB, Frongillo EA, Melgar‑Quiñonez HR,
Young SL. Best practices for developing and validating scales
response rate is a significant issue in self‑administered
for health, social, and behavioral research: A primer. Front Public
scales. The significant benefits of electronic format are the Health 2018;6:149.
reduction in cost, time, and data cleaning requirements. 4. Rubio DM, Berg‑Weger M, Tebb SS, Lee ES, Rauch S.
In contrast, paper‑based administration of questionnaire Objectifying content validity: Conducting a content validity
increases external generalization, paper feel, and no need study in social work research. Soc Work Res 2003;27:94–104.
of internet. As per Greenlaw and Welty, the response 5. Lawshe CH. A quantitative approach to content validity. Pers
rate improves with the availability of both the options to Pschol 1975;28:563‑75.
participants. However, cost and time increase in comparison 6. Zamanzadeh V, Ghahramanian A, Rassouli M, Abbaszadeh A,
Alavi‑Majd H, Nikanfar AR. Design and implementation content
to the usage of electronic format alone.[27]
validity study: Development of an instrument for measuring
Item order and weights patient‑centered communication. J Caring Sci 2015;4:165‑78.
7. Kishore K, Kapoor R. Statistics corner: Structured data entry.
There are multiple ways to order an item in a questionnaire. J Postgr Med Educ Res 2019;53:94–7.
The order of questions becomes more critical for a 8. Kishore K, Kapoor R, Singh A. Statistics corner: Data

274 Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021


Kishore, et al.: Developing a questionnaire

cleaning‑I. J Postgrad Med Educ Res 2019;53:130–2. 19. Dinno A. Exploring the sensitivity of Horn’s parallel analysis to
9. Kishore K, Kapoor R. Statistics corner: Reporting descriptive the distributional form of random data. Multivariate Behav Res
statistics. J Postgrad Med Educ Res 2020;54:66–8. 2009;44:362–88.
10. Jakobsen JC, Gluud C, Wetterslev J, Winkel P. When and how 20. DeVon HA, Block ME, Moyle‑Wright P, Ernst DM, Hayden SJ,
should multiple imputation be used for handling missing data Lazzara DJ, et al. A psychometric toolbox for testing validity
in randomised clinical trials–A practical guide with flowcharts. and reliability. J Nurs Scholarsh 2007;39:155–64.
BMC Med Res Methodol 2017;17:162. 21. Straub D, Boudreau MC, Gefen D. Validation guidelines for IS
11. Graham JW. Missing data analysis: Making it work in the real positivist research. Commun Assoc Inf Syst 2004;13:24.
world. Annu Rev Psychol 2009;60:549–76. 22. Revelle W, Zinbarg RE. Coefficients alpha, beta, omega, and the
12. Beavers AS, Lounsbury JW, Richards JK, Huck SW. Practical glb: Comments on Sijtsma. Psychometrika 2009;74:145.
considerations for using exploratory factor analysis in educational 23. Robinson MA. Using multi‑item psychometric scales for
research. Pract Assessment Res Eval 2013;18:6.
research and practice in human resource management. Hum
13. Rattray J, Jones MC. Essential elements of questionnaire design Resour Manag 2018;57:739–50.
and development. J Clin Nurs 2007;16:234–43.
24. Edwards P, Roberts I, Sandercock P, Frost C. Follow‑up by mail
14. Sindhuja T, De D, Handa S, Goel S, Mahajan R, Kishore K.
in clinical trials: Does questionnaire length matter? Control Clin
Pemphigus oral lesions intensity score (POLIS): A novel scoring
Trials 2004;25:31–52.
system for assessment of severity of oral lesions in pemphigus
vulgaris. Front Med 2020;7:449. 25. Sahlqvist S, Song Y, Bull F, Adams E, Preston J, Ogilvie D,
et al. Effect of questionnaire length, personalisation and reminder
15. Costello AB, Osborne J. Best practices in exploratory factor
analysis: Four recommendations for getting the most from your type on response rate to a complex postal survey: Randomised
analysis. Pract assessment Res Eval 2005;10:7. controlled trial. BMC Med Res Methodol 2011;11:62.
16. Wood ND, Akloubou Gnonhosou DC, Bowling JW. Combining 26. Edwards P. Questionnaires in clinical trials: Guidelines for
parallel and exploratory factor analysis in identifying relationship optimal design and administration. Trials 2010;11:2.
scales in secondary data. Marriage Fam Rev 2015;51:385–95. 27. Greenlaw C, Brown‑Welty S. A comparison of web‑based and
17. Yang Y, Xia Y. On the number of factors to retain in exploratory paper‑based survey methods: Testing assumptions of survey
factor analysis for ordered categorical data. Behav Res Methods mode and response cost. Eval Rev 2009;33:464–80.
2015;47:756–72. 28. Podsakoff PM, MacKenzie SB, Lee JY, Podsakoff NP. Common
18. Revelle W, Rocklin T. Very simple structure: An alternative method biases in behavioral research: A critical review of
procedure for estimating the optimal number of interpretable the literature and recommended remedies. J Appl Psychol
factors. Multivariate Behav Res 1979;14:403–14. 2003;88:879‑903.

Indian Dermatology Online Journal | Volume 12 | Issue 2 | March-April 2021 275

You might also like