Guidelines for Effective Questionnaire Development
Guidelines for Effective Questionnaire Development
266 © 2021 Indian Dermatology Online Journal | Published by Wolters Kluwer - Medknow
Kishore, et al.: Developing a questionnaire
validation. The scientific standards are available to select specific essential requirements which are not directly
items, subscales, and entire scales. The researchers can a part of scale development and evaluation; however,
broadly segment scientific criteria for a questionnaire into
reliability and validity.
Questionnaire
Development What to do?
Despite increasing usage, many academicians grossly Purpose
misunderstand the scales. The other complication is that Step-1 Item generation
Literature review
Expert and nonexpert
many authors in the past did not adhere to the rigorous Existing questionnaire
standards. Thus, the questionnaire‑based research was Step-2 Item Formatting
Unambiguous
criticized by many in the past for being a soft science.[2] Step-3 Preliminary
Unloaded
Unbarreled
The scale construction is also not a part of most of the Questionnaire
Covering letter
graduate and postgraduate training. Given the previous Step-4 Validation Order of items
discussion, the primary objective of this article is to Questionnaire layout
Step-5
sensitize researchers about the various intricacies and Pilot Testing Coefficient validity ratio
Coefficient validity index
importance of each step for scale construction. The Step-6 Data Collection Interrater agreement
emphasis is also to make researcher aware and motivate Floor & ceiling effect
to use multiple metrics to assess psychometric properties. Step-7 Evaluation
Rewording
Deletion
Table 1 describes a glossary of essential terminologies used
Data collection
in context to questionnaire. Data Entry
Data cleaning
The process of building a questionnaire starts with item
Descriptive analysis
generation, followed by questionnaire development, and Reliability
concludes with rigorous scientific evaluation. Figure 1 Validity
summarizes the systematic steps and respective tasks Figure 1: Flowchart demonstrating the various steps involved in the
at each stage to build a good questionnaire. There are development of a questionnaire
these improve the utility of the instrument. The indirect advantageous to speak with experts telephonically, face
but necessary conditions are documented and discussed to face, or electronically, requesting their participation
under the miscellaneous category. We broadly segment before mailing the questionnaire. It is good to explain to
and discuss the questionnaire development process under them right in the beginning that this process unfolds over
three domains, known as questionnaire development, phases. The time allowed to respond can vary from hours
questionnaire evaluation, and miscellaneous properties. to weeks. It is recommended to give at least 7 days to
respond. However, a nonresponse needs to be followed
Questionnaire Development up by a reminder email or call. Usually, this stage takes
The development of the list of items is an essential and two to three rounds. Therefore, it is essential to engage
mandatory prerequisite for developing a good questionnaire. with experts regularly; else there is a risk of nonresponse
The researcher at this stage decides to utilize formats such from the study. Table 2 gives general advice to researchers
as Guttman, Rasch, or Likert to frame items.[2] Further, for making a cover letter. The researcher can modify the
the researcher carefully identifies the appropriate member cover letter appropriately for their studies. The authors can
of the expert panel group for face and content validity. consult Rubio and coauthors for more details regarding the
Broadly, there are six steps in the scale development. drafting of a cover letter.[4]
Step I Step IV
It is crucial to select appropriate questions (items) to The responses from each round will help in rewording,
capture the latent trait. An exhaustive list of items is the rephrasing, and reordering of the items in the scale. Few
most critical and primary requisite to lay the foundation questions may need deletion in the different rounds of
of a good questionnaire. It needs considerable work in previous steps. Therefore, it is better to evaluate content
terms of literature search, qualitative study, discussion with validity ratio (CVR), content validity index (CVI), and
colleagues, other experts, general and targeted responders, interrater agreement before deleting any question in the
and other questionnaires in and around the area of interest. instrument. Readers can consult formulae in Table 2 for
General and targeted participants can also advise on items, calculating CVR and CVI for the instrument. CVR is
wording, and smoothness of questionnaire as they will be calculated and reported for the overall scale, whereas
the potential responders. CVI is computed for each item. Researchers need to
consult Lawshe table to determine the cutoff value for
Step II CVR as the same depends on the number of experts in
It is crucial to arrange and reword the pool of questions the panel.[5] CVI >0.80 is recommended. Researchers
for eliminating ambiguity, technical jargon, and loading. interested in detail regarding CVR and CVI can read
Further, one should avoid using double‑barreled, excellent articles written by Zamanzadeh et al. and
long, and negatively worded questions. Arrange all Rubio et al.[4,6] It is crucial to compute CVR, CVI,
items systematically to form a preliminary draft of the and kappa agreement for each item from the rating of
questionnaire. After generating an initial draft, review importance, representativeness, and clarity by experts.
the instrument for the flow of items, face validity The CVR and CVI do not account for a chance factor.
and content validity before sending it to experts. The Since interrater agreement (IRA) incorporates chance
researcher needs to assess whether the items in the factor; it is better to report CVR, CVI, and IRA
score are comprehensive (content validity) and appear to measures.
measure what it is supposed to measure (face validity).
Step V
For example, does the scale measuring stress is measuring
stress or is it measuring depression instead? There is The scholars require to address subtle issues before
no uniformity on the selection of a panel of experts. administering a questionnaire to responders for pilot
However, a general agreement is to use anywhere from a testing. The introduction and format of the scale play a
minimum of 5–15 experts in a group.[3] These experts will crucial role in mitigating doubts and maximizing response.
ascertain the face and content validity of the questionnaire. The front page of the questionnaire provides an overview
These are subjective and objective measures of validity, of the research without using technical words. Further,
respectively. it includes roles and responsibilities of the participants,
contact details of researchers, list of research ethics
Step III (such as voluntary participation, confidentiality and
It is advisable to prepare an appealing, jargon‑free, and withdrawal, risks and benefits), and informed consent for
nontechnical cover letter explaining the purpose and participation in the study. It is also better to incorporate
description of the instrument. Further, it is better to include anchors (levels of Likert item) in each page at the top or
the reason/s for selecting the expert, scoring format, and bottom or both for ease and maximizing response. Readers
explanations of response categories for the scale. It is can refer to Table 3 for detail.
Table 2: General overview and the instructions for rating in the cover letter to be accompanied by the questionnaire
Content Explanation
Construct Definition of characteristics of the measurement
Purpose To evaluate the content and face validity
How Please rate each item for its representativeness and clarity on a scale from 1 to 4
Evaluate the comprehensiveness of the entire instrument in measuring the domain
Please add, delete, or modify any item as per your understanding
Measure CVR CVI
Characteristics Importance Representative Clarity
Scoring 0-Not necessary 1-Not representative 1-Not clear
1-Useful 2-Need major revisions to be representative 2-Need major revisions to be clear
2-Essential 3-Need minor revisions to be representative 3-Need minor revisions to be clear
4-Representative 4-Clear
Formula CVR = (NE-N/2)/(N/2) CVIR=NR/N CVIC=NC/N
where NE=number of experts where CVIR=CVI for representativeness where CVIC=CVI for clarity
rated an item as essential NR=Number of experts rated an item as NC=Number of experts rated an
N=Total number of experts representative (3 or 4) item as clear (3 or 4)
N=Total number of experts N=Total number of experts
Table 3: A random set of questions with anchors at the top and bottom row
Items Strongly disagree Disagree Neutral Agree Strongly agree
(SD) (D) (N) (A) (SA)
Duration of disease (since onset) SD D N A SA
Number of relapse(s) of the disease SD D N A SA
Duration of oral erosions (present episode) SD D N A SA
Number of relapse(s) of oral lesions SD D N A SA
Persistence of oral lesions after subsidence of cutaneous lesions SD D N A SA
Change in size of existing lesion in last 1 week SD D N A SA
Development of new lesions in last 1 week SD D N A SA
Difficulty in eating normal food SD D N A SA
Difficulty in eating food according to their consistency SD D N A SA
Inability to eat spicy food SD D N A SA
Inability to drink fruit juices SD D N A SA
Excessive salivation/drooling SD D N A SA
Difficulty in speaking SD D N A SA
Difficulty in brushing teeth SD D N A SA
Difficulty in swallowing SD D N A SA
Restricted mouth opening SD D N A SA
Strongly disagree Disagree Neutral Agree Strongly agree
in this process is to calculate the appropriate sample size Table 4: A sample of data entry format
for administering a preliminary questionnaire in the target (a) Illustration of master sheet
group. The evaluations of various measures do not follow a Participant Age Religion Family Height Weight Q1 Q2 Q3
sequential order like the previous stage. Nevertheless, these 1 25 1 1 185.0 85.0 1 5 2
measures are critical to evaluate the reliability and validity 2 26 3 1 155.0 63.0 2 5 1
of the questionnaire. 3 22 2 2 155.0 57.0 4 2 1
4 35 2 1 158.5 67.5 3 2 2
Data entry
5 49 1 2 175.0 64.0 2 4 3
Correct data entry is the first requirement to evaluate the 6 40 4 1 159.0 78.0 2 4 3
characteristics of a manually administered questionnaire. Qi→ith Question in the questionnaire, where i=1,2,3, … n
The primary need is to enter the data into an appropriate (b) Illustration of coding sheet
spreadsheet. Subsequently, clean the data for cosmetic and Variable Description Coding and Measurement
logical errors. Finally, prepare a master sheet, and data label valid range scale
dictionary for analysis and reference to coding, respectively. Participant A random None String
Authors interested in more detail can read “Biostatistics serial number
Series.”[7,8] The data entry process of the questionnaire is to participant
like other cross‑sectional study designs. The rows and Age Age in years None Interval
columns represent participants and variables, respectively. (30‑70 years)
It is better to enter the set of items with item numbers. Religion Religion of 1=Hindu Nominal
the participant 2=Sikh
First, it is tedious and time‑consuming to find suitable
variable names for many questions. Second, item numbers 3=Muslim
help in quick identification of significantly contributing and 4=Others
non‑contributing items of the scale during the assessment Q Level of 1=Strongly Ordinal
of psychometric properties. Readers can see Table 4 for agreement in disagree
more detail. the question 2=Disagree
3=Neutral
Descriptive statistics
4=Agree
Spreadsheets are easy and flexible for routine data entry and 5=Strongly
cleaning. However, the same lack the features of advanced agree
statistical analysis. Therefore, the master sheet needs to be
exported to appropriate software for advanced statistical
type of missingness. The typically recommended threshold
analysis. Descriptive analysis is the usual first step which
for the missingness is 5%.[10] There are broadly three types
helps in understanding the fundamental characteristics of
the data. Thus, report appropriate descriptive measures of missingness, such as missing completely at random,
such as mean and standard deviation, and median and missing at random, and not missing at random. After
interquartile/interdecile range for continuous symmetric identification of a missing mechanism, impute the data with
and asymmetric data, respectively.[9] Utilize exploratory single or multiple imputation approaches. Readers can refer
tabular and graphical display to inspect the distribution to an excellent article written by Graham for more details
of various items in the questionnaire. A stacked bar chart about missing data.[11]
is a handy tool to investigate the distribution of data Sample size
graphically. Further, ascertain linearity and lack of extreme
multicollinearity at this stage. Any value of IQC >0.7 The optimum sample size is a vital requisite to build a
warrants further inspection for deletion or modification. good questionnaire. There are many guidelines in the
Help from a good biostatistician is of great assistance for literature regarding recruiting an appropriate sample size.
data analysis and reporting. Literature broadly segments sample size approaches into
three domains known as subject to variables ratio (SVR),
Missing data analysis minimum sample size, and factor loadings (FL). The factor
Missing data is the rule, not the exception. Majority of analysis (FA) is a crucial component of questionnaire
the researchers face difficulties of finding missing values designing. Therefore, recent recommendations are to
in the data. There are usually three approaches to analyze use FLs to determine sample size. Readers can consult
incomplete data. The first approach is to “take all” which Table 5 for sample size recommendations under various
use all the available data for analysis. In the second domains. Interested readers can refer to Beavers and
method, the analyst deletes the participants and variables colleagues for more detail.[12] The stability of the factors
with gross missingness or both from the analysis process. is essential to determine sample size. Therefore, data
The third scenario consists of estimating the percentage and analysis from questionnaires validates the sample size
after data collection. The Kaiser–Meyer–Olkin (KMO) sample size adequacy, IQC, and Bartlett’s test in step 7
criterion testing the adequacy of sample size is available in [Figure 1]. The value of EFA is used at the initial stages
the majority of the statistical software packages. A higher to extract factors while constructing a questionnaire. It is
value of KMO is an indicator of sufficient sample size for especially important to identify an adequate number of
stable factor solution. factors for building a decent scale. The factors represent
latent variables that explain variance in the observed data.
Correlation measures First and the last factor explain maximum and minimum
The strength of relationships between the items is an variance, respectively. There are multiple factor selection
imperative requisite for a stable factor solution. Therefore, criteria, each with its advantages and disadvantages. It
the correlation matrix is calculated and ascertained for is better to utilize more than one approach for retaining
same. There are various recommendations of correlation factors during the initial extraction phase. Readers can
coefficient; however, a value greater than 0.3 is a must.[13] consult Sindhuja et al. for the practical application of more
A lower value of the correlation coefficient will fail to form than one‑factor selection criteria.[14]
a stable factor due to lack of commonality. The determinant
and Bartlett’s test of sphericity can be used to ascertain the
Kaiser’s criterion
stability of the factors. The determinant is a single value Kaiser’s criterion is one of the most popular factor retention
which ranges from zero to one. A nonzero determinant criteria. The basis of the Kaiser criterion is to explain the
indicates that factors are possible. However, it is small in variance through the eigenvalue approach. A factor with
most of the studies and not easy to interpret. Therefore, more than one eigenvalue is the candidate for retention.[15]
Bartlett’s test of sphericity is routinely used to infer that An eigenvalue bigger than one simply means that a single
determinant is significantly different than zero. factor is explaining variance for more than one observed
variable. However, there is a dearth of scientifically
Validity rigorous studies to declare a cutoff value for Kaiser’s
Physical quantities such as height and weight are observable criterion. Many authors highlighted that the Kaiser criterion
and measurable with instruments. However, many tools over‑extract and under‑extract factors.[16,17] Therefore,
need regular calibration to be precise and accurate. The investigators need to calculate and consider other measures
standardization in context to the questionnaire development for extraction of factors.
is known as reliability and validity. The validity is the
Cattell’s scree plot
property which indicates that an instrument is measuring
what it is supposed to measure. Validation is a continuous Cattell’s scree plot is another widespread eigenvalue‑based
process which begins with the identification of domains and factor selection criterion used by researchers. It is popularly
goes on till generalization. There are various measures to known as scree plot. The scree plot assigns the eigenvalues
establish the validity of the instrument. Authors can consult on the y‑axis against the number of factors in the x‑axis.
Table 6 for different types of validity and their metrics. The factors with highest to lowest eigenvalues are plotted
from left to right on the x‑axis. Usually, the scree plots
Exploratory FA form an elbow which indicates the cutoff point for factor
FA assumes that there are underlying constructs (factors) extraction. The location or the bend at which the curve first
which cannot be measured directly. Therefore, the begins to straighten out indicates the maximum number of
investigator collects the exhaustive list of observed factors to retain. A significant disadvantage of the scree
variables or responses representing underlying constructs. plot is the subjectivity of the researcher’s perception of the
Researchers expect that variables or questions in the “elbow” in the plot. Researchers can see Figure 2 for detail.
questionnaire correlate among themselves and load on the
Percentage of variance
corresponding but a small number of factors. FA can be
broadly segmented in exploratory factor analysis (EFA) The variance extraction criterion is another criterion to
and confirmatory factor analysis. The EFA is applied on retain the number of factors. The literature recommendation
the master sheet after assessing descriptive statistics such varies from more than a minimum of 50–70%
as tabular and graphical display, missing mechanism, onward.[12] However, both the number of items and factors
Table 6: Scientific standards to evaluate and report for constructing a good scale
Psychometric Component Definition Indices
properties
Validity Content validity The items are addressing all the relevant aspect of construct Content validity ratio
Content validity indices
Interrater agreement
Face validity The test appears to measure the intended measure Expert opinion (qualitative)
Construct validity The strong (rs) and weak (rw) correlation between same and Exploratory factor analysis
different construct, respectively Correlation coefficient
Criterion validity The correlation between a predictor measure (teamwork) Correlation coefficient
and criterion measures (actual performance in team)
Convergent The correlation between a scale and conceptually similar Correlation coefficient
validity scales or subscales of a scale Multitrait‑multimethod matrix
Reliability Internal The cohesiveness of items in measuring the same variable Coefficient α
consistency consistently Coefficient β
Coefficient Ω
Test‑retest Consistency of score for stable characteristics on separate Correlation coefficient
times Intra‑class correlation coefficient
Alternate forms Consistency of scores among the same sample for similar Correlation coefficient
tests
Descriptive Tabular display Display of essential data characteristics in rows and Mean (SD)
analysis columns Median (IQR)
Graphical display Visual display of large data to exhibit trends, patterns, and Box plot
relationships Bar graph
Missing MCAR Missing data is independent of observed or unobserved data Little’s MCAR
mechanism MAR Missing data is related to observed but not unobserved data Listing and Schlittgen (LS) test
NMAR Missing data is related to unobserved data No standard test (based on
assumptions)
Factorability Sample size
Minimum number of participants required to measure study KMO criteria
outcomes
Correlation A matrix displaying the inter‑correlations among the Determinant
matrix variables
Sphericity Refers to equality of correlations between different items Bartlett’s test
MCAR: Missing completely at random; MAR: Missing at random; NMAR: Not missing at random; KMO: Kaiser‑Meyer‑Olkin; SD: Standard
deviation; IQR: Interquartile range
is the only technique which accounts for the probability correlation >0.70 represents high reliability.[21] The change
that a factor is due to chance. PA simulates data to generate in study condition (recovery of patients after intervention)
95th percentile cutoff line on a scree plot restricted upon the over time can decrease test–retest reliability. Therefore, it
number of items and sample size in original data. The factors is important to report the time between repeated measures
above the cutoff line are not due to chance. PA is the most while reporting test–retest reliability.
robust empirical technique to retain the appropriate number
of factors.[16,20] However, it should be used cautiously for Parallel forms and split‑half reliability
the eigenvalue near the 95th percentile cutoff line. PA is Parallel form reliability is also known as an alternate form
also robust to distributional assumptions of the data. Since of consistency. There are two types of option to report
different techniques have their fair share of advantages and parallel form reliability. In the first method, different
disadvantages, researchers need to assess information on but similar items make alternative forms of the test. The
the basis of multiple criteria. assumptions of both the assessment are that they measure
the same phenomenon or underlying construct. It addresses
Reliability the twin issues of time and knowledge acquisition of test in
Reliability, an essential requisite of a scale, is also known as test–retest reliability. In the second approach, the researcher
reproducibility, repeatability, and consistency. It identifies randomly divides the total items of an instrument into two
that the instrument is consistently measuring the attribute halves. The calculation of parallel form from two halves is
under identical conditions. Reliability is a necessary known as split‑half reliability. However, randomly divided
characteristic of a tool. The trustworthiness of a scale can half may not be similar. The parallel from and split‑half
be increased by increasing and decreasing the systematic reliability are reported with the correlation coefficient. The
and random component, respectively. The reliability of an recommendations are to use a value higher than 0.80 to
instrument can be further segmented and measured with assess the alternate form of consistency.[24] It is challenging
various indices. Reliability is important but it is secondary to generate two types of tests in clinical studies. Therefore,
to validity. Therefore, it is ideal to calculate and report researchers rarely report reliability from two analogous but
reliability after validity. However, there are no hard and separate tests.
fast rules except that both are necessary and important
measures. Readers may consult Table 6 for multiple types General Questionnaire Properties
of indices for reliability. The major issues regarding the reliability and validity of
scale development have already been discussed. However,
Internal consistency there are many other subtle issues for developing a good
Cronbach’s alpha (α), also known as α‑coefficient, is one questionnaire. These delicate issues may vary from a choice
of the most used statistics to report internal consistency of Likert items, length of the instrument, cover letter, web
reliability. The internal consistency using the interitem or internet mode of data collection, and weighting of
correlations suggests the cohesiveness of items in a scale. The immediately preceding issues demand careful
questionnaire. However, the α‑coefficient is sample‑specific; deliberation and attention from the researcher. Therefore,
thus, the literature recommends the same to calculate and the researcher should carefully think through all these
report for all the studies. Ideally, a value of α >0.70 is issues to build a good questionnaire.
preferred; however, the value of α >0.60 is also accepted
for construction of new scale.[21,22] Researchers can increase Likert items
the α‑coefficient by adding items in the scale. However, a The Likert items are the fixed choice ordinal items which
value can either reduce with the addition of non‑correlated capture attitude, belief, and various other latent domains.
items or deletion of correlated items. Corrected interitem The subsequent step is to rank the questions of the Likert
correlation is another popular measure to report for internal scale for further analysis. The numerals for ranking can
consistency. A value of α <0.3 indicates the presence of either start from 0 or 1. It does not make a difference. The
nonrelated items. The studies claim that coefficient beta (β) Likert scale is primarily bipolar as opposite ends endorse
and omega (Ω) are better indices than coefficient‑α, but the contrary idea.[2] These are the type of items which
there is a scarcity of literature reporting these indices.[23] express opinions on a measure from strong disagreement
to strong agreement. The adjectival scales are unipolar
Test–retest scale that tends to measure variables like pain intensity
Test–retest reliability measures the stability of an instrument (no pain/mild pain/moderate pain/severe pain) in one
over time. In other words, it measures the consistency direction. However, the Likert scale (most likely–least
of scores over time. However, the appropriate time likely) can measure almost any attribute. The Likert scale
between repeated measures is a debatable issue. Pearson’s can either have odd or even categories; however, odd
product‑moment and intraclass correlation coefficient categories are more popular. The number of classifications
measure and report test–retest reliability. A high value of in the Likert scale can vary from anywhere between 3 and
11,[2] although the scale with 5 and 7 classes have displayed lengthy questionnaire. There are different opinions about
better statistical properties for discriminating between either grouping or mixing the issues in an instrument.[24]
responses.[2,24] Grouping inflates intra‑scale correlation, whereas mixing
inflates inter‑scale correlation.[28] Both the approaches
Length of questionnaire have empirically shown to give similar results for at least
A good questionnaire needs to include many items to 20 or more items. The questions related to a particular
capture the construct of interest. Therefore, investigators domain can be assigned either equal or unequal weights.
need to collect as many questions as possible. However, the There are two mechanisms to assign unequal weights in
lengthier scale increases both time and cost. The response a questionnaire. In the first situation, researchers affix
rate also decreases with an increase in the length of the different importance to items. In the second method, the
questionnaire.[25] Although what is lengthy is debatable investigators frame more or fewer questions as per the
and varies from more than 4 pages to 12 pages in various importance of subscales in the scale.
studies,[26] the longer scales increase the false positivity
rate.[27] Conclusion
The fundamental triad of science is accuracy, precision,
Translating a questionnaire
and objectivity. The increasing usage of questionnaires in
Many a time, there are already existing reliable and valid medical sciences requires rigorous scientific evaluations
questionnaires for use. However, the expert needs to assess before finally adopting it for routine use. There are
two immediate and important criteria of cultural sensitivity no standard guidelines for questionnaire development,
and language of the scale. Many sensitive questions on evaluation, and reporting in contrast to guidelines such
sexual preferences, political orientations, societal structure, as CONSORT, PRISMA, and STROBE for treatment
and religion may be open for discussion in certain societies, development, evaluation, and reporting. In this article,
religions, and cultures, whereas the same may be taboo or we emphasize on the systematic and structured approach
receive misreporting in others. The sensitive questions need for building a good questionnaire. Failure to meet the
to be reframed considering regional sentiments and culture questionnaire development standards may lead to biased,
in mind. Further, a questionnaire in different language unreliable, and inaccurate study finding. Therefore, the
needs to be translated by a minimum of two independent general guidelines given in this article can be used to
bilingual translators. Similarly, the translated questionnaire develop and validate an instrument before routine use.
needs to be translated back into the original language by a
minimum of two independent and different bilingual experts Financial support and sponsorship
who converted the original questionnaire. The process Nil.
of converting the original questionnaire to the targeted
language and then back to the original language is known Conflicts of interest
as forward and backward translation. The subsequent steps There are no conflicts of interest.
such as expert panel group, pilot testing, reliability, and
validity for translating a questionnaire remain the same as References
in constructing a new scale. 1. Streiner DL, Norman GR, Cairney J. Health Measurement
Scales: A Practical Guide to their Development and Use. USA:
Web‑based or paper‑based Oxford University Press; 2015.
Broadly, paper and electronic format are the two modes 2. Chapple ILC. Questionnaire research: An easy option? Br Dent
of administering a questionnaire to the participants. Both J 2003;195:359.
techniques have advantages and disadvantages. The 3. Boateng GO, Neilands TB, Frongillo EA, Melgar‑Quiñonez HR,
Young SL. Best practices for developing and validating scales
response rate is a significant issue in self‑administered
for health, social, and behavioral research: A primer. Front Public
scales. The significant benefits of electronic format are the Health 2018;6:149.
reduction in cost, time, and data cleaning requirements. 4. Rubio DM, Berg‑Weger M, Tebb SS, Lee ES, Rauch S.
In contrast, paper‑based administration of questionnaire Objectifying content validity: Conducting a content validity
increases external generalization, paper feel, and no need study in social work research. Soc Work Res 2003;27:94–104.
of internet. As per Greenlaw and Welty, the response 5. Lawshe CH. A quantitative approach to content validity. Pers
rate improves with the availability of both the options to Pschol 1975;28:563‑75.
participants. However, cost and time increase in comparison 6. Zamanzadeh V, Ghahramanian A, Rassouli M, Abbaszadeh A,
Alavi‑Majd H, Nikanfar AR. Design and implementation content
to the usage of electronic format alone.[27]
validity study: Development of an instrument for measuring
Item order and weights patient‑centered communication. J Caring Sci 2015;4:165‑78.
7. Kishore K, Kapoor R. Statistics corner: Structured data entry.
There are multiple ways to order an item in a questionnaire. J Postgr Med Educ Res 2019;53:94–7.
The order of questions becomes more critical for a 8. Kishore K, Kapoor R, Singh A. Statistics corner: Data
cleaning‑I. J Postgrad Med Educ Res 2019;53:130–2. 19. Dinno A. Exploring the sensitivity of Horn’s parallel analysis to
9. Kishore K, Kapoor R. Statistics corner: Reporting descriptive the distributional form of random data. Multivariate Behav Res
statistics. J Postgrad Med Educ Res 2020;54:66–8. 2009;44:362–88.
10. Jakobsen JC, Gluud C, Wetterslev J, Winkel P. When and how 20. DeVon HA, Block ME, Moyle‑Wright P, Ernst DM, Hayden SJ,
should multiple imputation be used for handling missing data Lazzara DJ, et al. A psychometric toolbox for testing validity
in randomised clinical trials–A practical guide with flowcharts. and reliability. J Nurs Scholarsh 2007;39:155–64.
BMC Med Res Methodol 2017;17:162. 21. Straub D, Boudreau MC, Gefen D. Validation guidelines for IS
11. Graham JW. Missing data analysis: Making it work in the real positivist research. Commun Assoc Inf Syst 2004;13:24.
world. Annu Rev Psychol 2009;60:549–76. 22. Revelle W, Zinbarg RE. Coefficients alpha, beta, omega, and the
12. Beavers AS, Lounsbury JW, Richards JK, Huck SW. Practical glb: Comments on Sijtsma. Psychometrika 2009;74:145.
considerations for using exploratory factor analysis in educational 23. Robinson MA. Using multi‑item psychometric scales for
research. Pract Assessment Res Eval 2013;18:6.
research and practice in human resource management. Hum
13. Rattray J, Jones MC. Essential elements of questionnaire design Resour Manag 2018;57:739–50.
and development. J Clin Nurs 2007;16:234–43.
24. Edwards P, Roberts I, Sandercock P, Frost C. Follow‑up by mail
14. Sindhuja T, De D, Handa S, Goel S, Mahajan R, Kishore K.
in clinical trials: Does questionnaire length matter? Control Clin
Pemphigus oral lesions intensity score (POLIS): A novel scoring
Trials 2004;25:31–52.
system for assessment of severity of oral lesions in pemphigus
vulgaris. Front Med 2020;7:449. 25. Sahlqvist S, Song Y, Bull F, Adams E, Preston J, Ogilvie D,
et al. Effect of questionnaire length, personalisation and reminder
15. Costello AB, Osborne J. Best practices in exploratory factor
analysis: Four recommendations for getting the most from your type on response rate to a complex postal survey: Randomised
analysis. Pract assessment Res Eval 2005;10:7. controlled trial. BMC Med Res Methodol 2011;11:62.
16. Wood ND, Akloubou Gnonhosou DC, Bowling JW. Combining 26. Edwards P. Questionnaires in clinical trials: Guidelines for
parallel and exploratory factor analysis in identifying relationship optimal design and administration. Trials 2010;11:2.
scales in secondary data. Marriage Fam Rev 2015;51:385–95. 27. Greenlaw C, Brown‑Welty S. A comparison of web‑based and
17. Yang Y, Xia Y. On the number of factors to retain in exploratory paper‑based survey methods: Testing assumptions of survey
factor analysis for ordered categorical data. Behav Res Methods mode and response cost. Eval Rev 2009;33:464–80.
2015;47:756–72. 28. Podsakoff PM, MacKenzie SB, Lee JY, Podsakoff NP. Common
18. Revelle W, Rocklin T. Very simple structure: An alternative method biases in behavioral research: A critical review of
procedure for estimating the optimal number of interpretable the literature and recommended remedies. J Appl Psychol
factors. Multivariate Behav Res 1979;14:403–14. 2003;88:879‑903.