0% found this document useful (0 votes)
7 views89 pages

Guidelines for Effective Scale Development

Uploaded by

T
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views89 pages

Guidelines for Effective Scale Development

Uploaded by

T
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Guidelines in Scale Development

Chapter Review
This presentation provides a set of specific guidelines that we can use in developing
measurement scales through 8 steps as follows:

1 Determine Clearly What It Is You Want to Measure Consider Inclusion of Validation Items 5

2 Generate an Item Pool Administer Items to a Development Sample 6

3 Determine the Format for Measurement Evaluate the Items 7

4 Have Initial Item Pool Reviewed by Experts


Optimize Scale Length 8
In this presentation we focus on scales that intend to measure
elusive phenomena that cannot be observed directly.

Let us recall that:


Yükleniyor…
Researchers often select items and throw them together assuming
they constitute a suitable scale. But in fact they are not.

• We have to be aware whether the items:


1. share a common cause (thus constituting a scale); or
2. share a common consequence (thus constituting an
index);or
3. Are just examples of a shared category that does not imply
either a common cause or consequence (thus constituting
an emergent variable)
“ Step 1: Determine Clearly What It Is You Want to
Measure.

• Did you happen to think that you have a clear idea of something you wish to
measure, to discover that your ideas are vaguer than you thought?
• Do you think that the scale should be based in theory, or should you begin in new
intellectual directions?


Yükleniyor…
How specific do you think the measure should be?
Should some aspect of the phenomenon be emphasized more than others?
Theory as an Aid to Clarity

Clear ideas are important as a guide, and theory is a great aid to clarity.

If it happens that the extant theory offers no guide, then we may decide that a new
intellectual direction is necessary.

Even if there is no available theory to guide us, we must arrange our own
conceptual formulations prior to trying to operationalize them (e.g., a tentative
theoretical model that may serve as a guide to scale development).
Specificity as an Aid to Clarity

The level of specificity or generality at which a construct is measured is important.

Sometimes, a scale is intended to relate to very specific behaviours or constructs,


while at other times, a more general and global measure is sought.

For example, Locus of control construct. This construct can be applied broadly
(global behaviour), or narrowly (a specific context)

Examples of different locus of control scales are: (See page 104)


. Rotter’s (1966) Internal-External scale;
. Levenson (1973) multidimensional locus of control scale; and
. Wallston, Wallston, and DeVellis (1978) Multidimensional Health Locus of
Control (MHLC) scales.
Which of the mentioned (locus of
control) scales is more useful?
Being Clear About What to Include in a Measure

We should ask ourselves if the construct we wish to measure is distinct from other
constructs. For example, General anxiety scale might assess both test anxiety and
social anxiety.

If it matches the our goal, then it is okay. But if we are interested in only one
specific type of anxiety, then it should exclude all others.

Sometimes, apparently similar items may tap quite different constructs.


For example, a certain depression measure might have some items that tap
somatic aspects of depression (See page 105).
“ Step 2: Generate an Item Pool

The first step is to generate a large pool of items that are candidates for eventual
inclusion in the scale.
Choose Items That Reflect the Scale’s Purpose

Items should be selected with a specific measurement goal in mind.

The content of each item should primarily reflect the construct of interest (latent
variable), so we can call it a homogeneous scale.

Yükleniyor…
A good set of items is chosen randomly from the large universe of items relating to
the construct of interest (remember content validity)

The properties of a scale are determined by the items that make it up (poor/good
reflection).
Remember, items we select or create are manifestations of a common latent
variable that is their cause.

Just because items relate to a common category, that does not guarantee that they
have the same underlying latent variable (e.g., attitudes define categories of
constructs rather than the constructs themselves) (see examples on page 106-7)
Redundancy

Redundancy could be both a good and a bad feature. How?

Redundancy is not a bad thing when developing a scale.

The theoretical models that guide our scale development efforts are based on
redundancy.

By using multiple and seemingly unimportant items, the content that is common to
the items will summate across items while their irrelevant idiosyncrasies will
cancel out. Without redundancy, this would be impossible.
Redundancy could be related to the grammatical structure, and choice of words.
For example, consider these two items:
“A really important thing is my child’s success,” (original item).
“The really important thing is my child’s success.” (altered version).

Redundancy could also be related to the variable of interest.


For example, read these two items:
“I will do almost anything to ensure my child’s success”
“No sacrifice is too great if it helps my child succeed”

When irrelevant redundancies are avoided, relevant redundancies will yield more
reliable item sets.
Construct-irrelevant similarities in wording may lead the respondents’ to react
similarly to the items which may result in an inflated estimate of reliability.
For example, some scales contain several items begin with a common phrase such
as “When I think about it,…”

How global or specific the construct of interest is can alter the impact of redundancy.
For example, in an instrument that has been designed to capture all aspects of
emotion, several items about anxiety may pose a problem.

It may result in a number of problems (e.g., it may undermine the unidimensionality


of the item set).
Number of Items

It is impossible to specify the number of items that should be included in an initial


pool.

Internal consistency reliability is a function of how strongly the items correlate with
one another and how many items you have in the scale.

When we select lots of items we try to increase the insurance against poor internal
consistency.
But we should not forget that the more items we have in our pool, the fussier we
can be about choosing ones that will do the job we intend.
It would not be unusual to begin with a pool of items that is three or four times as
large as the final scale. Thus, a 10-item scale might evolve from
a 40-item pool.

The larger the item pool, the better.

We can eliminate some items based on a priori criteria, such as lack of clarity,
questionable relevance, or undesirable similarity to other items.
Beginning the Process of Writing Items

It is beneficial to begin with a statement that is a paraphrase of the construct you


want to measure. How?

The goal at this early stage is simply to identify a wide variety of ways that the
central concept of the intended instrument can be stated.

After generating perhaps three or four times the number of items you anticipate
including in the final instrument, you will look over what you have written. (Now is
the time to become critical) So;

Items can be examined for how well they capture the central ideas and for clarity of
expression.
Characteristics of Good and Bad Items

Again, clarity (a good item should be unambiguous).

Avoid exceptionally lengthy items (increases complexity and diminishes clarity).

We should consider the reading difficulty level at which the items are written.
Fry (1977) delineates several steps to quantifying reading level (see page 111).

Avoid double-barreled items (e.g., “I support civil rights because discrimination is


a crime against God).
Avoid ambiguous pronoun references (e.g., Murderers and rapists should not seek
pardons from politicians because they are the scum of the earth).

Avoid Misplaced modifiers. It create ambiguities similar to ambiguous pronoun


references (e.g., Our members of Congress should work diligently to legalize
prostitution in the House of Representatives).

Avoid using adjective forms instead of noun forms. It can create unintended
confusion (e.g., consider the differences in meaning between:

. “All vagrants should be given a schizophrenic assessment”


. “All vagrants should be given a schizophrenia assessment.”
Positively and Negatively Worded Items

Why some scale developers choose to write negatively worded items while others
choose to write positively worded items?

The goal is to arrive at a set of items, some of which indicate a high level of the
latent variable when endorsed and others that indicate a high level when not
endorsed. For example, The Rosenberg (1965) Self-Esteem Scale includes items
of high and low esteem (see page 112).

The purpose of wording items both positively and negatively within the same scale
is usually to avoid an acquiescence, affirmation, or agreement bias.

There may be a price to pay for including positively and negatively worded
items (See page 113).
“ Step 3: Determine the Format for Measurement

We will discuss common formats that depart from the pattern implied by the
theoretical models as well as ones that adhere to that pattern.
Thurstone Scaling

Consider the example of tuning fork. (A Thurstone scale is intended to work in the
same way).

The scale developer attempts to generate items that are differentially responsive to
specific levels of the attribute in question. When the “pitch” of a particular item
matches the level of the attribute a respondent possesses, the item will signal this
correspondence.
Items could be developed to correspond with different intensities of the attribute,
could be spaced to represent equal intervals, and could be formatted with agree-
disagree response options.

As Nunnally (1978) points out, developing a true Thurstone scale is considerably


harder than describing one. Finding items that consistently “resonate” to specific
levels of the phenomenon is quite difficult.

Although Thurstone scaling is an interesting and sometimes


suitable approach, it will not be referred to in the remainder of this text.
Guttman Scaling

A Guttman scale is a series of items tapping progressively higher levels of an


attribute. Thus, a respondent should endorse a block of adjacent items until, at
a critical point, the amount of the attribute that the items tap exceeds that
possessed by the subject.

Guttman scales can work quite well for objective information or in situations where
it is a logical necessity that responding positively to one level of a hierarchy implies
satisfying the criteria of all lower levels of that hierarchy.

Like Thurstone scales, Guttman scales undoubtedly have their place, but their
applicability seems rather limited.
Thurstone scale for measuring parents’ Guttman version of the preceding parental
aspirations for their children’s educational aspiration scale (page116)
and career attainments (page115)

In both approaches, the disadvantages and difficulties will often outweigh the advantages
Scales With Equally Weighted Items

An attractive feature of scales of this type is that the individual items can have a
variety of response-option formats.
How Many Response Categories?

Most scale items consist of two parts: a stem and a series of response options.
Explain ?

Some item-response formats allow the subject an infinite or very large number of
options, whereas others limit the possible responses.
For example, the response format for measuring anger could be measured in
different way:

It could be calibrated from “no anger at all” at the base of the thermometer
to “complete, uncontrollable rage” at its top.

It also could be measured by asking participants to indicate, using a


Yükleniyor…
number from 1 to 100, how much anger each situation provoked.

Another way to measure is restrict the response options to a few choices,


such as “none,” “a little,” “a moderate amount,” and “a lot,” or to a simple
binary selection between “angry” and “not angry.”
What are the relative advantages of
these alternatives in the previous
example?
A desirable quality of a measurement scale is variability.

A measure cannot covary if it does not vary.

How can we increase the opportunities for variability ?

. One way is to include lots of scale items.


2. Another is to provide numerous response options within items.

A 0-to-100 scale might reveal wide differences in reactions to these situations and
yield good variability for the two-item scale. Likewise, a 50-item scale with two
response option format might yield sufficient variability.
Another issue related to the number of response options is the respondents’ ability
to discriminate meaningfully.

How fine a distinction can the typical subject make? This obviously depends on
what is being measured.

Sometimes, the respondent’s ability to discriminate meaningfully between


response options will depend on the specific wording or physical placement of
those options.

For example, using vague quantity descriptors, such as “several,” “few,” and
“many,” may create problems.
When respondents are presented with an obvious continuum, they often seem to
understand what is desired.

An ordering such as (Many Some Few Very Few None) might be problematic.

The worst circumstance is to combine ambiguous words with ambiguous page


locations. (See the example on page 118-119).

Terms such as somewhat and not very are difficult to differentiate under the best of
circumstances.

Another issue to consider is the investigator’s ability and willingness to record


a large number of values for each item.
There is at least one more issue related to the number of responses.
Assuming that a few discrete responses are allowed for each item,
Should the number be odd or even?

Again, this depends on the type of question, the type of response option, and the
investigator’s purpose.

If the response options are bipolar, with one extreme indicating the opposite of the
other (e.g., a strong positive vs. a strong negative attitude), an odd number of
responses permits equivocation (e.g., “neither agree nor disagree”) or uncertainty
(e.g., “not sure”); an even number usually does not.
An odd number implies a central “neutral” point (e.g., neither a positive nor a
negative appraisal). An even number of responses, on the other hand, forces the
respondent to make at least a weak commitment in the direction of one or the other
extreme (e.g., a forced choice between a mildly positive or mildly negative
appraisal as the least extreme response).

We may want to prevent equivocation if it is felt that subjects will select a neutral
response as a means of avoiding a choice.

Neither format is necessarily superior.


• What do you think about the following two alternative formats, for a study of social
comparisons among people with arthritis?

1. Would you prefer information about:


(a) Patients who have worse arthritis than you have
(b) Patients who have milder arthritis than you have

2. Would you prefer information about:


(a) Patients who have worse arthritis than you have
(b) Patients who have arthritis equally as bad as you have
(c) Patients who have milder arthritis than you have

• A neutral option such as 2b might permit unwanted equivocation.


• A neutral point may also be desirable.
Specific Types of Response Formats

There are several ways to present items that are used widely and have proven
successful in diverse applications.
Likert Scale

Likert scaling is widely used in instruments measuring opinions, beliefs, and attitudes

Item is presented as a declarative sentence, followed by response options that


indicate varying degrees of agreement with or endorsement of the statement.

Using odd or even number of response options mainly depends on the phenomenon
being investigated and the goals of the investigator.

Generally six possible responses are included : “strongly disagree,” “moderately


disagree,” “mildly disagree,” “mildly agree,” “moderately agree,” and “strongly agree.”

A neutral midpoint can also be added.


A good Likert item should state the opinion, attitude, belief, or other construct under
study in clear terms.

It is often useful for these statements to be fairly (though not extremely) strong when
used in a Likert format.
Think about the strength of following sentences. Which is best for a Likert scale?
“Physicians generally ignore what patients say”

“Sometimes, physicians do not pay as much attention as they should to


patients’ comments,”

“Once in a while, physicians might forget or miss something a patient has


told them”

In general, very mild statements may elicit too much agreement when used in Likert
scales. For example, consider this sentence “The safety and security of citizens is
important.” Many people will strongly agree with it.
How to calibrate the strength of
a statement?
A useful way of calibrating how strongly or mildly a statement should be worded is to
do the following:

Imagine the typical respondent for whom the scale is intended.


Try to imagine how that person would respond to items of different forcefulness.
Now, think about what sort of item wording would be most likely to elicit
a response from that typical respondent that was at or near the centre of the
Likert scale response options you plan to use.
Semantic Differential

It is used in reference to one or more stimuli. The stimulus might be a group of


people, such as automobile salesmen. (e.g., attitude).

Identification of the target stimulus is followed by a list of adjective pairs. Each pair
represents opposite ends of a continuum, defined by adjectives (e.g., honest and
dishonest).
The semantic differential scaling method is chiefly associated with the attitude
research (e.g., Osgood & Tannenbaum, 1955).

The individual lines (seven and nine are common numbers) represent points along
the continuum defined by the adjectives.

The adjectives one chooses can be either bipolar or unipolar, depending, as always,
on the logic of the research questions the scale is intended to address.

Like the Likert scale, the semantic differential response format can be highly
compatible with the theoretical models.
Visual Analog

Visual analog scale is similar to the semantic differential.

Respondents are instructed to place a mark at a point on the line that represents
their opinions, experiences, beliefs, or whatever is being measured.

The visual analog scale, as the term analog in the name implies, is a continuous
scale.

Consider a visual analog scale for pain such as this:


A major advantage of visual analog scales is that they are potentially very sensitive
which in turn makes them useful for measuring phenomena before and after some
intervening event, such as an intervention or experimental manipulation, that exerts
a relatively weak effect.

Sensitivity may be more advantageous when examining changes over time within
the same individual rather than across individuals
Another potential advantage of visual analog scales when they are repeated over
time is that it is difficult or impossible for subjects to encode their past responses
with precision.

Visual analog scales have often been used as single-item measures. This has the
sizable disadvantage of precluding any determination of internal consistency.

How we can assess reliability with a single-item measure?


Numerical Response Formats and Basic Neural Processes

Certain response options may correspond to how the brain processes numerical
information (Zorzi, Priftis, and Umilitá, 2002).

As the typical Likert scale, numbers arrayed in a sequence express quantity not
only in their numerical values but in their locations.

It is suggested that the visual line of numbers is not merely a convenient


representation but corresponds to fundamental neural processes.
For example, when individuals were asked what would be midway between points
labelled “3” and “9,” errors were shifted to the right (i.e., to higher values).

The authors conclude that their work constitutes “strong evidence that the mental
number line is more than simply a metaphor” and that “thinking of numbers in
spatial terms (as has been reported by great mathematicians) may be more
efficient because it is grounded in the actual neural representation of numbers”
Binary Options

A response format that gives subjects a choice between binary options for each
item.

Subjects might, for example, be asked to check off all the adjectives on a list that
they think apply to themselves. Or they may be asked to answer “yes” or “no” to
a list of emotional reactions they may have experienced in some specified situation.

In both cases, responses reflecting items sharing a common latent variable (e.g.,
adjectives such as “sad,” “unhappy,” and “blue” representing depression) could be
combined into a single score for that construct.
A major shortcoming of binary responses is that each item can have only minimal
variability. Similarly, any pair of items can have only one of two levels of
covariation (agreement or disagreement).
With binary items, each item contributes precious little to that sum because of the
limitations in possible variances and covariances.

Binary items are usually extremely easy to answer. Therefore, the burden placed
on the subject is low for any one item.
Item Time Frames

• It is about the temporal features of the measures.

• Theory is an important guide to this process.

• The investigator should choose a time frame for a scale actively rather than
passively.

• The item formats, including response options and instructions, should reflect the
nature of the latent variable of interest and the intended uses of the scale.

• Is the scale intended to detect subtle variations occurring over a brief time frame
(e.g., increases in negative affect after viewing a sad movie) or changes that may
evolve over a lifetime (e.g., progressive political conservatism with increasing
age)?
“ Step 4: Have Initial Item Pool Reviewed by Experts

Having a group of people who are knowledgeable in the content area review the
item pool.

This review serves multiple purposes related to maximizing the content validity of
the scale:

. Confirming or invalidating your definition of the phenomenon.

. Evaluating the items’ clarity and conciseness (this bears on item reliability)

. Pointing out ways of tapping the phenomenon that you have failed to
include.
Sometimes, content experts might not understand the principles of scale
construction. This can lead to bad advice.

The final decision to accept or reject the advice of your experts is your responsibility
as the scale developer.
“ Step 5: Consider Inclusion of Validation Items

It might be possible and relatively convenient to include some additional items in
the same questionnaire that will help in determining the validity of the final scale.
There are at least two types of items to consider.

The first type serves to detect flaws or problems. (e.g., other motivations
that influence participants responses. such as social desirability).

The second type concerns the construct validity of the scale.

Rather than mounting a separate validation effort after constituting the final scale, it
may be possible to include measures of relevant constructs at this stage.

Step 6: Administer Items to a Development Sample

• The sample of subjects should be large. But how large is large?
• Is there a consensus on this issue?
What is the rationale for a large sample?

Nunnally (1978) points out that the primary sampling issue in scale development
involves the sampling of items from a hypothetical universe.

Nunnally suggests that 300 people is an adequate number.

However, practical experience suggests that scales have been successfully


developed with smaller samples.

The number of items and the number of scales to be extracted also have a bearing
on the sample size issue. If only a single scale is to be extracted from a pool of
about 20 items, fewer than 300 subjects might suffice.
There are several risks in using too few subjects.

. The patterns of covariation among the items may not be stable.

. The development sample may not represent the population for which the
scale is intended (we should consider both the size and composition of the
development sample)
• Not all types of non-representativeness are identical.

• There are at least two different ways in which a sample may not be representative
of the larger population.

1. The first involves the level of attribute present in the sample versus the
intended population. Explain?

2. A more troublesome type of non-representativeness involves a sample that


is qualitatively rather than quantitatively different from the target population
(e.g., items may have a different meaning for them than
for people in general) explain?

• The consequences of this second type of sample non-representativeness can


severely harm a scale development effort.
“ Step 7: Evaluate the Items

It is time to evaluate the performance of the individual items so that appropriate
ones can be identified to constitute the scale.
Initial Examination of Items’ Performance

The ultimate quality we seek in an item is a high correlation with the true score of
the latent variable.

We cannot directly assess the true score, thus, we cannot directly compute its
correlations with items. However, we can make inferences based on the formal
measurement models that have been discussed thus far.

So, we can learn about relationships to true scores from correlations among items.

The higher the correlations among items are, the higher the individual item
reliabilities
The more reliable the individual items are, the more reliable the scale that they
compose will be. So, the first quality we seek in a set of scale items is that they be
highly intercorrelated.

One way to determine how intercorrelated the items are is to inspect the correlation
matrix.
Reverse Scoring

If there are items whose correlations with other items are negative, then the
appropriateness of reverse scoring those items should be considered.

For example, if we initially anticipated two separate groups of items (e.g., happiness
and sadness) but decide for some reason that they should be combined into
a single group. So, we would reverse score the sadness item.

Sometimes, items are administered in such a way that they are already reversed.

For example, subjects might be asked to circle higher numerical values to indicate
agreement with a “happy” item and lower values to endorse a “sad” one.
• Consider these 2 example (What limitation could be here ?)

1. I am sad often.

2. Much of the time, I am happy


This process may confuse the subject.

People may ignore the words after realizing that they are the same for all items.
How we can eliminate?
It is probably preferable to altering the order of the descriptors (e.g., from “strongly
disagree” to “strongly agree” from left to right for some items and the reverse for
others).

Another option is to have both the verbal descriptions and their corresponding
numbers the same for all items but to enter different values for certain items at the
time of data coding.
The easiest method for reverse scoring is to do so electronically once the data
have been entered into a computer.
Item-Scale Correlations

• If we want to arrive at a set of highly intercorrelated items, then each individual


item should correlate substantially with the collection of remaining items.

• We can examine this property for each item by computing its item-scale
correlation.
There are two types of item scale correlation:

The corrected item-scale correlation (correlates the item being evaluated


with all the scale items, excluding itself)

The uncorrected item-scale correlation (correlates the item in question with


the entire set of candidate items, including itself)

In theory, the uncorrected value tells us how representative the item is of the whole
scale.

It is probably advisable to examine the corrected item-total correlation. An item with


a high value for this correlation is more desirable than an item with a low value.
Item Variances

Another valuable attribute for a scale item is relatively high variance. (of course,
increasing variance by adding to the error component is not desirable)

Comparing item variances may also be useful, especially if the goal is to develop a
tool that meets the assumptions of essential tau equivalence.
Item Means

A mean close to the centre of the range of possible scores is also desirable.

For example, If the response options for each item ranged from 1 (corresponding
with “strongly disagree”) to 7 (for “strongly agree”), an item mean near 4 would be
ideal. If a mean were near one of the extremes of the range, then the item might
fail to detect certain values of the construct.

Generally, items with means too near to an extreme of the response range will
have low variances, and those that vary over a narrow range will correlate poorly
with other items.

Remember, an item that does not vary cannot covary


Dimensionality

Items may have no common underlying variable (as in an index or emergent


variable) or may have several.
So, a set of items is not necessarily a scale.
Determining the nature of latent variables underlying an item set is critical.
The best means of determining which groups of items, if any, constitute
a unidimensional set is by factor analysis.
Although factor analysis requires substantial sample sizes, so does scale
development in general. If there are too few respondents for factor analysis, the
entire scale development process may be compromised.
Consequently, factor analysis of some sort should generally be a part of the scale
development process at this stage
Reliability

Recall that one of the most important indicators of a scale’s quality is the reliability
coefficient, alpha.

There are several options for computing alpha, differing in degree of automation.
Explain and mention the programs?

Another option for computing alpha is to do so by hand (Spearman-Brown formula)


Theoretically, alpha can take on values from 0.0 to 1.0

If alpha is negative, something is wrong (A likely problem is negative correlations -


or covariances- among the items.)

If this occurs, try reverse scoring or deleting items as described earlier.


What reduces alpha ?
• A noncentral mean, poor variability, negative correlations among items, low
item-scale correlations, and weak interitem correlations will tend to reduce alpha
and potentially warrant the use of an alternative, such as omega.

• Alpha is an indication of the proportion of variance in the scale scores that is


attributable to the true score.
What is the acceptable lower
bound of alpha?
• Nunnally (1978) suggests a value of .70 as an acceptable lower bound for alpha.
(It is not unusual to see published scales with lower alphas.)

• There are many arguments, anyways, they are personal and subjective groupings
of alpha values.

• When a scale consists of a single item, it will be impossible to use alpha as the
index of reliability. If possible, some reliability assessment should be made.

• Test-retest correlation may be the only option in the single-item instance.

• A preferable alternative, if possible, would be to constitute the scale using more


than a single item.

• Omega is an alternative to alpha that may be appropriate when the


assumptions for essentially tau-equivalent tests have not been met.
“ Step 8: Optimize Scale Length

At this stage of the scale development process, the investigator has a pool of
items that demonstrate acceptable reliability.
• A scale’s alpha is influenced by two characteristics:
1. The extent of covariation among the items
2. The number of items in the scale.

• For items that have item-scale correlations about equal to the average inter-item
correlation (i.e., items that are fairly typical), adding more will increase alpha and
removing more will lower it.
Generally, shorter scales are good because they place less of a burden on
respondents.

Longer scales, on the other hand, are good because they tend to be more reliable.

We should give some thought to the optimal balance between brevity and reliability.

If a scale’s reliability is too low, then brevity is no virtue (subjects may, indeed, be
more willing to answer a 3-item scale than a 10-item scale).

However, if we cannot assign any meaning to the scores obtained from the shorter
version, then nothing has been gained.
Effects of Dropping “Bad” Items

Does dropping “bad” items increase or lowers alpha?


It depends on how poor the items are, and on the number of items in the scale.

With fewer items, a greater change in alpha results from the addition or subtraction
of each item.
Yükleniyor…
If the average inter-item correlation among four items is .50, the alpha will equal
.80. If there are only three items with an average inter-item correlation of .50, alpha
drops to .75.
If an item has a sufficiently lower-than-average correlation with the other items,
dropping it will raise alpha.

If its average correlation with the other items is only slightly below (or equal to or
above) the overall average, then retaining (keeping) the item will increase alpha.
Tinkering with Scale Length

Items that contribute least to the overall internal consistency should be the first to
be considered for exclusion. But how to define such items?

The item whose omission has the least negative or most positive effect on alpha is
usually the best one to drop first.

The item-scale correlations can also be used as a barometer of which items are
expendable.
Those with the lowest item-scale correlations should be eliminated first.

SPSS also provides a squared multiple correlation for each item, obtained by
regressing the item on all the remaining items. This is an estimate of the item’s
communality, the extent to which it shares variance with the other items.

A poor item-scale correlation is typically accompanied by a low squared multiple


correlation and a small loss, or even a gain, in alpha when the item is eliminated.

The reliability of alpha as an estimate of reliability increases with the number of


items.
Split Samples

If the development sample is sufficiently large, it may be possible to split it into two
subsamples:
One can serve as the primary development sample (arrive at a final version
of the scale that seems optimal).
One can be used to cross-check the findings (replicate the findings).

If the alphas remain fairly constant across the two subsamples, you can be more
comfortable assuming that these values are not distorted by chance.

The two subsamples are likely to be much more similar than two totally different
samples.
The subsamples, divided randomly from the entire development sample, are likely
to represent the same population; in contrast, an entirely new sample might
represent a slightly different population.
Also, data collection periods for the two subsamples are not separated by time,
whereas a development sample and a totally separate sample almost always are.

Despite the unique similarity of the resultant subsamples, replicating findings by


splitting the developmental sample provides valuable information about scale
stability.

The most obvious way to split a sufficiently large sample is to halve it.
However, if the sample is too small to yield adequately large halves, you can split
unevenly (unequally).

The larger subsample can be used for the more crucial process of item evaluation
and scale construction and the smaller for cross-validation.
Thank you!

You might also like