0 ratings0% found this document useful (0 votes) 18 views20 pagesChapter5 Qualities R Instruments
Research instruments for academic research
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content,
claim it here.
Available Formats
Download as PDF or read online on Scribd
Chapter 5
QUALITIES OF A GOOD RESEARCH
INSTRUMENT
Research-made instruments such as tests, question-
neires, rating scales, interviews, observation schedule, ete., should
meet the qualities of a good research instrument before they are
used. These measuring instruments are used for gathering or
collecting data, and are important devices because the success
or failure of a study lies on the data gathered.
The qualities of a good research instument are (1) valid-
ity; (2) reliability, and (3) usability.
Validity
Validity means the degree to which a test or measuring
instrument measures what it intends to measure. The validity
of a measuring instrument has to do with its soundness, what
the tost or questionnaire measures its effectiveness, how well
it could be applied.
For instance, in the study entitled, “Correlates of NCEE
Percentile Rank and achievement of First Year BSE and BEEd
Students in SUC in Region 6 (Western Visayas)” achievement is
the dependent variable thus, the measuring instrument is an
achievement test on what level of achievement it measures and
how effective it manifests itself. In other words, the achievement
test must be valid.
Generally, no test or research instrument can be said to
have a “high” or “low” validity in the abstract. Its validity must
be determined with reference to the particular use for which the
test is being considered. The validity of test must always be
considered in relation to the purpose it serves. Validity is always
specific in relation to some definite situation. Likewise, a valid
test is always valid
SaQUALITIES OF GOOD RESEARCH INSTRUMENT 73
Types of Validity
Validity is classified under four types, namely, content
validity, concurrent validity, predictive validity, and construct
validity.
Content validity. Content validity means the extent to
which the content or topic of the test is truly representative
of the content of the course. It involves, essentially, the sys-
tematic examination of the test content to determine whether
it covers a representative sample of the behavior domain to
be measured, It is very important that the behavior domain
to be tested must be systematically analyzed to make certain
that all major aspects are covered by the test items and in
correct proportions. The domain under consideration should
be fully described in advance rather than defined after the test
has been prepared
Content validity is described by the relevance of a test
to different types of criteria, such as thorough judgments and
systematic examination of relevant course syllabi and text-
book, pooled judgments of subject-matter experts, statements
of behavioral objectives, analysis of researcher-made test
questions, among others. Thus, content validity depends on
the relevance of the individual's responses to the behavior area
under consideration rather on the apparent relevance of item
content.
Content validity is commonly used in evaluating achieve-
ment test. A well-constructed achievement test should cover
the objectives of instruction, not just its subject matter. The
Taxonomy of Educational Objectives by Bloom would be of
great help in listing the objectives to be covered in an achieve-
ment test. Three domains of behavior are included, namely,
cognitive. affective, and psychomotor.
Content validity is particularly appropriate for the cri-
terion-referenced measure. It is also applicable to certain oc-
cupational tests designed to select and classify employees. But
content validity is inappropriate for aptitude and personality
tests. These tests are not based on a specified course of
instruction from which the test content can be drawn, and they
bear less intrinsic resemblance to the behavior domain.14 METHODS OF RESEARCH AND THESIS WRITING
Mlustration
For example, a researcher wishes to validate a test or
questionnaire in Biology. He requests experts in Biology
to judge if the test items or questions measure the know-
ledge, skills and values supposed to be measured. Another
way of testing content validity is for the researcher to
check if the test items cr questions represent the know-
ledge, skills and values suggested in the Biology course con-
tent.
Good and Scates (1972) suggested the evidences of ques-
tionnaire validity which are as follows:
1. Is the question on the subject?
2. Is the question perfectly clear and unambiguous?
3. Does the question get at something stable, some-
thing relatively deep-seated, well-considered,
nonsuperficial, and not ephemeral, but something
which is typical of the individual or of the situa-
tion?
Does the question pull?
Do the responses show a reasonable range of vari-
ation?
6. Is the information obtained consistent?
Is the item sufficiently inclusive?
8. Is there a possibility of using an external criterion
to evaluate the questionnaire?
The researcher requires a selected group of experts to
validate the content of the questionnaire on the basis of the
foregoing questions. If answers of experts are all affirmative,
the questionnaire is valid 2
Concurrent validity. Concurrent validity is the degree to
which the test agrees or correlates with a criterion set up as
an acceptable measure. The criterion is always available at the
time of testing. It is applicable to tests employed for the di-
agnosis of existing status rather than for the prediction of
. future outcome:QUALITIES OF GOOD RESE.
Mlustration : a he
For instance, a researcher wishes to validate a Biology
achievement test he has constructed. He administers this test
to a-group of Biology students. The result of this test is cor-
related with an acceptable Biology test which has been pre-
viously proven as valid. If the correlation is “high,” the Biology
test he has constructed is valid.
Predictive validity. Predictive validity, as described by
Aquino and Garcia (1974), is determined by showing how well
predictions made from the test are confirmed by evidence
gathered at some subsequent time. The criterion measure
against this type of validity is important because the outcome
of the subjects is predicted.
Wustration
For example, the researcher wants to estimate how well
a student may be able to do in graduate school courses on the
basis of how well he has done on tests he took in the under-
graduate courses. The criterion measure against which the
test scores are validated and obtained are available after a
long period of interval.
Construct validity. The construct validity of a test is the
extent to which the test measures a theoretical construct or
trait. This involves such tests as those of understanding,
appreciation and interpretation of data. Examples are intel-
ligence and mechanical aptitude tests.
Illustration
For instance, a researcher wishes to establish the
validity of an IQ (Intelligence Quotient) using Wechsler
Adult Intelligence Scale (WAIS). He hypothesizes that
students with high IQ also have high achievement, and
those with low 1, low achievement. He therefore administers
both WAIS and achievement tests to two groups of students
with high and low IQ, respectively. If the results show that
those with high 1Q have high scores in the achievement tests
and those with low IQ have low scores in the achievement
tests, the test is valid. meMETHODS OF RESEARCH AND THESIS WRITING
Reliability
Reliability means the extent to which a “test is depend-
able, self-consistent and stable.” (Merriam, 1975). In other
words, the test agrees with itself. It is concerned with the
consistency of responses from moment to moment. Even if a
person takes the same test twice, the test yields the same
tesults, However, a reliable test may not always be valid.
For instance, a research student receives a grade of 1.25
in Méthods of Research. When ask by his friends, he says his
grade is only 1.5. In statistical sense, the story is reliable for
cis consistent, but not valid because there is no veracity or
truthfulness of the story. Hence, it is reliable but not valid.
Likewise, a reliable test or measuring instrument is not al-
ways valid even if it may be reliable.
Methods in Testing the Reliability
of Good Research Instrument
‘There are four methods in estimating the reliability of
good research instrument. These methods are as follows: (1)
test-retest method, (2) parallel-forms method, (3) split-half
method, and (4) internal-consistency method.
Test-retest method. The same research instrument is
administered twice to the same group of subjects and the
correlation coefficient is determined. The limitations of this
method are: (1) when the time interval is short, the subjects
may recall his previous responses and this tends to make thé
correlation coefficient high; (2) when the time interval is long,
such factors as unlearning, forgetting, among others, may
occur and may result in low correlation of the test; and (3)
regardless of the time interval separating the two administra-
tions, other varying environmental conditions such as noise,
temperature, lighting, and other factors may affect the corre-
lation coefficient of the research instrument.
A Spearman rank correlation of coefficient or Spearman
rho is a statistic used to measure the relationship between,
paired ranks assigned to indiyidual scores on two variables.
Thus, this is tised to correlate the scores in a test-retest method.
To obtain the valué of Spearman rho(r,) consider this formulaQUALITIES OF GOOD RESEARCH INSTRUMENT. n
@D
where r, = Spearman rho
ED? =sum of the squared differences between ranks
N_ = Total number of cases
‘To apply the foregoing formula (5.1), the steps are as
follows:
Step 1. Rank the scores of subjects from the highest to
the lowest in the first set of administration (X),
and mark these ranks as R,. The highest score
receives the rank of 1; the second highest, 2;
third highest, 3; and so on.
Step 2. Rank the second sct of scores (¥) in the same
manner as in Step 1 and mark as R,
Step 3. Determine the difference in ranks for every pair
of ranks.
Step 4. Square each difference to get D®
Step 5. Sum the square difference to find SD?
Step 6. Compute Spearman rho (r,) by applying the
formula (5.1),
For example, fourteen respondents are used as pilot
sample to test the reliability of an achievement test in Biology
Table 5.1 shows respondents’ scores and reliability coefficient
in two administrations using the Spearman rho. These four-
teen respondents are not the subjects of the study
Table 5.1, Spearman rho Computation of the First
and Second Administration of Achieve-
ment Test in Biology
ee ce
Respondents) | Xij0 Wie i Ag. AyD” | D;
ee ee ee
1 90' 270 9 20. “gg. S85 30.85
2 43° 31° «13.0 «125 0.50.28
3 S45 19: 6.57.3... a5 , 12:95
4 86 70 45 75 -30 9.0078 METHODS OF RESEARCH AND THESIS WRITING
55 4311.0 10.5 0.50.25
17 7 85 75 10 1.00
a as: 45° 30. 40
91. $8 “10 -10- bo 0:00
40. 21 140 125 15 2.25
75, 770° 7100 76 “25 6:25
8 68) 6450 20 25 6.25
2 89 75 3.0 45 -15 2.05
48 30 120 140 -20 4.00
17 43° 85- 10.5 .-20° - 4.00
EGBESeax24H
(High relationship)
‘The Spearman rho valuc obtained is 0.82 which denotes
high relationship. Thus, the respondents who got high score
in the first administration also got high score in the second
administration and those who got low score in the second
administration also got low score in the second administra-
tion. This means that their responses are reliable. According
to Garrett (1969), “a highly reliable test is always a valid
measure of some functions.” Since the coefficient of correlation
is reliable and very high, the research instrument is both valid
and reliable. |
Paraliel-forms method. Parallel or equivalent forms of a
test may be administered to the group of subjects, and the
paired observations correlated. “In estimating reliability by
the administration of parallel or equivalent forms of a test,
criteria parallelism is required.” (Ferguson and Takane, 1989).
The two forms of the test must be constructed so that the
content, type of item, difficulty, instructions for administra-
tion, and many others, are similar but not identical,
For instance, the item, “Convert 7,000 grams to kilo-
_ grams” in Form Ais parallel to “Convert 7 kilograms to grams”
ee~ QUALITIES OF GOOD RESEARCH INSTRUMENT 75
in Form B. Moreover, these two forms should have approxi-
mately the same average and variability of scores.
‘The correlation between the scores obtained on paired
observations of these two forms represents the reliability
coefficient of the test. If the coefficient correlation (r) value
obtained is high, the research instrument is reliable.
Split-half method. The test in this method may be admin-
istered once, but the test items are divided into two halves.
The common procedure is to divide a test into odd and even
items. The two halves of the test must be similar but not
identical in content, number of items, difficulty, means and
standard deviations. Each student obtains two scores, one on
the odd and the other on the even items in the same test. The
scores obtained in the two halves are correlated. The result
is a reliability coefficient for a half test. Since the reliability
holds only for a half test, the reliability coefficient for a whole
test may be estimated by using the Spearman-Brown formula
This formula is:
aoe (5.2)
1+ Thy
where ry, is the reliability of a whole test; and rp, reli-
ability of 2 half test.
For instance, a test is administered to ten students as
pilot sample to test the reliability coefficient of the odd and
even items. The results are shown below in Table 5.2
Table 5.2. Computation of Reliability Coefficient of .
Odd and Even Items
‘Scores Ranks Differences
we 2¥ RY R, oO Ce
Student (Odd) (Even)
1 a Ogata Lot 22a
2 Pio the ml Os wei eo £0
3 ee) 6° Ts: 15. 225.
4 35°40 5 4°10 1.0
5 48-55 3° 155 15 4 22520 METHODS OF RESEARCH AND THESIS WRITING
6 or 24 10 9.505 0.25
7 25, she 75 6 15 2.25
8 5051 2 3. 10. 20
9 28 3B 4 Se t08 1d
10 5555 7 To 06) 0.25,
Total 16.50
GED? 2rht)
ny 21 - Se aint
m NUON, | 8 | ent
= 1 - 8116.50) _ 2(0.90)
108 - 10 1 + 0.90
The = 0-90 Tuy = [0.95] (Very high
relationship)
Based on the foregoing example, the reliability of a half
test (r,,) is 0.90 and the reliability of a whole test is 0.95. Since
the reliability of the whole test (F,,,) obtained is very high, the
whole test is reliable.
split-halfis applicable for research instrument not highly
speeded. If the research instrument includes easy items and
the subject is able to answer correctly all or nearly all items
ccihin the time limit of the test, the scores on the two halves
vould be about similar and the correlation would be closed to
+1.00 (perfect positive correlation).
Internal-consistency method. This method is used
with psychological tests which consist of dichotomously
Scored items. The examinee cither passes or fails in an
item. A rating of 1 (one) is assigned for a pass and for 0 (zero)
2 failure. The method of obtaining reliability coefficient
in this method is determined by Kuder-Richardson Formula
20. This formula is a measure of internal consistency or
homogeneity of the measuring instrument. The formula
is:
ce
N_SD*= 2piaiy (5.3)
N-1 sD?
ores
Nis the number of items; SD? is the variance of S205) QUaLITiEs OF GOOD RESEARCH INSTRUMENT al
K-%
; 3 ; and p; q; is the product of the
proportion passing or failing in item i
‘The proportion of individuals passing item i is denoted
by the symbol p;, and the proportion failing, by q;, where gi
=1-p, . Ifall items are perfectly corrected, a situation which
can only arise when all have the same difficulty, r,, = 10.
‘The steps in applying the Kuder-Richardson formula 20
are as follows:
Step 1. Compute the variance (SD*) of the test scores
for the whole group.
Step 2. Find the proportion passing each item (pi) and
the proportion failing each item (q,). For in-
stance, twelve of the fourteen students passed
2
14
0.86); and two students failed in Item 1, (q;
= 2. - 0.14 or q, = 1p) 1 ~ 0.86 = 0.14).
or got the correct answer for Item 1, (p,
Step 3. Multiply p and q; for each item, ie., 0.86 x 0.14
0.1204; and sum for all items. This gives the
Ds value.
Step 4. Substitute the calculated values in formula 5.3.
For illustration purposes, consider the following exam-
ple: Suppose a test of 10 items has been administered to a
group of fourteen subjects. Table 5.3 shows the computation
of Kuder-Richardson formula 20.
Table 5.3 Computation of Kuder-Richardson
Formula 20
Cee ee ee
Students
hems 1234567 891011121315 fT BG PA
1 2212111111110 0 12 086 0.14 0.1204
2 1211111111110 0 12 086 0.14 0.1208.
grove athe Por 201 Vr 00" O° 110.79 0.12 0.1659
4212122112112 1-000 0 100.71 0.28 0.1988
BoD deed aoe died iO, 0-020, 10,071 028" 0-1ane
6 1L112112121 1000 0 10071 0.28 0.1988
7 WY at 1a 0-0-0: 0 0 06s oie ORS0Y82 METHODS OF RESEARCH AND THESIS WRITING
1101101000 8 957 0.43 0.2451
2111000000 8 os7 0.43 0.2451
10 0110100000 4 029 071 0.2089
2 22 100000 5 029 0.71 0.2089,
9 0 1.9296
i 9296
By
7
9
10
sp? = 2X xX
N-1
ot = 88.8092
Wt wW-1
=[ 671 sp? = [ 6.83
Kuder-Richardson Formula 20 Computation
N 10titems)
SD? = 6.83,
10_}1,6.83 — 1.9296 ) pig, = 1.9296,
w-1 6.93
(2.12(-4.9004 | a
6.83 : %
1.11(0.7174816)
0.796 or | 0.80 High relationship
0QUALITIES OF GOOD RESEARCH INSTRUMENT. 83
‘The reliability coefficient (r,,) value obtained is 0.80,
which means that the test is reliable, Moreover, 3 can be
sleaned that responses of subjects on Table 5.3 are internally
consistent where students 13 and 14 failed the exsiest item
(Item 1) with a difficulty valuc of 86 percent aud failed the
rest of the items. On the other hand, students 2, 4, 5,7, and
8 passed Items 1 to 9 but failed the most difficult item (Item
10) with a difficulty value of 0.29 or 29 per cent. According
to Ferguson and Takane (1989), “If an individual obtains a
passing score on the casier items of the test and fails in the
more difficult ones, his performance contains no irconsisten-
cies. If all the individuals taking a test obtain their scores in
this way, and no inconsistencies are present in the response
pattern, the response may be spoken of as an internally con-
sistent pattern.”
Interpretation of Correlation Coefficient Value
To interpret the correlation coefficient value (r) obtained,
the following classifications may be applied
An x from 0.00 to + 0.20 denotes negligible correlation
An r from + 0.21 to + 0.40 denotes low or slight corre-
lation
An r from + 0.41 to + 0.70 denotes marked or moderate
relationship
Anr from + 0.71 to + 0.90 denotes high relationship
Anz from + 0.91 to + 0.99 denotes very high correlation
An r from + 1.00 denotes perfect correlation
Usability
Usability means the degree to which the research instru-
ment can be satisfactorily used by teachers, researchers,
supervisors and school managers without undue expenditure
of time, money, and effort. In other words, usability means
practicability.
Factors That Determine Usability
There are five factors that determine usability, namely:
(1) ease of administration, (2) ease of scoring, (3) ease of4
METHODS OF RESEARCH AND THESIS WRITING
interpretation and application, (4) low cost, and (5) proper
mechanical make-up.
1
3.
Ease of administration. To facilitate the administra-
tion of a research instrument, instructions should
be complete and precise. As a rule, group tests are
easier to administer than individual tests. The former
is easier to admihister because directions are given
only once and the instrument is simultaneously
administered to a group of students, thus saving
time and energy on the part of the examiner or
researcher.
Ease of scoring. Ease of scoring a rescarch instru-
ment depends upon the following aspects
2.1 Construction of the test in the objective
type;
2.2 Answer keys are adequately prepared; and
23 Scoring directions are fully understood.
Moreover, scoring is easier when all subjects
are instructed to write their responses in one col-
uma in numerical form or word and with separate
answer sheets for their responses.
Ease of interpretation and application. Results of
tests are easy to interpret and apply if tables are
provided. All scores must be given meaning from the
tables of norms without the necessity of computa-
tion, As a rule, norms should be based both on age
and year level, as in the case of school achievement
tests. It is also desirable if all achievement tests
should be provided with separate norms for rural
and urban subjects as well as for learneis of various
degrees of mental ability.
Low cost. It is more practical if the test is low cost,
material-wise. It is more cconomical also if the
research instrument is of low cost and can be reused
by future researchers.
Proper mechanical make-up. A good research in-
stryment should be printed clearly in an appropri-
ernieI. Multiple Choice:
options.
1
QUALITIES OF GOOD RESEARCH INSTRUMENT
85
ate size for the grade or year level for which the
> instrument is intended. Careful attention should be
given to the quality of pictures and illustrations on
the lower grade subjects of the study.
EXERCISES
Choose the best answer among the
Write only the number of the chosen answer on the
blank at the right column.
The interpretation of 0.92 correla-
tion value obtained by X & Y is
1
3
4,
moderate relationship
high relationship
Very high relationship
Nogligible correlation
Which of the correlation (r) value
below has slight relationship?
1
3
4
0.19
0.45,
0.25
0.20
The extent which the topic or sub-
stance of a re;
arch instrument is
truly a representative of the sub-
stance of the course is
1;
2
3.
4
construct validity
predictive validity
concurrent validity
content validity86
METHODS OF RESEARCH AND THESIS WRITING
“4. The degree: to which the research
instrument measures what it pur-
ports to measure is
1. Reliability
2. Validity
2
Objectivity ~
Usability
-
5. The degree to which the test agrees
with or correlates with a criterion
set up as an acceptable measure is
1. Concurrent validity
2. Construct validity
3. Predictive validity
4. Content validity
6. The extent to which the research
instrument measures a theoretical
trait is
1. Concurrent validity
2. Construct validity
3. Predictive validity
4. Content validity
It is concerned with the consistency
of responses froxi moment to mo-
ment even if the subjects take the
same instrument twice is
1. Validity
2. Objectivity
3. Usability
4. Reliability
8. A method of testing reliability of a
“research instrument in which the
4.enters
| QUALITIES OF GOOD RESEARCH INSTRUMENT
scores of the first and second admin-
istration of the test are determined
by correlation coefficient is
1. parallel-forms
2. internal consistency
3, test-retest
4, © split-half
9. A method of reliability in which a
subject receives a point of one or zero
for each item is
1, split-half
2. internal consistency
3. test-retest
4. parallel-forms E
10. The most practical quality of a good
research instrument is
1. Usability
2. Objectivity
3. Reliability
4. Validity
11. Split-half method is determined
by
1, Kuder-Richardson Formula 20
2. Spearman rho
Spearman-Brown formula
9.
10,
11
87
3.
4, Pearson product moment coefficient of correla-
tion
12. The formula of arithmetic mean is
12.88 METHODS OF RESEARCH AND THESIS WRITING
rye)
T+ ty
2
ender
NS -
4 Xt
N
13. A method of estimating the reliabil- 13.
ity of a research instrument by cor-
relating the scores in the odd and
even items is
1. split-half
2. test-retest
3. parailel-forms
4. internal consistency
14. Internal consistency method is de- 14,
termined by
1 Pearson Product-Moment Correlation Coefficient.
2. Kuder-Richardson Formula 20
3. ~ Spearman rho
4. Spearman-Brown Formula
15. Which of the following does not be- 15. ___
long to the group?
1. Validity
2. Objectivity... -
3. Reliability’
4. Usability
16. - Test-re
ethod is determined by 16. : 13
: i
1. _: Pearson Product-Moment Correlation Cooffecient i
Spearman-Brown, FormulaWw.
18.
19.
20.
3. _ Spearman rho
4. Kuder-Richardson Formula 20
‘The formula of Kuder- Richardson”
Formula 20 is
The correlation coefficient (r) value
vf 0.75 is interpreted as
1. high relationship
2. moderate relationship
3. slight correlation
4. very high relationship
The degree to which the research
instrument gives proper mec-
hanical make-up to the researcher
is ‘
1, Reliability
2. Usability
3. Validity
4. Objectivity
‘The formula of variance is
= -6ED?
1.
Ne-N
It
18.
19.
20.20 METHODS OF RESEARCH AND THESIS WRITING
Bee
1¥ Th
=x
ay
2 LX-X
Noi.
Il. Essay
1, How do you account this statement: “A valid test
is always valid but a’ reliable test is not always
valid.”
Which of the qualities of a good research instrument
is most practical? Why?
s
3. How are the parallel-forms method of estimating
reliability of a. test constructed?
4,” Discuss briefly each of the 5 factors that determine
the usability of a test?
5. How do you administer test-retest method of testing
u reliability of a test?
IIL. Problem Solving
1. Using the data below, find out if the research instru-
ment is reliable, applying the test-retest method
administered to 14 respondents as pilot sample.
First Second
Respondent Administration Administration
1 90 91
2 85 85
3 77 78
4 55 55
5 Baie 86.
6 )
% 91 ol.
8
Sits he = 80,QUALITIES OF GOOD RESEARCH INSTRUMENT,
9
10
a
12
13
14
48
55
80
80
90
50
53
53
80
81
89
53
31
2. Find out if there is internal consistency in the re-
sponses of the 10 students as pilot sample in a 15-
item test in Psychology.
Items
1
2
3
4
5
6
7
8
9
10
SBN
12
13
14
15
Total
1
1
4,
1
0
0
oO
0
0
0
0
0
0
0
0
0
3
Answer to Exercises
apoan
Hee Ow
6.
1
8.
9.
10.
Hoye Hb
wl oooo COP HEHEHE HEN
Students.
45 6
Tea,
i eigrtaet
laid
a ealeeast
1a 2
fe
Dt #1
101
106
10° @
10; 0
100
000
00 0
000
8
1.
12.
13.
14.
15,
eoocc eo COO OH OHHH A
wn loooococ ooo oH EHD
16.
17.
18.
19.
20.
aloooccocCCOOHHHHHSe
aye RO