0% found this document useful (0 votes)
11 views10 pages

Research Measurement Concepts Explained

Chapter 4 discusses measurement concepts in research, emphasizing the importance of designing a proper measurement system to assign numbers or labels to attributes. It covers the identification of variables, measurement issues, the development of measurement scales, and criteria for good measurement, including reliability and validity. The chapter also addresses potential errors in measurement and the need for sensitivity and relevance in research instruments.

Uploaded by

daothao2005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views10 pages

Research Measurement Concepts Explained

Chapter 4 discusses measurement concepts in research, emphasizing the importance of designing a proper measurement system to assign numbers or labels to attributes. It covers the identification of variables, measurement issues, the development of measurement scales, and criteria for good measurement, including reliability and validity. The chapter also addresses potential errors in measurement and the need for sensitivity and relevance in research instruments.

Uploaded by

daothao2005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Chapter 4: Measurment concepts in research

4.1 Measurement
Before collecting data, a researcher must design a proper measurement system.
 Measurement = assigning numbers or labels to attributes of people, objects, or
concepts according to rules.
 It can represent attributes quantitatively (numbers) or qualitatively (categories).
 In research, measurement tells us how much of a particular attribute an object or
person has.
So, measurement is basically turning abstract things (like satisfaction, motivation) into
something we can record and analyze.
4.2 Identifying and Deciding on the Variables to Be Measured
First, the researcher must clarify what is to be measured:
 Concepts: general, abstract ideas (e.g. motivation, satisfaction, education).
 Constructs: more complex concepts, often made up of several simpler concepts,
created to simplify and explain complicated situations (e.g. “job satisfaction” as a
combination of pay, environment, relationships, etc.).
A constitutive definition of a concept or construct:
 States the central theme of the study.
 Defines the boundaries of the research.
 Helps form clear research questions.
Example: “Education system in Rajasthan” is too broad. The researcher must define whether
it’s primary, secondary, higher education, adult education, etc. Only then can measurement
decisions be made properly.
4.3 Research Measurement Issues
When measuring, the researcher must address several issues:
1. Does the concept allow nominal/ordinal levels or higher (interval/ratio)?
2. Is the concept discrete or continuous? (Continuous allows more powerful statistics.)
3. How many indicators are needed?
o Simple concepts may use one indicator.

o Abstract concepts (like motivation) need multiple indicators.

4. Are there valid and reliable measures available?


5. Is the measurement free from errors such as invalidity and unreliability?
6. Measurement validity is separate from internal and external validity (those belong to
research design).
7. Does the researcher know what data exist in various agencies and sources, and use
them correctly?
These decisions shape how well a concept is actually captured.
4.4 Need for Development of Measurement Scales
A scale is a set of numbers or symbols with rules for assigning them to units being studied.
 Some things are easy to measure quantitatively (sales, profit, employee productivity).
 Others are harder, like motivation, attitudes, acceptance of a new product.
Difficulties:
 Respondents may struggle to express their feelings in words.
 Scales may fail to “pull out” true responses.
 Respondents may be unwilling to reveal their true opinions.
To address these:
 The researcher must build trust and cooperation.
 Explain clearly what information is needed and why.
 Use appropriate scaling techniques (e.g. to measure customer retention, attitudes,
etc.).
So, good scale development is critical for measuring subjective, psychological, or intangible
variables.
4.5 Measurement Scales
The design of a measurement scale depends on:
1. The objective of the study.
2. The statistical techniques the researcher intends to use.
3. Whether the goal is simply to classify, rank, or compare and predict.
Measurement = observing and recording using one of four levels:
1. Nominal
2. Ordinal
3. Interval
4. Ratio
4.5.1 Nominal Scale
 Qualitative, no order.
 Data are placed in categories without any ranking or distance.
 Numbers are just labels (e.g. 1 = public sector, 2 = private, 3 = entrepreneur, 4 =
jobless, 5 = others).
 Mathematical operations (add, subtract, etc.) are not meaningful.
 Only basic statistics:
o Counts/Frequencies, Percentages, Mode, Chi-square, contingency
measures, etc.
Nominal scales simply tell which category each case belongs to.
4.5.2 Ordinal Scale
 Qualitative with order.
 Objects can be ranked (1st, 2nd, 3rd…) but the distance between ranks is unknown.
 Gives information on relative position, not exact differences.
 No equal intervals, no true zero.
Examples:
 Ranking students by academic performance (1st, 2nd, 3rd…).
 Income groups (e.g. low, medium, high) ordered from low to high.
Statistical operations possible:
 Median, Mode, Range, Percentiles, Rank correlation, non-parametric tests.
Ordinal scale answers: “Who is higher or lower?” but not “How much higher?”
4.5.3 Interval Scale
 Has all properties of nominal + ordinal scales.
 Equal intervals between points on the scale.
 But no true zero (zero is arbitrary, not “absence”).
Examples:
 Temperature in °C or °F. Differences between 30° and 50° = differences between 15°
and 35°.
 A 1–7 satisfaction scale (dissatisfied → satisfied) with equal distances between points.
 Score ranges (0–10, 10–20, etc., with equal class intervals).
Statistics possible:
 Mean, Standard Deviation, Correlation, Regression, ANOVA, t-test, F-test, factor
analysis, etc.
 Coefficient of variation is not appropriate (no true zero).
Interval scales allow both order and meaningful differences, but not ratio comparisons like
“twice as much”.
4.5.4 Ratio Scale
 Has all properties of nominal + ordinal + interval.
 Has an absolute zero (true absence of the attribute).
 Supports all mathematical operations (add, subtract, multiply, divide).
Examples:
 Height, weight, distance, money.
 Distance: not only is the difference between 3 and 6 miles equal to difference between
6 and 9 miles, but 9 miles = 3 × 3 miles.
Statistics: all parametric methods plus ratio-based comparisons.
Note: Interval and ratio data are often called parametric, while nominal and ordinal are
non-parametric.
4.6 Criteria for Good Measurement
Three main criteria:
1. Reliability – consistency.
2. Validity – accuracy (does it measure what it should measure?).
3. Practicality – feasibility (economical, convenient, interpretable).
4.7 Reliability
4.7.1 Meaning of Reliability
Reliability = the degree to which a measurement is consistent and dependable.
 If the same construct is measured repeatedly under the same conditions, it should give
similar results.
 A reliable instrument has stable, reproducible outcomes.
Examples:
 A tea/coffee vending machine giving the same quantity each time → reliable.
 A stress test that gives similar scores when repeated on the same person (assuming no
real change) → reliable.
Sources of low reliability:
 Poor data collection methods.
 Respondents not understanding questions.
 Random errors in administration.
Reliability is not directly measured, but estimated, usually via correlation between repeated
measures.
4.7.2 Methods of Estimating Reliability
External Consistency Procedures – compare results across repeated or parallel
measurements.
[Link].1 Test–Retest Reliability
 Administer the same test to the same sample at two different times.
 Correlate the two sets of scores.
 High correlation = high reliability.
Problems:
 Hard to locate or get cooperation from all respondents again.
 Real changes in respondents or environment can affect scores.
 Memory effect (remembering previous answers) and practice effect (improvement
due to repetition).
 Absence of some participants in the retest.
[Link].2 Parallel Forms Reliability
 Use two equivalent but different versions of a test measuring the same construct.
 Administer both to the same group, correlate the scores.
 Very rigorous but hard to create truly equivalent forms.
Because of the difficulty, many researchers prefer internal consistency methods.
Internal Consistency Procedures – look at consistency within a single test.
Example: Split the test into two halves and see if both halves produce similar results.
Issues:
 Different ways of splitting items may give different reliability estimates.
Methods:
[Link].1 Split-Half Reliability
 Randomly divide the test items into two sets.
 Administer the whole test once, calculate scores for each half, correlate the two.
 Adjust correlation using Spearman–Brown formula to estimate full-length reliability.
[Link].2 Kuder–Richardson (KR-20)
 Used when items are dichotomous (0/1, true/false).
 Computes reliability by looking at item difficulties and variances.
 Essentially the average of all possible split-half coefficients.
[Link].3 Cronbach’s Alpha (α)
 Generalization of KR-20 for items not scored 0/1 (e.g. Likert scale).
 Treats alpha as the mean of all possible split-half reliabilities.
 Ranges from 0 to 1 (higher = better internal consistency).
Used widely for scales with multiple items (e.g. attitude scales).
4.8 Validity
Validity = the degree to which an instrument measures what it is supposed to measure.
 A valid measure reflects true differences among individuals on the construct of
interest.
 Validity is typically expressed as a correlation with a criterion (validity coefficient).
Examples:
 An exam that only tests memorization, not understanding → questionable validity.
 Measuring morale only through absenteeism → invalid, as absenteeism may be due to
other factors (illness, family issues).
Types of validity:
1. Content validity
2. Criterion-related validity
o Concurrent

o Predictive

3. Construct validity
o Convergent

o Discriminant

4. Face validity
5. Internal validity
6. External validity
4.8.1 Content Validity
 How well the scale covers all relevant aspects of the concept.
 Are the selected variables appropriate and adequate?
Example:
If evaluating school facilities, measuring only hoardings, alumni meets, and canteen snacks
has no content validity. Relevant items would be classrooms, labs, water, toilets, teachers,
playground, etc.
4.8.2 Criterion-Related Validity
 How well the measure correlates with an external criterion or known standard.
Two types:
[Link] Concurrent Validity
 Predictor and criterion are measured at the same time.
 Example: A test of nervousness should accurately reflect a person’s current level of
nervousness.
[Link] Predictive Validity
 The measure is used to predict future performance on a criterion.
 Example: A selection test that predicts future job performance; or a builder focusing
repairs that will attract future tenants.
4.8.3 Construct Validity
 Degree to which a test truly represents the theoretical construct and relates to other
variables as theory predicts.
 Focuses on why people behaved in a certain way, not just how.
Example:
Not just whether a consumer bought the product, but why they didn’t buy it.
Two subtypes:
[Link] Convergent Validity
 High correlation among different measures of the same concept.
[Link] Discriminant Validity
 Low correlation between the measure and different, unrelated constructs.
Example:
A scale measuring the tendency to stay in low-cost hostels should correlate with related traits
(self-confidence, low need for status), but not with unrelated constructs (brand loyalty,
aggressiveness).
Good construct validity = high convergent + low discriminant correlations.
4.8.4 Face Validity
 Superficial, judgment-based validity: does it look like it measures what it should?
 Based on expert/researcher judgment.
 Weakest form of validity but still useful as a first check.
4.8.5 Internal Validity
 Concerns causal logic: can we confidently say that changes in the dependent variable
were caused by the independent variable?
 Higher internal validity = stronger basis for causal inference.
Threats to internal validity include:
1. Confounding – extraneous variables not controlled.
2. Selection bias – groups differ systematically before treatment.
3. History – external events influence outcomes.
4. Maturation – natural changes in subjects over time.
5. Repeated testing – practice or memory effects.
6. Instrument change – changing measurement tools mid-study.
7. Regression to the mean – extreme scores naturally move toward average.
8. Mortality – dropout of participants in a non-random way.
9. Diffusion of treatment – treatment “leaks” to control group.
10. Compensatory rivalry / resentful demoralization – control group alters effort
(works extra hard or becomes demoralized).
11. Experimenter bias – researchers unknowingly influence participants differently in
different groups.
4.8.6 External Validity
 The degree to which results can be generalized to other people, settings, times, etc.
 If samples are small, local, or biased (e.g. only volunteers), external validity may be
weak even if internal validity is strong.
4.9 Practicality
A good measure should be:
 Economical (not too costly),
 Convenient (easy to administer),
 Interpretable (results easy to understand),
 Not too long or complex, and not requiring extremely specialized personnel to
administer.
4.10 Sensitivity
Sensitivity = the instrument’s ability to detect small differences or changes.
 Scales with only “agree/disagree” are less sensitive.
 Adding more categories (strongly agree, mildly agree, etc.) increases sensitivity and
captures subtle variations.
4.11 Generalizability
This refers to how flexibly the data collected with a particular scale can be interpreted across
different designs, respondents, and conditions.
 Multiple-item scales with wide applicability and robust interpretation have higher
generalizability.
4.12 Relevance
Relevance = appropriateness of a scale for measuring a particular variable.
 Often expressed as:
Relevance = Reliability × Validity
 If either reliability or validity is low, relevance is low, even if the other is high.
4.13 Errors in Measurement
Two main types:
 Measurement invalidity – the measure doesn’t correctly capture the concept.
 Unreliability – the measure is inconsistent across repeated uses.
4.13.1 Respondent-Associated Errors
Most research depends on respondents. If they don’t cooperate or tell the truth, errors arise.
Two kinds:
 Nonresponse error – some units or items are missing.
 Response bias – respondents misrepresent the truth.
4.13.2 Nonresponse Errors
Occurs when:
 Some people don’t participate at all, or
 They skip certain questions.
Results may be biased if nonrespondents differ systematically from respondents (e.g. refusing
to answer sensitive income questions).
4.13.3 Response Bias
Occurs when respondents:
 Intentionally or unintentionally give false or misleading answers.
 Do this to avoid embarrassment, appear knowledgeable, or protect privacy.
4.13.4 Errors Associated with Instrument
Arise from:
 Poor questionnaire design.
 Inadequate space for answers.
 Ambiguous or complex wording.
 Improper sample selection.
These can confuse respondents and cause wrong answers.
4.13.5 Situational Errors
Caused by context of the interview:
 Poor rapport between interviewer and respondent.
 Presence of a third person.
 Interview location (home vs noisy public place).
 Lack of assurance of confidentiality.
These factors may make respondents hold back or alter their answers.
4.13.6 Measurer as Error Source
Errors caused by the researcher/interviewer:
 Wrong coding.
 Faulty tabulation and statistical calculations.
 Rewording or shortening respondents’ answers incorrectly.
 Using wrong statistical techniques.
 Influencing respondents with body language (smiles, frowns, tone).
All these distort the data and harm measurement quality.

You might also like