0% found this document useful (0 votes)
7 views9 pages

Understanding Objectivity and Reliability in Testing

Uploaded by

khadija
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views9 pages

Understanding Objectivity and Reliability in Testing

Uploaded by

khadija
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Objectivity (Simplified)

o Objectivity means keeping your personal opinions / biases out of the test scoring.

o In psychology testing, a test is objective if different trained evaluators give similar


scores for the same responses.
o Tests with a clear scoring key (right/wrong answers) are usually more objective than
tests that rely on personal judgment.
o Think back to school: some teachers were stricter than others, and your grade could
change depending on who marked your paper. That’s a lack of objectivity.

 To make tests more objective:


Ø Use clear scoring guidelines so personal bias is reduced.
Ø Check interrater agreement, which measures how much different scorers agree with
each other.

 The method to calculate this depends on:


1. The type of data (Nominal, Ordinal, Interval).
2. How many people are scoring the test.

 Example 1: Multiple-choice test


• A personality test asks 20 questions with fixed answer choices.
• Two different psychologists score the test using the answer key.
• Both give the same results, showing high objectivity because personal opinions don’t
affect the score.
 Example 2: Essay-based assessment
• A creativity test asks participants to write a story.
• If scorers don’t have a clear rubric, one might give high marks for imaginative ideas
while another focuses on grammar.
• Scores vary depending on the scorer, showing low objectivity.

Reliability
 Reliability is about how consistent a test’s results are.
 A reliable test gives similar results every time it is used under the same conditions.
 Main ways to measure reliability:
 Internal consistency (Cronbach’s Alpha)
o Checks if items on a test measure the same thing.
o Uses the average correlation between items.

o Works best for unidimensional constructs (one clear concept).

o Software can calculate this automatically.

o Rule of thumb: alpha ≥ 0.70 is acceptable; for intelligence tests, aim for 0.90–
0.95.

Internal Consistency (Cronbach’s Alpha)


o A job satisfaction survey with 10 questions all meant to measure overall satisfaction.
Cronbach’s Alpha checks whether the questions are consistently capturing the same
underlying concept.
o A depression questionnaire with items like “I feel sad,” “I have low energy,” and “I
don’t enjoy activities.” Alpha shows if these items all measure the same construct
(depression).

 Internal consistency basically asks: “Do all the questions on this test really measure
the same underlying construct?”

 So in your example:

 If it’s a logic test, all items should test reasoning, pattern recognition, or problem-
solving.
 If you suddenly include a simple multiplication question, it doesn’t really fit —
because that’s basic arithmetic, not logic. That would lower the internal consistency.

 Think of it this way:

 High internal consistency = all items are pulling in the same direction.
 Low internal consistency = some items are “off-topic,” measuring something
different.

 Strong internal consistency (all items measure logic):

1. “If all cats are animals and some animals are pets, can we conclude all cats are
pets?” → This tests reasoning.
2. “Complete the sequence: 2, 4, 8, 16, __ ?” → This tests pattern recognition.

👉 Both are about logical reasoning/problem-solving, so they support the same construct.

 Weak internal consistency (one item is off-topic):

1. “If all squares are rectangles and some rectangles are red, can we conclude some
squares are red?” → Logical reasoning.
2. “What is 37 × 12?” → This is just arithmetic, not logic.
Test–Retest Reliability
 Give the same test to the same participants at two different times.
 Measures stability over time.
 Best for traits that don’t change, like knowledge / aptitude.
 Reliability may drop if the time between tests is long.

1. An IQ test given to the same group of students in January and then again in March.
Scores should be stable if the test is reliable.
2. A typing speed test administered on Monday and again two weeks later. If the test is
consistent, results should be similar (unless participants improved with practice).

Parallel Forms Reliability


 Two equivalent versions of a test are given to the same participants.
 Checks if both versions give similar results.
 Useful when practice effects could affect scores, but creating parallel forms can be
hard.

 Parallel forms reliability = two different versions of the same test should give
similar results for the same people.

 Why?

 Imagine if people could memorize answers or get better just by practicing. If you give
them the same test again, scores might improve just because they’ve seen the
questions before — not because they actually learned more.
 To avoid this, you make two equivalent versions (same difficulty, same content area,
just different questions). If both versions measure the construct equally well, scores
should come out very similar.

🔹 Good example (strong parallel forms):

 TOEFL English test has Form A and Form B. Form A asks: “What’s the main idea
of this passage about global warming?”
 Form B asks: “What’s the main idea of this passage about renewable energy?”
👉 Both measure reading comprehension in English — just with different texts. A
student should score about the same on both.

o (my example) you know at university, if you fail one test, they let you do another
test for the same course, so the equivalency is the same, it's just the questions
won't be to check your understanding.

🔹 Bad example (weak parallel forms):

 Form A of a math test has mostly algebra questions.


 Form B has mostly geometry questions.
👉 Even if a student does great in algebra but struggles with geometry, their scores will
differ. That means the two forms aren’t really equivalent, so reliability drops.

1. A math placement test with two different but equivalent sets of questions (Form A
and Form B). Students should score similarly on both.
2. A language proficiency test (English grammar) with two parallel versions. If both
versions give nearly the same results, the test is reliable.

Validity (Simplified)
o Validity is about whether a test really measures what it is supposed to measure.

o It’s the most important feature of a test because valid scores let us make
meaningful interpretations.
o Validity isn’t permanent—it’s an ongoing process of collecting evidence to show
that test results are accurate.
o No test is valid for every person or situation—context matters.

 Three main ways to check validity:

1. Content validity
 Are the test questions relevant to what you want to measure?
 Example: A math test should cover the topics it claims to assess.
2. Criterion-related validity
 Does the test agree with other measures of the same thing?
 Example: Scores on a new depression questionnaire correlate with scores on
an established depression scale.
3. Construct validity
 Does the test behave as expected within the theoretical framework?
 Example: A test of anxiety should correlate with stress measures and not
correlate with unrelated traits like intelligence.

Content Validity

What it is:

 Content validity checks whether the items on your test truly represent what you
want to measure.
 Example: If you’re testing math ability, your questions should cover addition,
subtraction, multiplication, etc., not reading comprehension.

How it’s evaluated:

1. Experts review the items – often called judges.


2. They give their opinion on whether each item is relevant and representative of the
domain.
3. The level of agreement among experts matters:
o 2–5 experts → 1.0 (perfect agreement needed)
o 6–8 experts → 0.83
o 9+ experts → at least 0.78 (in psychology, 0.78 is preferred)

Additional considerations:

 Who the experts are (their qualifications).


 How many experts agreed.
 Whether the test items cover the full domain of the concept.
 Whether the test was structured properly (e.g., using a table of specifications or
similar method).

Quick example:

 You’re creating a stress questionnaire.


 Experts in psychology review your items: “I feel tense,” “I have trouble sleeping,” “I
get headaches under stress.”
 If most experts agree that these items reflect stress, the test has good content validity.
 If some items are off-topic (like “I like ice cream”), content validity drops.
Content - Related Validity

What it is:

 Criterion-related validity checks whether a test relates to an outcome or criterion


that it’s supposed to predict or align with.
 Basically, it asks: “Does this test actually correspond with real-world performance or
other measures of the same thing?”

Main types & methods:

1. Known-groups method
o Compare two groups expected to differ on the trait.
o Example: A depression scale
 Group A: Individuals hospitalized for depression
 Group B: Individuals with no history of depression
 If the test scores are higher in Group A, it shows validity.
2. Concurrent validity
o Compare your test with another established test measuring the same or
related construct.
o Example: A new math knowledge test compared with an existing, validated
math test.
3. Predictive validity
o Check whether test results predict future outcomes.
o Example: An ability test predicting future job performance — high scores
today should correspond to good performance later.

Quick tip to remember:

 Known-groups → distinguishes groups now


 Concurrent → matches another test now
 Predictive → predicts the future
Construct Validity

What it is:

 Construct validity is about whether a test actually measures the theoretical


concept it’s supposed to measure.
 It’s more complex than criterion-related validity because it often involves multiple
pieces of evidence and statistical methods.

Key ways to assess construct validity:

1. Item-total correlations
o Check whether each item correlates with the overall test score.
o Example: In a depression scale, “I feel sad” should correlate with the total
depression score.
2. Factor analysis
o A statistical method to see whether items cluster together into factors
representing underlying dimensions.
o Example: A personality test might have items forming clusters for
Extraversion, Agreeableness, etc.
o Helps verify that items intended to measure the same construct actually share
a common factor.

Factor analysis analogy (vectors in space):

 Each item = a vector in multidimensional space


 Vectors aligned closely together → strong positive correlation (R ≈ 1)
 Vectors perpendicular → no correlation (R ≈ 0)
 Vectors pointing opposite directions → strong negative correlation (R ≈ ‐1)
 Factor analysis = finding latent dimensions (fewer lines) that explain these
correlations

Example:

 Black items cluster together → Factor 1


 Blue items cluster together → Factor 2
 Factor lines show which items align with which underlying construct

Quick tip to remember:

 Construct validity = does this test really measure the theoretical concept it’s
supposed to?
 Factor analysis and item correlations are tools to prove or examine this.
Tracking changes over time:

 Some constructs are expected to change. Observing whether test scores change
accordingly supports construct validity.
 Examples:
1. Vocabulary test & age: Older participants should generally score higher,
reflecting vocabulary growth.
2. Pretest-posttest for interventions: A depression scale should show lower
scores after therapy, or a language test should show improvement after a
course.

Convergent & discriminant evidence:

1. Convergent evidence:
o The test should correlate highly with other measures of the same or related
construct.
o Example: A new anxiety scale correlates strongly with an established anxiety
questionnaire.
2. Discriminant evidence:
o The test should show low or no correlation with measures of unrelated
constructs.
o Example: That same anxiety scale should not correlate strongly with, say, a
math ability test — because they measure different things.

Quick summary for memory:

 Construct validity = does the test measure the concept it’s supposed to?
 Evidence comes from:
 Item correlations / factor analysis (structure of the test)
 Changes over time (scores evolve as expected)
 Convergent evidence (matches related tests)
 Discriminant evidence (does not match unrelated tests)

Multitrait-Multimethod (MTMM) Matrix

What it is:

 The MTMM matrix is a tool to test construct validity by examining multiple traits
(concepts) measured with multiple methods.
 It helps determine whether your test is really measuring the intended trait (convergent
validity) and not mixing it up with other traits (discriminant validity).
How it works:

1. Multiple traits → At least two concepts you want to measure (e.g., Anxiety,
Depression).
2. Multiple methods → At least two ways to measure each trait (e.g., self-report
questionnaire, interviewer rating).
3. Matrix of correlations →
o Diagonal entries = reliability of each measure (e.g., Cronbach’s alpha or test-
retest).
o Off-diagonal entries = correlations between different measures.

What to look for:

1. Convergent evidence:
o Measures of the same trait across different methods should correlate
strongly.
o Example: Anxiety questionnaire score correlates highly with anxiety
interviewer rating.
2. Discriminant evidence:
o Measures of different traits should have low correlations, even if measured
by the same method.
o Example: Anxiety score should correlate weakly with depression score.
3. Method effects:
o Sometimes measures using the same method correlate higher than expected,
even for different traits.
o This shows that the method itself can inflate correlations and slightly reduce
discriminant validity.

Quick visual analogy:

 Imagine traits = colors (red = anxiety, blue = depression)


 Methods = shapes (circle = questionnaire, triangle = interview)
 MTMM matrix checks:
o Red circles & red triangles → should correlate highly (convergent)
o Red circles & blue triangles → should correlate low (discriminant)
o Red circle & blue circle → might be slightly higher due to method effect

You might also like