0% found this document useful (0 votes)
13 views20 pages

Module 2 - PSY 102

This midterm module on Psychological Statistics focuses on reliability, validity, and hypothesis testing for first-year psychology students. It covers essential concepts such as measurement quality, types of reliability and validity, and the importance of hypothesis testing in research. The module includes lectures, laboratory activities using SPSS, and learning outcomes aimed at enhancing students' understanding of statistical decision-making in psychology.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views20 pages

Module 2 - PSY 102

This midterm module on Psychological Statistics focuses on reliability, validity, and hypothesis testing for first-year psychology students. It covers essential concepts such as measurement quality, types of reliability and validity, and the importance of hypothesis testing in research. The module includes lectures, laboratory activities using SPSS, and learning outcomes aimed at enhancing students' understanding of statistical decision-making in psychology.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PSYCHOLOGICAL STATISTICS MODULE

(Midterm)

Module 2: Reliability, Validity, and Hypothesis Testing


Course/Subject: Psychological Statistics (Lecture & Laboratory)
Target Learners: 1st Year College (Psychology Program)
Estimated Time: Lecture: 3 hours | Laboratory: 6 hours (1 Week)

MODULE OVERVIEW
In psychology, we often measure things we cannot directly see—like stress, anxiety, depression,
motivation, self-esteem, and intelligence. So we use questionnaires, scales, tests, and
observations.

But here’s the question: Should we trust the results?


Because one “score” can affect real decisions.
●​ Who needs counseling support?
●​ Who qualifies for a scholarship?
●​ Which intervention works?
●​ Is there really a difference between two groups, or did it just happen by chance?

This module helps you understand two big foundations of research:


1.​ Measurement Quality
If our tool is poor, our conclusions will be poor. That’s why we study reliability and validity
first.
2.​ Statistical Decision-Making
Once we have good measurement, we test claims through hypothesis testing using
p-values, significance levels, confidence intervals, and standard error.

This module is designed for off-site learning, so expect guide questions, examples, and
activities that help you practice thinking like a researcher in the comfort of your own spaces.

LEARNING OUTCOMES
At the end of this module, students should be able to:

1.​ Explain reliability and validity using examples.


2.​ Differentiate test-retest, inter-rater, parallel-forms, and internal consistency reliability.
3.​ Explain the common measures of validity and why they matter in research.
4.​ Define hypothesis, null hypothesis, and alternative hypothesis.
5.​ Differentiate Type I and Type II errors using realistic scenarios.
6.​ Explain p-value and level of significance in simple language.
7.​ Explain confidence interval and standard error as tools for making decisions.
8.​ Perform reliability and standard error procedures using SPSS and interpret output
correctly.

MATERIALS NEEDED
● Calculator
● Laptop with IBM SPSS Statistics (for lab)
● Pen and paper/yellow pad
● Prepared dataset (questionnaire items + group variables)

WARM-UP (5–10 minutes)


Quick Check (Pre-Assessment)

Answer briefly. Don’t worry if you’re not sure yet—this is only to check what you already know.

1.​ If a test gives you almost the same score every time you take it, what does that suggest?
2.​ If a test gives consistent scores but measures the wrong thing, is it still “good”? Why?
3.​ What do you think the word “null” means in null hypothesis?
4.​ If SPSS shows p = .03, what do you think that number is telling you?

A. LECTURE MODULE
LESSON 1: RELIABILITY

1. What is Reliability?
Reliability simply means consistency.

When you hear “reliable,” think: “Masaligan.”


If a tool is reliable, it produces results that are stable, repeatable, and consistent.
Here’s a simple example:
Imagine you have a bathroom weighing scale. You step on it three times.
●​ If it shows: 52kg, 52kg, 52kg → consistent → reliable
●​ If it shows: 52kg, 60kg, 48kg → inconsistent → unreliable

In psychology, reliability matters because our tools measure traits and behaviors. If your stress
scale is unreliable, then a person’s score may change wildly—even if their stress level did not
actually change. That’s unfair and unscientific.

Reliability answers:
“Does this tool give consistent results?”

One important reminder:


A test can be reliable but still wrong. Consistent does not automatically mean correct. That idea
becomes clearer when we discuss validity.

2. Measures (Types) of Reliability


A. Test-Retest Reliability (Consistency Over Time)

This is reliability across time.

If you give the same test today, then give it again later (like after one week or two weeks), the
scores should still be similar—assuming the trait did not truly change.

Example:
A self-esteem scale is given on Monday and again next Monday.
If students’ scores are very similar, the test has good test-retest reliability.

Why it matters:
Psychological traits like intelligence or personality usually don’t change overnight. So if your test
changes drastically in a short period, the tool may be the problem—not the person.

B. Inter-Rater Reliability (Consistency Across Observers)


This is reliability between different observers or raters.

Example:
Two psychologists observe a child’s classroom behavior.
If one rates the child as “highly disruptive” and the other rates “slightly disruptive,” we have a
problem.
Why it matters:
Observation involves human judgment. Inter-rater reliability helps ensure fairness and
consistency—especially in diagnosis, performance tasks, and behavioral assessments.

C. Parallel-Forms Reliability (Consistency Between Test Versions)


Sometimes we create two versions of a test: Form A and Form B.

Both forms should measure the same construct and have similar difficulty.

Example:
Two versions of a stress test are given.
If students score similarly across both versions, we can say the test has good parallel-forms
reliability.

Why it matters:
This is common when teachers create multiple test forms to prevent cheating. It ensures that
Form B is not easier or harder than Form A.

D. Internal Consistency (Consistency Among Test Items)


This asks:
Do the items in the test “belong together”?

If the test claims to measure anxiety, then the questions should all reflect anxiety-related
thoughts, feelings, and behaviors—not random unrelated topics.

Most common statistic: Cronbach’s Alpha


Interpretation guide (rule of thumb):
●​ ≥ .90 = Excellent
●​ .80–.89 = Good
●​ .70–.79 = Acceptable
●​ Below .70 = Needs improvement

Why it matters:
A test can look professional but still be messy inside. Cronbach’s Alpha helps us check whether
the items are measuring one concept consistently.

Quick Self-Check (Not graded)


1.​ Which reliability checks stability over time?
2.​ Which reliability checks agreement between observers?
3.​ Which reliability is most connected to Cronbach’s Alpha?
Activity 1 (Lecture – Reliability)
Given Scenario A (Test-Retest)
1.​ A stress scale was given to the same students twice within 2 weeks. The correlation
between scores is r = .86.
2.​ What type of reliability is this?
3.​ Is r = .86 strong or weak reliability? Explain in 2–3 sentences.
4.​ Give one reason why test-retest reliability might become low even if the test is good.

Given Scenario B (Inter-Rater)


Two guidance staff rated the same student’s behavior during class observation.
●​ Rater 1 gave 9/10 (highly disruptive).
●​ Rater 2 gave 3/10 (slightly disruptive).
1.​ What type of reliability is affected?
2.​ Give two possible reasons why their ratings are far apart.

Given Scenario C (Internal Consistency)


A 10-item “Depression Scale” includes items like:

●​ “I feel hopeless.”
●​ “I feel tired most days.”
●​ “I enjoy parties and social gatherings.”

1.​ What problem might this scale have?


2.​ What reliability measure would help check this?

Output Requirements (Submit):


●​ Answers to all questions in complete sentences
●​ 1 short conclusion (2–3 sentences): Why reliability matters in psychological testing

LESSON 2: VALIDITY
Reliability is about consistency. Validity is about accuracy—meaning the test measures what it is
supposed to measure.

Think of a dartboard:

●​ If your darts hit the same spot repeatedly, you are consistent (reliable).
●​ If your darts hit the bullseye, you are accurate (valid).
A test can be consistent yet still miss the target.
That is why validity is essential.
Validity answers:
“Are we measuring the right thing?”

Measures (Types) of Validity


A. Face Validity

This is the simplest.


Does the test look like it measures what it claims to measure?

Example:
A math test that contains math problems has face validity.

But remember: face validity is not proof. Something can look correct but still be poorly designed.

B. Content Validity

This asks:
Does the test cover the full content of the construct?

Example:
A stress test should not measure only “schoolwork stress.”
It should also include items about physical symptoms, emotional stress, and cognitive stress.

Content validity is usually checked by experts. In real research, panels or professionals review
items to ensure completeness.

C. Construct Validity

This is deeper and more scientific.

It checks whether the test truly measures the theoretical concept.


It often uses patterns like:

●​ Convergent validity: It should relate to similar constructs


●​ Discriminant validity: It should not strongly relate to unrelated constructs

Example:
A depression scale should correlate with anxiety (related concept).
But it should not strongly correlate with height (unrelated).
D. Criterion Validity
This checks if the test relates to a real-world outcome.

Two types:
1.​ Predictive validity
Example: An entrance exam predicting future GPA.
2.​ Concurrent validity
Example: A new anxiety test matching results of an existing clinical diagnosis at the
same time.

Quick Reminder
Reliability is needed before validity.
Because if your tool is inconsistent, it cannot measure accurately.

Activity 2 (Lecture – Validity)


Scenario 1
A teacher made a “Study Habits Test” but all items only ask about time management (e.g., “I
make a schedule”). It does not ask about note-taking, focus, motivation, or study environment.
1.​ What type of validity is weak? Explain.
2.​ If you were to improve the test, what kinds of items should be added? Give 3 examples.

Scenario 2
A scholarship screening test is strongly related to students’ academic performance the following
year.
1.​ What type of validity is this?
2.​ Why is this important for fairness in scholarships?

Reflection
Why should we check reliability first before we claim a test is valid?

Output Requirements (Submit):


●​ Complete answers in sentences
●​ Provide examples where requested
●​ Short conclusion: “What I learned about validity” (3–4 sentences)

B. LABORATORY MODULE (SPSS)


Lab Activity 1: Calculating Reliability and Validity in SPSS
Laboratory Objectives

Students will be able to:

●​ Compute Cronbach’s Alpha in SPSS and interpret it.


●​ Use correlations as basic evidence for construct validity (convergent validity).
●​ Write a short interpretation paragraph like in research papers.

Part A: Internal Consistency (Cronbach’s Alpha)

Given: Dataset with 10 items measuring Academic Stress (AS1 to AS10)

Step-by-Step Procedure

Step 1: Open IBM SPSS Statistics


Step 2: Encode or open the dataset
Step 3: Click Analyze → Scale → Reliability Analysis
Step 4: Move AS1–AS10 into the Items box
Step 5: Model: Alpha
Step 6: Click OK

Tasks
●​ Write the Cronbach’s Alpha value.
●​ Interpret it using the guide (Excellent/Good/Acceptable/Needs improvement).
●​ If alpha is below .70, give TWO possible reasons and TWO suggestions to improve the
scale.

Part B: Evidence for Construct Validity (Convergent Validity using


Correlation)

Given:
Total Stress Score (Stress_Total)
Anxiety Score (Anxiety_Total)

Step-by-Step Procedure
Step 1: Click Analyze → Correlate → Bivariate
Step 2: Move Stress_Total and Anxiety_Total into Variables
Step 3: Choose Pearson (default)
Step 4: Click OK

Tasks
●​ Report the correlation coefficient (r) and p-value.
●​ Interpret in words: Is it significant? Is it positive/negative? Weak/moderate/strong?
●​ Explain: Does this support convergent validity? Why?

Output Requirements (Submit)


●​ Submit ONE document containing:
●​ Screenshot of SPSS output (Alpha and Correlation table)
●​ Answers to all tasks
●​ One short interpretation paragraph (5–7 sentences) as if you are writing a results section

PART C: HYPOTHESIS TESTING (LECTURE)

LESSON 3: HYPOTHESIS TESTING (With p-value, Null,


Alternative, Errors)

In research, we do not rely on feelings, assumptions, or guesses.


We rely on evidence.

But here’s the challenge:

When we collect data from a sample, how do we know if what we observed is real… or just
random chance?

This is where hypothesis testing begins.

Before we go into formulas or SPSS, we need to clearly understand the logic behind it. So let’s
move step by step.

1. What is a Hypothesis?
Let us start with the most basic question:

What is a hypothesis?
A hypothesis is a testable statement about a population.
It is a prediction that can be supported or rejected using data.

In psychology, examples include:


●​ Students who sleep less will have higher stress.
●​ There is a relationship between self-esteem and academic performance.
●​ There is a difference in anxiety levels between males and females.
Notice something important:

A hypothesis is not just an opinion. It must be measurable and testable.

Now, in hypothesis testing, we actually work with two hypotheses — not just one.

And this is where things become interesting.

2. The Null Hypothesis (H₀)


In statistics, we do not start by proving our idea.
Instead, we start by assuming that nothing is happening.
This assumption is called the null hypothesis.

The null hypothesis means no effect.


It states:
●​ There is no difference.
●​ There is no relationship.
●​ There is no effect.

Symbol: H₀

Examples:
1.​ If we are studying stress levels between working and non-working students:
H₀: There is no difference in stress levels.

2.​ If we are studying correlation:


H₀: There is no relationship between sleep and stress.

Why do we start with “no effect”?

Because statistics works by testing whether the evidence is strong enough to reject this
assumption.

It is easier to test “no effect” using probability than to directly prove something exists.

Very important:
We never say we “accept” the null hypothesis.

We only:
Reject H₀
or
Fail to reject H₀
Failing to reject does NOT mean it is true. It only means there is not enough evidence.

3. The Alternative Hypothesis (H₁ / Hₐ)


This is the hypothesis we actually care about.

The alternative hypothesis states:


●​ There IS a difference.
●​ There IS a relationship.
●​ There IS an effect.

Symbol: H₁ or Hₐ

Using the same example:


H₁: There is a difference in stress levels between working and non-working students.

There are two types of alternative hypotheses:


a.​ Non-directional (Two-tailed)
i.​ There is a difference (but no direction specified).
b.​ Directional (One-tailed)
i.​ Working students have higher stress than non-working students.

The difference matters because it affects how we test the hypothesis.

Now that we understand H₀ and H₁, we move to the core of hypothesis testing:

How do we decide which one to keep?

THE LOGIC OF HYPOTHESIS TESTING

Here is the logic:

1.​ Assume the null hypothesis is true.


2.​ Collect sample data.
3.​ Calculate a test statistic.
4.​ Determine how likely our result is under H₀.

If the result is very unlikely under the assumption of “no effect,”


we reject H₀.

If it is not unlikely,
we fail to reject H₀.
Everything now depends on probability.

And this brings us to one of the most misunderstood concepts:

The p-value.

4. What is a p-value?
The p-value is one of the most important numbers in research.

The p-value stands for probability value.

It tells us:
If the null hypothesis is true, how likely is it to get results like ours (or more extreme)?

Suppose we run a study and SPSS gives


Example: p = .03

This means:
If there is truly no difference in the population, there is only a 3% chance that we would observe
a difference this large in our sample.

Since 3% is small, we consider that unlikely.

So we reject H₀.

Important clarification:
The p-value does NOT tell us:
●​ The probability that the null hypothesis is true.
●​ The probability that the study is correct.
●​ The probability that the results happened by chance.

Instead, it tells us:


Assuming no effect exists, how unusual is our data?

Small p-value → data is unusual under H₀


Large p-value → data is consistent with H₀
5. Level of Significance (α)
Now we introduce another key concept: alpha (α).

Alpha is the level of significance.

It is the threshold we set BEFORE analyzing data.

Common values:
α = .05
α = .01

If α = .05, we accept a 5% risk of rejecting a true null hypothesis.

Decision Rule:
If p < α → Reject H₀
If p > α → Fail to reject H₀

Example:
If α = .05 and p = .03

Since .03 < .05 → Reject H₀


If α = .05 and p = .20

Since .20 > .05 → Fail to reject H₀

Alpha represents how strict we are.

Smaller alpha → more strict → harder to reject H₀

6. EFFECT SIZE (WHY SIGNIFICANCE IS NOT ENOUGH)


Statistical significance does not tell us how big the effect is.

With large samples, even very small differences can become statistically significant.

Example:
Group A mean = 80
Group B mean = 81

p = .01
Statistically significant.
But is 1 point meaningful?

Maybe not.

Effect size measures magnitude.

Examples:
Cohen’s d
Pearson’s r

Interpretation for Cohen’s d:


0.20 = small
0.50 = medium
0.80 = large

Always ask:
Is the result statistically significant?
Is the result practically meaningful?

7. Types of Errors
Whenever we make decisions, mistakes are possible.

In hypothesis testing, there are two types.

1.​ Type I Error (False Positive):


Rejecting a true H₀
You conclude there is an effect, but there is none.

Symbol: α

Example: Concluding that a therapy works when it actually does not.

2.​ Type II Error (False Negative):


Failing to reject a false H₀
You conclude there is no effect, but there is one.

Example: Concluding that therapy does not work when it actually does.

Notice:

If we reduce alpha to avoid Type I error, we increase the risk of Type II error.
There is always a trade-off.
Activity 3 (Lecture – Hypotheses, p-values, and Errors)
1.​ Write H₀ and H₁ for each:

a. “There is a difference in anxiety levels between males and females.”


b. “Social media use is related to self-esteem.”

2.​ A researcher claims a program works. The truth is: it does not work. The researcher
concluded it works.
What error occurred? Explain in 2–3 sentences.
3.​ A researcher concluded there is no difference between two teaching methods. The truth
is: there is a difference.
What error occurred? Explain in 2–3 sentences.
4.​ Interpret this in your own words: p = .04, α = .05

LAB ACTIVITY 2: Differentiating Type I and Type II Errors


Before answering, remember:

Type I Error (False Positive) → You conclude there IS an effect when in reality there is NONE.
Type II Error (False Negative) → You conclude there is NO effect when in reality there IS one.

These errors are not just theoretical — they happen in real research, medicine, education, and
psychology practice. That is why understanding them is important.

Read each scenario carefully. Think about what the researcher concluded and what the truth
actually is.

Scenario 1: Anxiety Intervention Program


A school implemented a new anxiety-reduction workshop for senior high school students. After
conducting a study, the researcher concluded that the workshop significantly reduced anxiety
levels.
However, in reality, the workshop had no actual effect. Students’ anxiety levels did not truly
change; the difference was only due to random variation.

Questions:
1.​ What type of error occurred?
2.​ Explain your reasoning in 2–3 sentences.

Scenario 2: Sleep and Academic Performance


A researcher wanted to examine whether sleep duration affects GPA. After analyzing the data,
the researcher concluded that there was no significant relationship between sleep and GPA.
However, in reality, sleep does significantly affect academic performance — the study just failed
to detect it.

Questions:
1.​ What type of error occurred?
2.​ Why might this error happen in research?

Scenario 3: New Therapy for Depression


A psychologist tested a new therapy for mild depression. The results showed a statistically
significant improvement compared to a control group.
Later, a larger and better-designed study showed that the therapy was no better than the
standard treatment.

Questions:
1.​ What type of error likely occurred in the first study?
2.​ What might have contributed to this error? (Think about sample size, design, chance.)
3.​ identify error type and explain briefly.

Output:
Compile answers in one document.

LESSON 4: STANDARD ERROR, CONFIDENCE


INTERVAL, AND WHY THEY MATTER
When we compute a mean from a sample, that mean is an estimate. If we took another sample,
we might get a slightly different mean. This natural variation is normal. Statistics helps us
describe how stable our estimate is.

Standard Error (SE)

Standard error tells us how much sample means tend to vary.


Formula:

SE = SD / √n

Meaning:
Larger sample size → smaller SE
Smaller SE → more precise estimate

Confidence Interval (CI)


A confidence interval gives a range of plausible values for the population parameter.
A 95% CI means the method is expected to capture the true value in about 95% of repeated
sampling.
Activity 4 (Lecture – SE, CI, α)
Answer in complete sentences:

1.​ What happens to SE if sample size increases? Why?


2.​ Explain 95% CI in your own words.
3.​ If α = .01 and p = .03, what is the decision? Explain.
4.​ Why is it helpful to report CI instead of only reporting p-values?

LAB ACTIVITY 3: Calculating and Interpreting Standard


Error in SPSS
Scenario
You are conducting a small study among 20 first-year Psychology students to examine their
Academic Stress Level during midterm week.

Each student answered a 30-item stress questionnaire. The total possible score ranges from 0
to 120, where:

Higher score = Higher stress


Lower score = Lower stress

You encoded the following total stress scores into SPSS:

72 65 80 75 69 85 77 74 68 82 79 71 76 88 70 83 67 73 78 81

Sample Size (n) = 20 students

Step-by-step Procedure in SPSS


Analyze → Descriptive Statistics → Explore
Move variable into Dependent List
OK

What to Look for in the Output


In the Descriptives table, locate:
●​ Mean
●​ Std. Deviation (SD)
●​ Std. Error (SE)
Tasks:
Write Mean, SD, SE
Interpret SE in 4–6 sentences (what it suggests about precision)
Guide for Understanding Standard Error
Remember:
SE = SD / √n
This means:
Larger n → Smaller SE
Smaller SE → More stable estimate of population mean
Standard Error does NOT describe individual variability.
It describes how much the sample mean would vary across repeated samples.
That is why SE is used when constructing Confidence Intervals.

LESSON 5: STEPS IN HYPOTHESIS TESTING (Complete


Process)
A correct conclusion is not just “p < .05.” Good research follows a process.

1.​ State H₀ and H₁ clearly


2.​ Choose α (risk level)
3.​ Select correct statistical test
4.​ Compute test statistic and p-value
5.​ Compare p with α
6.​ Make decision: reject or fail to reject H₀
7.​ Interpret results in words (context matters)
Example conclusion:

Since p = .02 is less than α = .05, we reject the null hypothesis. There is sufficient
evidence of a difference in stress levels.

Notice the wording: We do not say “proven.” We say “sufficient evidence.”


8.​ Report effect size when applicable (practical meaning)

LAB ACTIVITY 4: Problem Set for Hypothesis Testing


Now you will practice decision-making using p-values and alpha (α).

Remember:
Decision Rule:
If p < α → Reject H₀ (Statistically Significant)
If p > α → Fail to Reject H₀ (Not Statistically Significant)

Also remember:
Rejecting H₀ increases risk of Type I Error
Failing to reject H₀ increases risk of Type II Error
For each:

1.​ Reject or Fail to Reject H₀


2.​ Significant or not
3.​ Which error is more possible based on the decision? (Type I or II) and why

Case 1
p = .02
α = .05
Should we Reject or Fail to Reject H₀?
Is the result statistically significant?
Which error becomes more possible after this decision (Type I or Type II)? Why?

Case 2
p = .08
α = .05
Decision?
Significant or not?
Which error becomes more possible? Why?

Case 3
p = .001
α = .01
Decision?
Significant or not?
Which error becomes more possible? Why?

SUMMARY (Key Takeaways)


● Reliability is consistency: stable results matter.
● Validity is accuracy: the tool must measure the right concept.
● Hypothesis testing is decision-making under uncertainty.
● Null hypothesis assumes no effect; alternative suggests an effect.
● p-value tells how unusual results are if H₀ were true.
● Alpha is the risk threshold for deciding.
● Standard error and confidence intervals explain precision of estimates.
● A good conclusion includes interpretation, not just numbers.

FINAL CONNECTION
Let us connect everything:

●​ Hypothesis → what we want to test


●​ Null hypothesis → assumes no effect
●​ Alternative hypothesis → assumes there is an effect
●​ P-value → how unusual our data is if H₀ is true
●​ Alpha → our decision threshold
●​ Effect size → how big the effect is
●​ Decision → reject or fail to reject H₀
●​
When you understand this flow, hypothesis testing becomes logical — not scary.

REFLECTION QUESTIONS FOR SELF-STUDY


1.​ Why do we start by assuming the null hypothesis is true?
2.​ What does a p-value of .04 actually mean?
3.​ If alpha is .01 and p is .03, what is the decision?
4.​ Why is effect size important in psychology research?
5.​ What is the difference between statistical significance and practical significance?

POST-ASSESSMENT
Explain the difference between reliability and validity using your own example.

1.​ A test has alpha = .60. What does it suggest and what can be improved?
2.​ Interpret: p = .04, α = .05. What is the decision and what does it mean in words?
3.​ Why does a larger sample size usually produce a smaller standard error?

You might also like