STATISTICAL
ANALYSIS
MEASURES, DISTRIBUTIO NS, AND RELATIO NSHI PS
Measures of Central Tendency
•Definition: A measure of central tendency is a descriptive statistic that describes the
average, or typical value of a set of scores.
•There are three common measures of central tendency:
•The Mode
•The Median
•The Mean
The Mode: The Most Frequent Score
• Definition: The mode is
the score that occurs
most frequently in a set
of data.
Advanced Modal Concepts: Bimodal and
Multimodal Distributions
Bimodal Distribution:
When a distribution has
two "modes," or two
scores with the same
highest frequency, it is
called bimodal.
Multimodal Distribution: If a distribution has
more than two "modes," it is called
multimodal.
Multimodal Distribution: If a distribution has more than two
"modes," it is called mu
Use Cases and Limitations of the Mode
•Primary Use Case: The mode is primarily used with nominally scaled data.
It is the only measure of central tendency appropriate for data that is
categorical without a natural order (e.g., favorite color, type of car).
•Key Limitations:
•The mode is often not a very useful measure of central tendency for
numerical data because of its instability.
•It is insensitive to large changes in the data set. Two datasets that are
structurally very different from each other can have the same mode.
The Median: The Middle Score
•Definition: The median is another name for the 50th percentile.
•Conceptual Understanding: It is the score that sits in the exact
middle of a dataset after it has been arranged in order. Half of the
scores are larger than the median, and half of the scores are
smaller than the median.
How to Calculate the Median
•The process for finding the median involves two main steps:
[Link] the data from highest to lowest (or lowest to highest).
[Link] the score in the middle position.
•Formula for Middle Position: The position of the median can be
found using the formula:
middle=(N+1)/2
where N is the number of scores.
•Special Case (Even Number of Scores): If N is an even number, the formula
will result in a.5 position (e.g., 3.5). In this case, the median is the average of
the two scores on either side of that position (e.g., the average of the 3rd and
4th scores).
•Example 1: Odd Number of Scores (N=11)
•Data: 10, 8, 14, 15, 7, 3, 3, 8, 12, 10, 9
•Sorted Data: 15, 14, 12, 10, 10, 9, 8, 8, 7, 3, 3
•Middle Position: (11+1)/2=6th position
•Median: The score in the 6th position is 9.
•Example 2: Even Number of Scores (N=6)
•Data: 24, 18, 19, 42, 16, 12
•Sorted Data: 42, 24, 19, 18, 16, 12
•Middle Position: (6+1)/2=3.5th position
•Median: The average of the 3rd and 4th scores:
(19+18)/2=18.5.
When to Use the Median
•The median is the preferred measure of central tendency
when the distribution of scores is skewed, either positively
or negatively.
•The reason for this preference is its robustness. A few
extremely large scores (in a positively skewed distribution)
or extremely small scores (in a negatively skewed
distribution) will not overly influence the median, unlike the
mean. This makes the median a more "typical"
representation of the center in such cases.
The Mean: The Arithmetic Average
•The mean can be understood in several ways:
[Link] is the arithmetic average of all the scores, calculated as
(ΣX)/N.
[Link] is the number, m, that makes the sum of deviations from it
equal to zero: Σ(X−m)=0. This identifies the mean as the
"balance point" of the distribution.
[Link] is the number, m, that makes the sum of squared deviations a
minimum: Σ(X−m)2. This property is the foundation for least
squares regression.
•Statistical Notation:
•The mean of a population is represented by the Greek letter μ (mu).
•The mean of a sample is represented by X (X-bar).
Calculating the Mean: Example
•Problem: Calculate the mean of the following data: 1, 5, 4, 3, 2.
•Step 1: Sum the scores (ΣX):
1+5+4+3+2=15
•Step 2: Divide the sum by the number of scores (N=5):
15/5=3
•Result: The mean, X, is 3.
When to Use the Mean
•The mean should be used under the following conditions:
•The data are measured on an interval or ratio scale.
•The data are not significantly skewed.
•Primary Advantage: The mean is the most common measure of central
tendency because it is sensitive to every score in the dataset. If any single
score is changed, the value of the mean will also change. This property makes
it powerful for statistical inference but also vulnerable to the influence of
outliers.
Relationship Between Measures of Central
Tendency
The shape of a distribution influences the
relationship between the mean, median, and
mode.
•Symmetrical Distributions: In a perfectly
symmetrical distribution (like the normal
distribution), the mean, median, and mode are all
equal and located at the center of the
distribution.
•Mean=Median=Mode
•Positively Skewed Distributions: The tail of
the distribution extends to the right. Extreme high
scores pull the mean to the right.
•Mode<Median<Mean
•Negatively Skewed Distributions: The tail of
the distribution extends to the left. Extreme low
scores pull the mean to the left.
•Mean<Median<Mode
Central Tendency for Grouped Data -
Introduction
•Frequency Table: A frequency table is a
method of organizing raw data in a compact
form. It displays a series of scores (x) along
with their corresponding frequencies (f), which
is the number of times each score occurs.1
•Example Table Structure:
X (Score) f (Frequency)
19 1
18 2
... ...
Frequency Distribution Table
Grouped Data :Mode
• Mode = Highest Frequency
• MODE = 13
GROUPED DATA: MEDIAN
Median class = (N/2)th value
= (50/2)th value
= 25th value
Median class = 30 to 40
l = 30, N//2 = 25, m = 24, f = 10 and c = 10
Substitute.
Median = 30 + ([25 - 24]/10) x 10
= 30 + 1
≈ 31
GROUPED DATA :MEAN
• Mean= 1132/16
• Mean=70.75
Central Tendency for Continuous Grouped
Data
• When data is grouped into class intervals (e.g., 90-100), we lose the exact values
of the original scores. The formulas used to find the central tendency are
therefore estimations based on the assumption that data points are
distributed evenly within their respective intervals.
MODE FORMULA
• The formula to find the mode of grouped data is:
𝑓1 −𝑓0
• Mode = 𝑙 + ×ℎ
2𝑓1 −𝑓0 −𝑓2
• where:
• 𝑙 = lower limit of the modal class (class interval with the highest
frequency)
• 𝑓1 = frequency of the modal class
• 𝑓0 = frequency of the class preceding the modal class
• 𝑓2 = frequency of the class succeeding the modal class
• ℎ = size (width) of the class interval
Size of Family 1-3 3-5 5-7 7-9 9-11
Number of
7 8 2 2 1
Families
•Modal class (highest frequency) = 3-5
•𝑙 = 3 (lower limit of modal class)
•ℎ = 2 (class interval size)
•𝑓1 = 8) frequency of modal class)
•𝑓0 = 7) frequency of preceding class, 1-3)
•𝑓2 = 2) frequency of succeeding class, 5-7)
Using the formula:
8−7 1 1 2
Mode = 3 + ×2=3+ ×2=3+ ×2=3+
2×8−7−2 16 − 7 − 2 7 7
= 3.286
Thus, the mode of the grouped data is approximately 3.286.
MEDIAN FORMULA
• Median=l + [(n/2−cf)/f] × h
• where:
• 𝑙 = lower limit of the median class (the class containing the median)
• 𝑛 = total number of observations
• 𝑐𝑓 = cumulative frequency of the class before the median class
• 𝑓 = frequency of the median class
• ℎ = width (size) of the class interval
MEDIAN
Total frequency, N = 45
Cumulative
Class Interval Frequency (f) N/2 = 22.5
Frequency (CF)
The cumulative frequency just greater than
0 - 10 5 5 22.5 is 28, which corresponds to the class 20
- 30. So, this is our median class.
10 - 20 8 13
S
•L = 20 Median = 20+(22.5- 13 /15)*10
20 - 30 15 28
•CF = 13 Median = 26.33
30 - 40 10 38 •f = 15
•h = 10
40 - 50 7 45
MEAN FORMULA
• Mean=∑ fi xi / ∑fi
• where:
𝑓𝑖 = frequency of the 𝑖 𝑡ℎ class interval
𝑥𝑖 = midpoint of the 𝑖 𝑡ℎ class interval (average of lower and upper
class limits)
MEAN
• For class intervals 10–20, 20–30, 30–40 with frequencies 3, 7, 5, the
midpoints are 15, 25, 35. The mean is calculated as:
3×15+7×25+5×35
•
3+7+5
• This results in a weighted average representing the data's central
value.
• Mean = 19.66
Introduction to Measures of Dispersion
•Definition: The measure of dispersion indicates the variability, spread, or
scattering of data. It describes the extent to which values in a distribution
differ from the average.
•Core Concept:
•The more similar the scores are to each other, the lower the measure of
dispersion will be.
•The less similar the scores are to each other (i.e., the more spread out
they are), the higher the measure of dispersion will be.
•
The Practical Importance of Dispersion
•Scenario: Choosing between two brands of paint. Both have the same average
lifespan before fading, but their consistency differs.
•Data:
Paint A (months) Paint B (months)
10 35
60 45
50 30
30 35
40 40
20 25
•Analysis:
•Average: Both paints have an average life of 35 months. The mean does not help in
the decision.
•Spread (Range):
•Paint A: 60 - 10 = 50months
•Paint B: 45 - 25 = 20 months
•Conclusion: Paint B has a smaller variation, which means it performs more consistently.
Dispersion is a measure of consistency and predictability.
The Ran The Range
•Definition: The range is the difference between the largest score (X_L) and the
smallest score (X_S) in a set of data.1
Range = X_L - X_S
•Example: For the data 4, 8, 1, 6, 6, 2, 9, 3, 6, 9, the largest score is 9 and the
smallest is 1. The range is:
9-1=8
•Limitations:
•The range is rarely used in scientific work as it is fairly insensitive.
•It depends on only two scores in the entire dataset. Two very different sets of
data can have the same range (e.g., 1, 1, 1, 1, 9 vs. 1, 3, 5, 7, 9).
The Semi-Interquartile Range (SIR)
•Definition: The semi-interquartile range (SIR) is defined as half the difference between
the third and first quartiles.1
SIR = (Q_3 - Q_1) / 2
•Quartiles:
•The first quartile (Q_1) is the 25 percentile (25% of observations are smaller).
•The second quartile (Q_2) is the median (50 percentile).
•The third quartile (Q_3) is the 75 percentile (75% of observations are
smaller).
•Use Case: The SIR is often used with skewed data because, by focusing on the middle
50% of the data, it is insensitive to the extreme scores in the tails.
Calculating Quartiles and SIR: Example
•Sample Data (Ordered): 11, 12, 13, 16, 16, 17, 18, 21, 22 (n=9).
•Calculate Q_1 Position:
•Position = (n+1)/4 = (9+1)/4 = 2.5th position.
•Q_1 is the average of the 2^{nd} and 3^{rd} values: (12+13)/2 = 12.5.
•Calculate Q_3 Position:
•Position = 3(n+1)/4 = 3(9+1)/4 = 7.5th position.
•Q_3 is the average of the 7^{th} and 8^{th} values: (18+21)/2 = 19.5.
•Calculate SIR:
SIR = (Q_3 - Q_1) / 2 = (19.5 - 12.5) / 2 = 7 / 2 = 3.5
Measure 3: Variance and Standard Deviation
(Conceptual)
•Deviation Score: The starting point is the deviation score, (xi - x̄), which measures how far a single score
is from the mean.
•The Problem with Averaging Deviations: The sum of all deviation scores, Σ(xi - x̄), is always zero because
the mean is the balance point. A simple average of deviations would therefore always be zero.
•The Solution: Squaring: To overcome this, each deviation score is squared. This makes all values positive
and also gives greater weight to scores that are further from the mean. A deviation of 4 (squared = 16)
has a much larger impact than a deviation of 2 (squared = 4). This property makes variance highly
sensitive to outliers, just like the mean.
•Variance: The mean of the squared deviation scores.
•Standard Deviation: The square root of the variance. Taking the square root returns the measure of
dispersion to the original units of the data (e.g., from pounds² back to pounds).
Calculating Variance and Standard Deviation
(Formulas)
•Population Variance (sigma^2):
•Definitional Formula:
σ² = Σ(X-μ)² / N
•Computational Formula:
σ² = (ΣX² - ((ΣX)²/N)) / N
1
•Sample Variance (s^2):
•Formula:
s² = Σ(X-X̄)² / (N-1)
•The use of $N-1$ in the denominator for a sample is a correction that provides a better,
unbiased estimate of the true population variance.1
•Standard Deviation:
•Population: σ = √σ²
•Sample: s = √s² 1
• Worked Example: Calculating Variance
• Data and Calculation Table:
X X−μ (X−μ)2 X2
9 2 4 81
8 1 1 64
6 -1 1 36
5 -2 4 25
8 1 1 64
6 -1 1 36
Sigma=42 Sigma=0 Sigma=12 sigma=306
Calculation (Definitional): N=6, μ = 42/6 = 7
σ² = Σ(X-μ)² / N = 12 / 6 = 2
Calculation (Computational):
σ² = (306 - (42)²/6) / 6 = (306 - 1764/6) / 6 = (306 - 294) / 6 = 12 / 6 = 2
Dispersion for Grouped Data: Range and
Quartile Deviation
•Range:
•For continuous grouped data, the range is the upper boundary of the highest class
minus the lower boundary of the lowest class.
•Example (BP Data): R = 180 - 90 = 90.1
•Quartile Deviation (SIR):
•The formulas for $Q_1$ and $Q_3$ are used for grouped data.
•Q₁ = Lₒ + ( (N/4 - m₁) / f₁ ) * C
•Q₃ = Lₒ + ( (3N/4 - m₃) / f₃ ) * C
Example (BP Data):
•N/4=17. Q_1 class is 120-130. Q₁ = 120 + ((17-15)/10) * 10 = 122.
•3N/4=51. Q_3 class is 140-150. Q₃ = 140 + ((51-40)/11) * 10 = 150.
•Semi Quartile Range = (150 - 122) / 2 = 14.
Dispersion for Grouped Data: Standard
Deviation
•Formula:
s = √ ( Σf(x-X̄)² / N )
where $x$ is the midpoint of the class interval and $\overline{X}$ is the mean of the distribution.1
•Calculation Table (BP Data):
•Instructions for visual: Recreate the large calculation table from page 63. This table must include columns
for Class Intervals, f, Midpoint (x), f*x, (x-\overline{X}), (x-\overline{X})^2, and f(x-\overline{X})^2.
•Calculation Steps (BP Data):
[Link] the mean: $\overline{X} = 135.6.
[Link] each class, calculate the deviation from the mean $(x-\overline{X})$, square it, and multiply by the
frequency f.
[Link] the final column: Sigma f(x-\overline{X})^2 = 26376.5.
[Link] into the formula:
s = √(26376.5 / 68) = √387.89 ≈ 19.7
Interpreting Standard Deviation: The
Empirical Rule
•If a dataset has a distribution that is approximately "mound-shaped" or "bell-
shaped" (i.e., normal), then we can use the standard deviation to estimate the
proportion of data within certain ranges of the mean.
•The Rule:
•Approximately 68% of the data lies within one standard deviation of the mean (μ
± 1σ).
•Approximately 95% of the data lies within two standard deviations of the mean
(μ ± 2σ).
•Approximately 99.7% of the data lies within three standard deviations of the
mean (μ ± 3σ).
Coefficient of Variation (CV): A Relative
Measure
•Purpose: The Coefficient of Variation (CV) is a measure of relative variation. It is used to compare the
variability of two or more groups, especially when their means are different. It expresses the standard
deviation as a percentage of the mean.
•Formula:
CV = (SD / X̄) * 100%
•Example: Comparing two stocks.
•Stock A: Average Price = 50, Standard Deviation = 5
•CV_A = (5/50) * 100% = 10%
•Stock B: Average Price = 100, Standard Deviation = 5
•CV_B = (5/100) * 100% = 5%
•Conclusion: Although both stocks have the same standard deviation (absolute variation), Stock A is
relatively more volatile because its variation is larger compared to its average price.
Measures of Distribution Shape: Skewness
•Definition: Skew is a measure of the asymmetry in the distribution of scores. 1
•Types of Skew:
•Normal (Symmetrical) Distribution: Skew = 0.
•Positive Skew: The tail of the distribution extends to the right. The scores are clustered at the
lower end.
•Negative Skew: The tail of the distribution extends to the left. The scores are clustered at the
higher end.
•Interpretation of Skew Value ($s^3$):
•If s³ < 0, the distribution has a negative skew.
•If s³ > 0, the distribution has a positive skew.
•If s³ = 0, the distribution is symmetrical.1
Measures of Distribution Shape: Kurtosis
•Definition: Kurtosis measures the “ peakedness " of a distribution, or
whether the scores are spread out more or less than they would be in a
normal (Gaussian) distribution.
•Types of Kurtosis:
•Mesokurtic (s⁴ = 3): The distribution has a kurtosis equal to that of a
normal distribution.
•Leptokurtic (s⁴ > 3): The distribution is more peaked and has "heavier"
tails than a normal distribution. This indicates more outliers.
•Platykurtic (s⁴ < 3): The distribution is flatter and has "lighter" tails than a
normal distribution. This indicates fewer outliers.
Summarizing Distribution Shape
•Collectively, the variance (s²), skew (s³), and kurtosis (s⁴) are moments
of the distribution that describe its shape.
•Variance (s²): Measures the spread or width.
•Skew (s³): Measures the asymmetry or lopsidedness.
•Kurtosis (s⁴): Measures the peakedness and tail weight.
Data Presentation: Frequency Tables
•Objective: To understand different ways to summarize and present data effectively.
•Frequency Distribution Table: A simple way to summarize data by listing categories and
their numerical counts (frequencies).
•Example: Constructing a Frequency Distribution
Categories Tally Frequencies Percentage
Anger (A) IIIII 5 14.29%
Excitement (E) IIIII 5 14.29%
Fear (F) IIIII I 6 17.14%
Happiness (H) IIIII IIIII II 12 34.29%
Interest (I) III 3 8.57%
Sadness (S) IIII 4 11.42%
Total 35 100.00%
•Grouped Frequency Distribution: Categories can be grouped for simplification.
Categories Frequency Percentage
Positive Emotions 20 57.15%
Negative Emotions 15 42.85%
Data Presentation: Charts and Graphs
•Purpose: Charts and graphs provide a visual representation of data, making it easier to see trends,
relationships, and comparisons.1
•Basic Guidance for Graphics:
•Ensure the graphic has a clear title.
•Label all components (axes, series, etc.).
•Indicate the source of the data and the date.
•Provide the number of observations ($n$) as a reference point.1
•Common Types of Graphics:
•Bar Chart: For comparisons across categories.
•Line Graph: To display trends over time.
•Pie Chart: To show percentages or proportional shares of a whole.1
BAR CHART
Chart Title
Category 4
Category 3
Category 2
Category 1
0 1 2 3 4 5 6
Series 3 Series 2 Series 1
LINE CHART
Chart Title
6
0
Category 1 Category 2 Category 3 Category 4
Series 1 Series 2 Series 3
PIE CHART
Sales
1st Qtr 2nd Qtr 3rd Qtr 4th Qtr
Common Statistical Graphs for Frequency
Distributions
•Histogram: A vertical bar chart of frequencies for continuous data.
The bars touch to indicate that the variable is continuous.
•Frequency Polygon: A line graph of frequencies. A point is plotted for
the frequency of each class at its midpoint, and the points are
connected by lines.
•Ogive: A line graph of cumulative frequencies. It is useful for finding
percentiles.
Frequency Polygon
Ogive Curve
Introduction to Hypothesis Testing
•Purpose: To use statistical methods to determine whether sample data
supports or rejects a specific idea or hypothesis about a population.
•Core Problem: We often collect data that seems to support our ideas.
Statistical testing provides an objective framework to prevent us from
uncritically accepting a pet idea. It keeps the scientific process honest.
•Key Principle: If a research paper suggests an alternative hypothesis should be
accepted but provides no statistical test, the claim should be viewed with
skepticism.
The Hypothesis Testing Framework
•Hypothesis testing involves choosing between two competing statements about a population based on
sample data.1
•The Two Hypotheses:
•Null Hypothesis (H_0): This is the statement we assume to be true to begin with. It typically represents
the status quo or a statement of "no effect," "no difference," or "no relationship". 1
•Alternative (Research) Hypothesis (H_A): This is what we aim to gather evidence for. It typically
represents the presence of an effect, a difference, or a relationship.
•The Courtroom Analogy:
•H_0: The person is innocent.
•H_A: The person is not innocent (i.e., guilty).
•The jury can only reject the null hypothesis of innocence if there is evidence "beyond a reasonable
doubt." We can only reject our statistical H_0 if our sample data is sufficiently unlikely to have occurred
if H_0 were true.
Type I and Type II Errors
Because we make decisions based on sample data, not the entire population, our conclusions can be wrong. There are
two types of errors we can make.1
•Type I Error ($\alpha$):
•Definition: Rejecting the null hypothesis (H_0) when it is actually true.
•Analogy: Concluding the drug works when it doesn't. Convicting an innocent person.
•The probability of a Type I error is called the level of significance ($\alpha$), which is typically set at 5% (0.05).1
•Type II Error ($\beta$):
•Definition: Failing to reject the null hypothesis (H_0) when it is actually false.
•Analogy: Concluding there is insufficient evidence that the drug works when it actually does. Acquitting a guilty
person.
•The probability of correctly rejecting a false null hypothesis is called the Power of the test ($1-\beta$).1
•Decision Table:
True State of Null Hypothesis: H0 True State of Null Hypothesis: H0
Your Statistical Decision
is True is False
Reject $H_0$ Type I Error ($\alpha$) Correct Decision (Power)
Do not reject $H_0$ Correct Decision Type II Error ($\beta$)
Introduction to Correlation and Regression
•Purpose: This area of inferential statistics involves determining whether a
relationship exists between two or more numerical variables.
•Key Questions Answered:
[Link] two or more variables related?
[Link] so, what is the strength of the relationship?
[Link] type of relationship exists (e.g., positive, negative, linear)?
[Link] kind of predictions can be made from the relationship?
•The Two Main Tools:
•Correlation: A statistical method used to determine whether a relationship
between variables exists and to measure its strength and direction.
•Regression: A statistical method used to describe the nature of the relationship
and to create a model for prediction.
Scatter Plots
•Definition: A scatter plot is a graph of ordered pairs (x, y) that visually
represents the relationship between two numerical variables. The
independent variable (x) is plotted on the horizontal axis, and the
dependent variable (y) is plotted on the vertical axis.
•Purpose: The scatter plot is the first step in correlation and regression
analysis. It allows us to visually inspect the data to see if a relationship
exists and to identify its nature (e.g., positive linear, negative linear,
curvilinear, or no relationship).1
The Correlation Coefficient (r)
•Definition: The Pearson product-moment correlation coefficient (PPMC), denoted
by r for a sample and rho for a population, is a numerical measure of the strength
and direction of a linear relationship between two variables.
•Range: The value of r ranges from -1 to +1.
•r close to +1: Indicates a strong positive linear relationship.
•r close to -1: Indicates a strong negative linear relationship.
•r close to 0: Indicates no linear relationship or a very weak one.
•Strength of Association:
•Strong: |r|>= 0.75
•Intermediate: 0.25 =<|r| =< 0.75
•Weak: |r| < 0.25
Calculating the Correlation Coefficient (r)
Student Absences (x) Final Grade (y) xy x² y²
A 6 82 492 36 6724
B 2 86 172 4 7396
C 15 43 645 225 1849
D 9 74 666 81 5476
E 12 58 696 144 3364
F 5 90 450 25 8100
G 8 78 624 64 6084
Totals (Σ) Σx=57 Σy=511 Σxy=3745 Σx²=579 Σy²=38993
Calculating the Correlation Coefficient (r)
•Formula:
r = (n(Σxy) - (Σx)(Σy)) / √([n(Σx²) - (Σx)²][n(Σy²) - (Σy)²])
where n is the number of data pairs.
•Example: Absences and Final Grades
•Values: n=7, Sigma x=57, Sigma y=511, Sigma xy=3745, Sigma x^2=579 , Sigma
y^2=38993.
•Calculation:
r = (7(3745) - (57)(511)) / √([7(579) - (57)²][7(38993) - (511)²]) = (26215 - 29127) /
√([4053-3249][272951-261121]) = -2912 / √((804)(11830)) = -0.944
•Conclusion: The value of r = -0.944 suggests a very strong negative linear
relationship between absences and final grades.
Introduction to Regression
•Purpose: If the correlation coefficient is found to be significant, the next step is to
determine the equation of the regression line, which is the line of best fit for the data.
•Line of Best Fit: This is the line that minimizes the sum of the squares of the vertical
distances from each data point to the line. It represents the trend in the data and allows
for prediction.
•Regression Equation:
•The equation of the regression line is written as:
ŷ = a + bx
•(y-hat) is the predicted value of y for a given value of x.
•a is the y-intercept (the predicted value of y when x=0).
•b is the slope of the line (the change in y for a one-unit change in x).
Calculating the Regression Line Equation
•Formulas for Slope (b) and Intercept (a):
b = (n(Σxy) - (Σx)(Σy)) / (n(Σx²) - (Σx)²)
a = (Σy)/n - b((Σx)/n) = ȳ - bx̄
•Example: Absences and Final Grades
•Using the sums from the correlation calculation:
•Slope (b):
b = (7(3745) - (57)(511)) / (7(579) - (57)²) = -2912 / 804 = -3.622
•Intercept (a):
a = (511/7) - (-3.622)*(57/7) = 73 - (-29.49) = 102.493
•Regression Equation:
ŷ = 102.493 - 3.622x
Using the Regression Line for Prediction
•The primary use of the regression equation is to make
predictions.
•Example: Predicting Final Grade
•Using the equation ŷ = 102.493 - 3.622x:
•What is the predicted final grade for a student with 5 absences
(x=5)?
ŷ = 102.493 - 3.622(5) = 102.493 - 18.11 = 84.383
•The predicted grade is approximately 84.4.
•Graphing the Line:
•To graph the line, select two values for x within the range of the
data, calculate their corresponding y-hat values, plot the two
points, and draw a line through them