0% found this document useful (0 votes)
5 views5 pages

Regression Analysis of Student Scores

The assignment focuses on calculating the least square regression lines for Statistics and Economics scores of 12 students. The regression line of Y on X is Y=0.5424X+19.54, while the regression line of X on Y is X=1.236Y−3.51. The document includes detailed calculations for means, deviations, and regression coefficients.

Uploaded by

chata4669
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views5 pages

Regression Analysis of Student Scores

The assignment focuses on calculating the least square regression lines for Statistics and Economics scores of 12 students. The regression line of Y on X is Y=0.5424X+19.54, while the regression line of X on Y is X=1.236Y−3.51. The document includes detailed calculations for means, deviations, and regression coefficients.

Uploaded by

chata4669
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

VIRTUAL UNIVERSITY OF PAKISTAN

BWP CAMPUS
ASSIGNMENT NO: 1

COURSE TITLE :Statistics & Probability

COURSE CODE : STA 301

SEMESTER : Spring

STUDENT NAME : Muhammad Umair Faryad

STUDENT ID :bc250225601

ASSIGNMENT TITLE :Least Square Regression

SUBMISSION DATE :12 May 2025


QUESTION:

The following table shows the final marks in Statistics and Economic obtained by 12
students selected at random from a large group of students.

Statistics Economics
(X) (Y)
64 57
71 59
53 49
67 62
55 51
58 50
77 55
57 48
56 52
51 42
76 61
68 57

Require: Find the least square regression lines of Y on X and X on Y

n=12 ,
Xi=Statistics scores
Yi=Economics scores
Step 1: Calculate Mean of X and Y
Formula

X=∑X /n
Solve

∑X=753
= 753/12
( =62.75)

Formula

Y=∑Y/n
Solve

∑Y=643
=643/12
( =53.58)
Step 2: Create Table for Deviation Calculations

X Y

64 57 1.25 3.42 1.5625 11.6964 4.275


71 59 8.25 5.42 68.0625 29.3764 44.715
53 49 -9.75 -4.58 95.0625 20.9764 44.655
67 62 4.25 8.42 18.0625 70.8964 35.785
55 51 -7.75 -2.58 60.0625 6.6564 19.995
58 50 -4.75 -3.58 22.5625 12.8164 17.005
77 55 14.25 1.42 203.0625 2.0164 20.235
57 48 -5.75 -5.58 33.0625 31.1364 32.085
56 52 -6.75 -1.58 45.5625 2.4964 10.665
-
51 42 -11.58 138.0625 134.0964 136.035
11.75
76 61 13.25 7.42 175.5625 55.0564 98.285
68 57 5.25 3.42 27.5625 11.6964 17.955
 X  X ^ 2  ( X  X )(Y  Y )

Step 3: Calculate Required Sums

 ∑(X−Xˉ)2=888.1875

 ∑(Y−Yˉ)2=389.6727

 ∑(X−Xˉ)(Y−Yˉ)=481.7

Step 4: Calculate Regression Coefficients

1:Regression of Y on X:
or

= 481.7/888.1875
=0.5424
Y−Yˉ=bYX(X−Xˉ)⇒Y=0.5424X+(Yˉ−0.5424Xˉ)=0.5424X+(53.58−0.5424×62.75)
Y=0.5424X+53.58−34.0416=0.5424X+19.5384
Regression Line of Y on X

Y=0.5424X+19.54

:Regression of X on Y:
2
X  X bxy (Y  Y ) or X bxy.Y  bxy.Y  X

= 481.7/ 389.6727

= 1.236
X−Xˉ=bXY(Y−Yˉ)⇒X=1.236Y+(Xˉ−1.236Yˉ)=1.236Y+(62.75−1.236×53.
58)
X=1.236Y+62.75−66.2599=1.236Y−3.5099

Regression Line of X on Y:
X=1.236Y−3.51

Final Answer:

Regression line of Y on X:

Y=0.5424X+19.54

· Regression line of X on Y:

X=1.236Y−3.51

Scatter Diagram:
The red line representing the regression of Y on X
he green line representing the regression of X on Y

Common questions

Powered by AI

The regression line of Y on X (Y=0.5424X+19.54) is used to predict the value of Y based on a given value of X, assuming that X is the independent variable influencing Y. Conversely, the regression line of X on Y (X=1.236Y−3.51) is used to estimate X based on a known Y, considering Y as the predictor. These lines differ primarily in their orientation of prediction and their roles in analysis, with each offering specific insights depending on which variable is considered independent or dependent.

A regression coefficient greater than 1, such as the coefficient of 1.236 for X on Y, indicates that for every one-unit increase in the independent variable (Y), the dependent variable (X) increases by more than one unit. This implies a strong positive association where changes in Y strongly impact X, suggesting a greater sensitivity of X to changes in Y.

Calculating the mean of X and Y is fundamental in least square regression as it serves as a balance point or central value around which deviations are measured. This facilitates the calculation of the regression coefficients, allowing the regression line to predict values by minimizing the squares of deviations between observed and predicted values. The mean ensures that the sum of deviations (errors) from the regression line equals zero, maintaining an overall balance in prediction errors.

The regression coefficient for Y on X (bYX) is calculated by dividing the sum of the product of the deviations (∑(X−X̄)(Y−Ȳ)) by the sum of squared deviations of X (∑(X−X̄)^2). In this case, bYX = 481.7 / 888.1875 = 0.5424. This coefficient represents the change in variable Y for a one-unit increment in variable X, signifying that, on average, an increase of 1 unit in Statistics marks causes an increase of 0.5424 units in Economics marks. This indicates a moderate positive linear relationship between Statistics and Economics marks.

A scatter diagram visually displays the relationship between two variables, with each point representing paired data values. By plotting regression lines, such as the regression of Y on X and X on Y, on a scatter diagram, one can visually assess the correlation and strength of the relationship. The lines indicate the predicted trends, while the scatter of points around these lines reveals the variability and goodness of fit, helping to intuitively understand the degree of linear association between the variables.

Understanding regression analysis benefits decision-making by providing quantitative insights into relationships between variables, enabling prediction of outcomes and identification of key factors affecting results. It aids in resource allocation by highlighting efficient predictors, supports risk assessment through estimation of variance and uncertainty, and informs policy analysis by quantifying impacts of changes in variables. In business and research, it contributes to strategic planning and improves forecasting accuracy, ultimately enhancing data-driven decision-making processes across various domains.

The least squares method minimizes error by finding the line through data points that minimizes the sum of the squares of the vertical distances (errors) between observed values and predicted values. By squaring deviations, the method penalizes larger errors and focuses on reducing overall prediction inaccuracies. This minimization results in a line, the best linear unbiased estimator (BLUE), which represents the most accurate model of the relationship between variables according to the least squares criterion.

A negative regression coefficient indicates an inverse relationship between the predictor and response variable. This means that as the independent variable increases, the dependent variable tends to decrease. For example, in a scenario where variables are students' study hours and stress levels, a negative coefficient would suggest that increased study hours lead to decreased stress levels. This interpretation reveals the direction of a relationship crucial for understanding causality or correlation nuances in regression analysis.

The sums of squared deviations, such as ∑(X−X̄)^2 and ∑(Y−Ȳ)^2, measure the variance within each variable, essential for assessing the spread and consistency of data points relative to their means. The cross products ∑(X−X̄)(Y−Ȳ) capture the extent and direction of simultaneous deviations from their means, reflecting the degree of correlation between variables. Together, these sums inform the calculation of regression coefficients, indicating the strength and nature of the relationship and providing key insights for predictive modeling.

Deviation calculations are crucial in regression analysis as they measure the extent to which individual data points deviate from the mean. These deviations are used to calculate the sums required to determine the regression coefficients, which define the line of best fit. Specifically, calculating (X−X̄)^2 helps find the total variance in X, while (X−X̄)(Y−Ȳ) provides the covariance, indicating how X and Y co-vary. These calculations form the foundation of the least squares method, minimizing the error between observed and estimated values.

You might also like