S
TATISTICAL METHOD
MULTIPLE LINEAR
REGRESSION
BY
UBAY LAKSAEGA SUBROTO 142241210
MADE AGNITYA PUTRI 142241066
MUHAMMAD FADELH 142231302
BAIHAQI
MANAGEMENT
FACULTY OF ECONOMIC AND BUSINESS
AIRLANGGA UNIVERSITY
SURABAYA
2025/2026
CHAPTER I...........................................................................................................................................3
INTRODUCTION.................................................................................................................................3
CHAPTER II..........................................................................................................................................5
METHODOLOGY................................................................................................................................5
2.1 Research Design and Data Collection.........................................................................................5
2.2 Variable Definition and Measurement........................................................................................5
2.3 Structural Equation Modeling (SEM) Procedure........................................................................6
2.4 Data Analysis Software and Estimation Method........................................................................6
CHAPTER III........................................................................................................................................7
RESULT & ANALYSIS........................................................................................................................7
3.1 Descriptive Statistics of Respondents and Variables..................................................................7
3.2 Measurement Model Evaluation.................................................................................................7
3.2.1 Outer Loadings and Indicator Reliability.................................................................................7
3.2.2 Internal Consistency Reliability...............................................................................................7
3.2.3 Convergent Validity.................................................................................................................8
3.2.4 Discriminant Validity...............................................................................................................8
3.3 Structural Model Evaluation.......................................................................................................9
3.3.1 Coefficient of Determination...................................................................................................9
3.3.2 Path Coefficients and Hypothesis Testing...............................................................................9
3.4 Mediation Analysis (Indirect and Special Indirect Effects)........................................................9
3.5 Discussion of Findings..............................................................................................................10
CHAPTER I
INTRODUCTION
Multiple linear regression (MLR) is one of the most widely used statistical techniques for
analyzing the relationship between one dependent variable and multiple independent variables.
In modern research—especially in economics, business, social sciences, and data-driven
decision–making—MLR provides a systematic framework to understand how several predictors
jointly influence an outcome. This method is essential not only for explanation but also for
prediction, allowing researchers to estimate the magnitude and direction of relationships while
controlling for other variables in the model.
The increasing availability of quantitative data in academic and professional environments
has made MLR a critical analytical tool. Organizations often use regression-based insights to
make strategic decisions, forecast performance, identify key influencing factors, and design
more effective policies. Within the academic context, MLR helps students and researchers
break down complex real-world problems into measurable components, enabling them to test
hypotheses using empirical evidence rather than assumptions.
This project aims to apply multiple linear regression to examine the influence of several
independent variables on a selected dependent variable (to be described later in the
methodology section). The focus of this study is to build a valid and reliable model that
demonstrates how these variables interact statistically. Through the use of structured analysis—
including descriptive statistics, measurement model evaluation, reliability assessment,
coefficient of determination (R²), path coefficient analysis, and mediation testing—this proposal
seeks to produce a comprehensive and academically rigorous explanation of the regression
findings.
In addition, this project is designed to strengthen analytical skills by implementing a
complete research cycle: from defining variables, collecting data, preparing the dataset,
conducting regression analysis, interpreting results, and presenting findings in a clear and
professional manner. Ultimately, the results of this study are expected to contribute to a deeper
understanding of how multiple factors jointly affect an outcome, and how regression techniques
can be used to support evidence-based decision-making.
CHAPTER II
METHODOLOGY
2.1 Research Design and Data Collection
This study uses a quantitative research design with a cross-sectional approach, meaning data
are collected at one point in time. Multiple linear regression is employed to examine the
relationship between one dependent variable and several independent variables.
Data collection involves distributing a structured questionnaire to respondents using a non-
probability sampling method, typically purposive or convenience sampling, depending on the
population characteristics. Respondents are selected based on predefined criteria relevant to the
research topic (e.g., students, consumers, employees). All responses are recorded using Likert
scales to ensure measurability and consistency across variables.
Before analysis, the data undergo cleaning procedures, such as removing incomplete responses,
detecting outliers, and checking normality. This ensures the dataset meets the statistical
assumptions required for regression analysis.
2.1 Variable Definition and Measurement
Variables are operationalized based on established theories and prior empirical research. Each
variable is measured using multiple indicators to ensure accuracy.
● Dependent Variable (Y):
Defined as the main outcome the study seeks to predict. Measurement uses Likert-scale statements
reflecting the construct’s characteristics.
● Independent Variables (X₁, X₂, X₃, …):
These are factors hypothesized to influence the dependent variable. Each is measured through
several validated items adapted from previous studies.
All indicators use a 5-point Likert scale (1 = Strongly Disagree, 5 = Strongly Agree) to ensure
uniformity. Reliability and validity tests are conducted to confirm measurement accuracy.
2.2 Data Analysis
Data analysis is conducted using statistical software such as SPSS, R, or Python.
Procedures include:
● Descriptive statistics (mean, standard deviation, frequency)
● Assumption testing
● Normality
● Multicollinearity (VIF & Tolerance)
● Heteroskedasticity
● Linearity
● Multiple Linear Regression Analysis
● Estimation of regression coefficients
● Evaluation of model significance (F-test)
● Evaluation of individual predictors (t-test)
● Interpretation of β coefficients
● Effect size & interpretation
● R² and Adjusted R²
● Standardized beta values
The output from these tests determines whether the independent variables significantly influence
the dependent variable.
CHAPTER III
RESULT &
ANALYSIS
3.1 Descriptive Statistics of Respondents and Variables
Descriptive statistics summarize the demographic profile of respondents (age, gender,
education) and provide an overview of the central tendencies of each research variable.
For variables, mean values describe overall respondent tendencies, while standard deviation
reflects data dispersion. This initial profile helps interpret how respondents generally perceive
each construct.
3.2 Measurement Model Evaluation
This analysis evaluates whether the measurement items accurately reflect the intended
variables. Although regression does not require a strict measurement model like SEM, reliability
and validity assessments ensure that each construct is statistically sound.
3.2.1 Outer Loadings and Indicator Reliability
Outer loadings assess the strength of each indicator in representing its latent variable.
Values above 0.70 indicate good convergent validity, while values between 0.50–0.70 may still be
acceptable if overall construct reliability is strong.
Indicators with very low loadings may be removed to improve validity and measurement
precision.
3.2.2 Internal Consistency Reliability
Internal consistency ensures all items measure the same underlying construct.
Tests used:
● Cronbach’s Alpha ( 0.70)
● Composite Reliability (CR 0.70)
● High values indicate consistent responses across indicators.
3.3 Structural Model Evaluation
The structural model examines the strength of relationships between variables.
Key criteria include:
● Path coefficients (β values)
● Significance levels (p-values)
● t-statistics from bootstrapping or regression output
This evaluation tests whether each independent variable exerts a statistically significant influence
on the dependent variable.
3.3.1 Coefficient of Determination
R² measures how much the independent variables collectively explain the variation in the
dependent variable.
● R² 0.25 = weak
● R² 0.50 = moderate
● R² 0.75 = strong
Adjusted R² provides a more accurate estimate by adjusting for the number of predictors.
3.3.2 Path Coefficients and Hypothesis Testing
Hypothesis testing uses t-tests and p-values to determine if each relationship is supported.
If p < 0.05: the hypothesis is accepted
If p > 0.05: the hypothesis is rejected
The sign of β indicates the direction of influence (positive or negative). The magnitude of β shows
how strongly each variable contributes to predicting the outcome.
3.1 Mediation Analysis
If the model includes a mediating variable, the analysis tests both:
Direct Effects: X Y
Indirect Effects: X M Y
Bootstrapping is used to evaluate the significance of mediation. If the indirect effect is
significant, mediation exists; if both direct and indirect effects are significant, partial
mediation occurs.
3.2 Discussion of Findings
This section interprets the statistical results in the context of theory and previous research. Key
points include:
● Whether findings support the hypotheses
● How each independent variable contributes to predicting the dependent variable
● Practical implications of the research
● Alignment or contradiction with past studies
● Recommendations for practitioners, organizations, or policymakers
The discussion connects numerical data to meaningful insights, ensuring the study provides value
beyond statistical output
CHAPTER IV
FINDINGS AND
INTERPRETATION
3.3 Histogram of FDI
This histogram displays the frequency distribution of the FDI (Foreign Direct Investment)
variable for all 55 observations
3.4 Normal Q-Q Plot of FDI
A Q–Q plot compares the actual distribution of FDI to a theoretical normal distribution. In this plot,
the points fall close to the diagonal line with only small deviations at the ends, indicating that FDI is
approximately normally distributed and meets the normality assumption for MLR.
3.5 Descriptive Statistics, Correlations, and Model Summary
This output presents the descriptive statistics for FDI, LEXPORT, and LGDP, showing their
mean and standard deviation to describe the basic characteristics of each variable. The
correlation matrix indicates that FDI has very weak relationships with both LEXPORT and
LGDP, while LEXPORT and LGDP are highly correlated, suggesting possible
multicollinearity. The variables entered/removed section confirms that both predictors were
included in the model. Meanwhile, the model summary shows a weak overall relationship (R
= 0.349) with the predictors explaining only 12.2% of the variation in FDI (R² = 0.122), and
the Durbin–Watson value of 0.813 suggests some positive autocorrelation in the residuals.
3.6 ANOVA, Coefficients Table, Diagnostics, Residual Statistics
This output shows the overall significance of the regression model, the effects of each
predictor, and diagnostic checks for data quality. The ANOVA table indicates that the model
is statistically significant at the 5% level (Sig. = 0.034), meaning at least one predictor
significantly influences FDI. The coefficients table shows that LEXPORT has a significant
negative effect on FDI (B = –1.187, Sig. = 0.012), while LGDP has a significant positive
effect (B = 2.119, Sig. = 0.010), and the constant is also significant. The casewise diagnostics
identify extreme residuals, with case 31 standing out as a potential outlier (residual = –3.746).
Finally, the residual statistics help assess regression assumptions, showing standardized
residuals ranging from –3.13 to 2.86, which is generally acceptable but indicates a few
unusually large values.
3.7 Normal P-P Plot of Regression Residual
This plot compares the observed and expected cumulative probabilities of the standardized
residuals, and the points fall very close to the diagonal line, indicating that the residuals are
normally distributed. This normality supports a key MLR assumption and provides evidence
that the model is statistically valid even though its explanatory power is relatively low.