0% found this document useful (0 votes)
8 views8 pages

Understanding Scatter Diagrams & Regression

Uploaded by

sujatagautam378
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views8 pages

Understanding Scatter Diagrams & Regression

Uploaded by

sujatagautam378
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Scatter Diagram:

The paired observations on two variables and are denoted by (x, y). The graph of observed values of
two variables is used to identify the relationship between two variables. Such a graph is called scatter
diagram. Usually, one variable depends to some degree on the other, then the dependent variable is
taken on the vertical (y) axis and the other variable is taken on horizontal (x) axis. The pattern of
points plotted on the scatter diagram indicates whether the variables are related.
The following scatter diagrams illustrate the possible form and relationship between the two
variables.

Simple linear regression:


Simple linear regression has only one x and one y variable i.e. one dependent and only one
independent variable. Multiple linear regression has one y and two or more x variables i.e. only one
dependent variable and two or more independent variables. For instance, when we predict rent based
on square feet alone that is simple linear regression. When we predict rent based on square feet and
age of the building that is an example of multiple linear regression.
The sum of squares of errors is used to represent the total error committed. Therefore, the linear
equation is obtained by minimizing sum of squares of errors. This method of deriving the equation is
called method of least squares.
The method of least squares obtains the values of the constant of regression equation y on x by
minimizing sum of squares of errors as s = ∑e 2 = ∑ (y – a - bx)2
And differentiating this by a and b to get regression coefficient and y intercept.
Regression coefficient y on x
𝑐𝑜𝑣(𝑥,𝑦)
b= denoted by b yx
𝑣𝑎𝑟(𝑥)
Regression coefficient x on y
𝑐𝑜𝑣(𝑥,𝑦)
b= 𝑣𝑎𝑟(𝑦)
denoted by b xy
Independent variable:
The variable which is used for prediction or estimation of dependent variable is called independent variable.
Dependent variable:
The variable who’s value is to be predicted or estimated from the known value of the other variable(s) is called
dependent variable.

SIMPLE LINEAR REGRESSION


When there is only single independent variable in the linear regression model it is called simple linear regression
model.
Simple linear regression model
Y = β0 + β1X + ε
where Y = dependent variable
X = independent variable
β0 = intercept of regression line
β1 = slope of regression line
ε =error component
Suppose we have 'n' pairs of observation of 'x' and 'y'
ie. (xi, yi) i = 1,2,..., n
For the individual observation, we can write the above model as
yi = β0 + β1xi+ εi i = 1, 2, ..., n
where β0 = intercept of a regression line.
β1 = slope of regression line
For every unit increase in X, the slope tells us how much y would increase.
The slope of the regression line tells us the change in dependent variable Y per unit increase in independent
variable x

Multiple Linear regression


A linear regression model that involves more than one explanatory (independent) variables is called a
multiple linear regression model.
Simple linear regression model
Y = β 0 + β 1 X 1 + β 2X 2 + ε
where Y = dependent variable
X = independent variable
β0 = intercept of regression line
β1, β2 = regression coefficients
ε =error component
Factor analysis

Factor analysis is a data reduction technique. It is used with cross-sectional data to identify variables
in the data that form theoretically meaningful subgroups. These subgroups are expected to be
independent from each other. Variables in a subset are related with each other and are unrela ted with
variables in other subsets. These subsets are called factors.
Initially, data is obtained for large number of participants on many variables. Then the correlation
matrix for the data is obtained and either using PCA or FA, the factors are extracted. Then a decision
about the number of factors to be retained is made and those many factors are rotated. The results are
then interpreted. Researchers often try different number of factors to understand interpretability and
interpretation of factors is the most important part. If the solutions are interpretable, then it is a
useful exercise.

Correlation matrix: This matrix (or covariance matrix) generally refers to the observed correlations
(or covariances) among variables obtained on the data. The truncated factor analysis solutions always
have less information than the information in a correlation matrix. After factor analysis, a well-fitting
factor analysis solution should be able to give us the correlation matrix back that is closer to the
observed correlation. The correlation matrix obtained from factor analysis solution is called a
reproduced correlation matrix. The difference between the observed and reproduced correlation
matrices is called a residual correlation matrix. The residual matrix contains the variance that the
factor analysis did not explain.

Extraction: This is the process of transforming variance from a correlation matrix into factors or
components. The PCA method provides components and FA provides factors. The PCA extracts from
the observed correlation matrix, i.e. all observed variance is analyzed. The FA extracts from shared
variance or common variance. A shared or common variance which is common to all variables.

Rotation: This is a process of making the selected factors or components more interpretable. Rotated
and unrotated matrices are mathematically identical, but with an appropriate-rotation method of
correct number of factors, the rotation improves interpretability. There are two types of rotation
methods: orthogonal and oblique. The orthogonal rotation method produces factors that are
independent or uncorrelated, and oblique rotation produces factors that are dependent or correlated.

Factor pattern matrix: This matrix is a "variables by factors" matrix. If we retain two factors for a
six-variable correlation matrix, then the factor pattern matrix is a 6x2 matrix, and 12 elements of this
matrix are called factor loadings. The number of factors is expected to be much smaller than the
variables. Loading refers to the coefficient describing the unique relation between a variable and a
factor. The factor pattern is evaluated for interpretation in the case of orthogonal rotation. The factor
pattern and factor structure matrices are identical for orthogonal factors.

Factor structure matrix: This matrix is a "variables by factors" matrix describing the correlations of
variables with factors. In the case of oblique rotations, the factor pattern and factor structure matrices
are different. In the case of orthogonal rotations, they are identical.
Phi matrix: This is the inter-factor correlation matrix. In the case of orthogonal factors, the phi matrix
is an identity matrix and in the case of orthogonal factors, off-diagonal elements describe inter-factor
correlations.

Confirmatory Factor Analysis (CFA)


The CFA is a special case of Structural Equation Modeling (SEM). Classically, the SEM has two
types of models: one, measurement model and, two, full-SEM model. The measurement model of the
SEM is CFA. The CFA is a method of factor analysis that tests the theory. The exploratory factor
analysis does not utilize any statistical theory to test a hypothesis that the obtained factor structure
resembles the theoretically expected factor structure. At the most, the theory guides in deciding the
number of factors to be rotated in the Exploratory factor analysis(EFA). The CFA is a process that
"tests" a theory about the structure of the variables. The CFA in that sense a "scientific" method. The
purpose of EFA is primarily to describe the underlying structure, whereas the purpose of CFA is to
verify a hypothesis about the underlying structure. Unlike EFA, the expected factor structure is pre-
specified in the CFA.
The logic of CFA is based on the assumption (which is very true in the case of psychological
variables and their measures) that measured variables (also called observed variables, manifest
variables, or indicators) are imperfect indicators of latent variables. For example, Intelligence is the
internal structure of mind that is not observable. Suppose an intelligence test has 30 items. An
individual can pass or fails on each item. The item's scores are not intelligence themselves (and for
that matter nor is the sum of scores on all the items or any standardized conversion of that total score
like, z, T, and so on). Someone getting a score of IQ=120 is an indication of just a relative position.
The item scores of individuals are the manifest behavior and not intelligence themselves. The central
idea of latent variables is that the score on 30 items is not intelligence, but it is intelligence (a latent
variable) that is causing the score on these 30 items. A latent variable is an unobservable, underlying
construct that causes the observable and measured variables (also called manifest variables). Since
the latent variables are unobservable, the measurement of the latent variables is not possible.
However, the observed scores are not just caused by the latent variables, but they can also be a
function of random errors. So, the complete model looks as follows.
The latent variables are primary cause of manifest variables. The manifest variables can also be
caused by random errors and so are considered as imperfect indicators.
Observed Variable = Latent Variable + Random Error
Steps in CFA:
The researcher employing CFA needs to take the following steps:
1. Have theory: The CFA is a hypothesis-testing method. The CFA cannot be conducted in the
absence of a theory-driven hypothesis (or at least a hunch). The theory further specifies that the two
latent variables are unrelated.
2. Get data: In the wake of causal theory, the researcher should get data on observable variables or
indicators in the model. The sample size should be large. In addition to test the CFA hypothesis, the
CFA method can be used to test various other questions (mean structures, multiple-groups, and do
on), and accordingly suitable modifications are made in the data- collection process.
3. Specify model: The theoretical model is specified as a linear model.
4. Test for identification: Identification is an important issue in CFA. A correct solution of
identification problem leads to appropriate estimation of model parameters .
5. Estimate model parameters: The model has various parameters. The parameters are esti mated
by one of the parameter estimation methods. Most popular method of parameter estimation is ML.
However, other methods might do the job better under specific circumstances.
6. Statistical test and fit indices: The statistical test of the "fit" between model and data is carried
out. The statistical test of CFA has certain limitations.
7. Compare different models: One of the alternatives is to compare different competing theoretical
models and choose the best among them.
8. Interpret and conclude: Once the results are obtained, the researcher has to carefully evaluate the
results and decide whether the hypothesis in question is to be retained or not.

Assumptions
Certain assumptions are required to estimate the structural coefficients. The assumptions are
1. E(X) = E(ξ) = 0. The mean of the observed and latent variables is zero.
2. The relationships between the observed (X) and latent variables (ξ) are linear.
3. There are assumptions about measurement error:
a) E(δ)=0. The errors have mean zero.
b) The errors have constant variance across observations.
c) The errors are independent across observations.
d) Ε(ξδ') = Ε(δ ξ') = 0. The errors are uncorrelated with latent variables.

ANOVA Introduction:
A composite procedure for testing simultaneously the difference between several sample means is
known as the analysis of variance. It helps us to know whether any of the differences between the
means of the given samples are significant. If the answer is 'yes', we examine pairs (with the help of
the test to see just where the significant differences lie. If the answer is 'no', we do not proceed
further.
In such a test, as the name implies, we usually deal with the analysis of the variances, Variance is
simply the arithmetic average of the squared deviation from their means. In other words, it is the
square of the standard deviation (variance = σ²). Variance has a quality which makes it especially
useful. It has an additive property, which the standard deviation with its square root does not possess.
Variance on this account can be added up and broken down into components. Hence, the term
'analysis of variance' deals with the task of analyzing of breaking up the total variance of a large
sample or a population consisting of a number of equal groups or sub-samples into two components
(two kinds of variances), given as follows:
1. "Within-groups" variance. This is the average variance of the members of each group around their
respective group means, i.e. the mean value of the scores in a sample (as members of each group may
vary among themselves).
2. "Between-groups" variance. This represents the variance of group means around the total or grand
mean of all groups, i.e. the best estimate of the population mean (as the group means may vary
considerably from each other)

You might also like