0% found this document useful (0 votes)
4 views25 pages

Factor Analysis

The document discusses factor analysis, a statistical technique used to identify underlying relationships between variables in a dataset. It outlines the process of creating a score matrix, calculating factor loadings, and interpreting results through methods such as centroid and principal components. Additionally, it emphasizes the importance of rotation techniques to reveal different data structures and the computation of eigenvalues and communalities to assess the significance of factors.

Uploaded by

dumplingace638
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
4 views25 pages

Factor Analysis

The document discusses factor analysis, a statistical technique used to identify underlying relationships between variables in a dataset. It outlines the process of creating a score matrix, calculating factor loadings, and interpreting results through methods such as centroid and principal components. Additionally, it emphasizes the importance of rotation techniques to reveal different data structures and the computation of eigenvalues and communalities to assess the significance of factors.

Uploaded by

dumplingace638
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
‘pinferfrom these some factor (such as social Class) which summarises the comm, Sidfourvariables. The technique used for such purpose is generally described a jonality of all the ‘tween the particular variable and the in the variable and the factors. score matrix), med ‘48S. The matrix contains the scores of N persons © Thus a, is the score of ‘onmeasure a, a, is the score of person 2 on measure 2, The score matrix then take the form as shown following of person Von & Research Meth, ay SCORE MATRIX (ot Matrix 5) Measures (variables) [ee | 1 a, bg, k, | 2 a, 4G ky | | 5 a, an) & | Persons (objects) . | | y __ It is assumed that scores on each measure are standardized [i.e., x, = (x _ X,)?/6,1. This ( being so, the sum of scores in any column of the matrix, §, is zero and the variance Of scores in any column is 1.0. Then factors (a factor is any linear combination of the Variables in a data ‘Matrix and can be stated in a general way like: A = Wa + Mb +... + WR) are obtained (by any method of E, factoring). After this, we work out factor loadings (i.., factor-variable Correlations). Then communal, 8) symbolized as /?, the eigen value and the total sum of squares are obtained and the. results interpreted. ig q "pr For realistic results, we Tesort to the technique of rotation, because such rotations reveal different Structures in the data. Finally, factor scores are obtained which hel; in explaining what the factors P. pl mean. They also facilitate comparison amon; roups of items as roups. With factor scores, one can 18 group: Sroup: also perform several other multivariate analyses such as multiple regression, cluster analysis, multiple (vi) discriminant analysis, etc, | Woes FACTOR ANALYSIS ‘ suc ‘ There are several methods of factor analysis, but they do not necessarily give same result, ie : ‘ factor analysis is not a Single unique method but a set of techniques, Important methods of fac F analysis are: ct (the centroid method: Mise ') the principal components method; = (0) the maximum likelihood method, . We Per =e 6. riate that some Sean ; Before we describe these different methods of factor analysis, it seems appropri , t asic tess relating to facto, analysis be well understood. vars S24 Ce, . jpserved ve bat n (0 Factor: A factor is an tunderlying dimension that account for ee aniiersad= Mh et, There can be one of more factors, depending upon the nature of the study so ae Noa q Of variables involved int ok. a : F analysis E he ot-loadings are those v, y alues whi ¢ to cach one of the factors discoverey sree xplain pe relate ae lations. In fact core absolute size (ra pretation ofa factor. itist in the inten y Ur) Communality, symbolized 5 li a8 PP, sho ‘-counted for the underlying factor taken feather a ‘hat not much of the variable is Teft over after whates vnsideration. It is worked out in respect of each varabie Jp of the ih variable = (ith factor loading of f + (ith factor loading of f 2, Y%,) en value (0F latent root): When we take the sum of squared Scope, «This relating to a factor, then such sum is referred to as Eigen \ he tot loading 2 mat indicates the relative importance of each factor in accountinn (eo Eiaen Ym Writ ang variables being analysed inting for the particular s. ethog omtmunal vf squares: When eigen values of all factors are totalled, the result inte a, termed as the total sum of squares. This value, when divided by the manhe no, ot (involved ina study), results in an index that shows how the particultr solution scans ke ifereny ‘what all the variables taken together represent. Ifthe variables are all yery different from © factors each other, this index will be low. If they fall into one or more highly redundant groups, and 8, one can ifthe extracted factors account for all the groups, the index will then approach unity S Multiple ) Rotation: Rotation, in the context of factor analysis, is something like staining a microscope slide, Just as different stains on it reveal different structures inthe tissue, different rotatio reveal different structures in the data, Though different rotations give results that appearto be entirely different, but from a statistical point of view, all results ace taken as equal, none ee superior or inferior to others. However, from the standpoint of making sense ofthe results of factor analysis, one must select the right rotation, If the factors are independent othogons! Assuch rotation is done and if the factors are correlated, an oblique rotation is made, Com nly Qf factor for each variables will remain undisturbed regardless of rotation but the eigen values change as result of rotation, es 40 each respondent gets high ii) Factor scores: Factor score represents the degree ae ees help explain seores onthe group of items that load high on each clr Fe scan be What the factors mean. With such scores, severs! performed. rs dl We can now take up the important methods of factor anty® some 15.2.1 Centroid Method x4 until about Jes quite frequently 3 nize ble Thurstone, was hoards roma nber —__tismethod of factor analysis, develoned by Lt gmputes. Tecan ess oF OH 95 ly ‘ before the advent of arg eapacity iene nod which etme the sum of loadings, disregarding signs: iis £360) Research Methoc combinations in which all weights are either ‘ple, can be easily understood and becomes easy to understand the loadings for each factor in turn. It is defined by linear +1.0 or=1.0. The main merit ofthis method is that itis relatively sim involves simpler computations. If one understands this method, it mechanics involved in other methods of factor analysis. Various steps involved in this method are as follows: (}) This method starts with the computation of a matrix of correlations, R, wherein unities die place in the diagonal spaces. The product moment formula is used for working out the correlation coefficients. (id) Ifthe correlation matrix so obtained happens to be positive manifold (ie. disregarding the diagonal elements each variable has a large sum of positive correlations than of negativg correlations), the centroid method requires that the weights for all variables be =1.0_ In other words, the variables are not weighted; they are simply summed. But in case the correlation matrix is not a positive manifold, then reflections must be made before the first centroid factor is obtained. (iii) The first centroid factor is determined as under: (a) The sum of the coefficients (including the diagonal unity) in each eolumn of the correlation matrix is worked out. (b) Then the sum of these column sums (7) is obtained. (c) The sum of each column obtained as per (a) above is divided by the square root of T obtained in (b) above, resulting in what are called centroid loadings. This way each centroid loading (one loading for one variable) is computed. The full set of loadings so obtained constitute the first centroid factor (say A). (iv) To obtain second centroid factor (say B), one must first obtain a matrix of residual coefficients. For this purpose, the loadings for the two variables on the first centroid factor are multiplied. This is done for all possible pairs of variables (in each diagonal space is the square of the particular factor loading). The resulting, matrix of factor cross products may be named as Q,. Then Q, is subtracted clement by element from the original matrix of correlation, R, and the result is the first matrix of residual coefficients, R,. One should understand the nature of the elements in R, matrix. Each diagonal element isa partial variance ic, the variance that remains after the influence of the first factor is partialed. Each off diagonal element is a partial co-variance i.e., the covariance between two variables after the influence of the first factor is removed. This can be verified by looking at the partial correlation coefficient between any two variables say | and 2 when factor 4 is held constant Tig Naa Na = ‘Analysis @ found i in R, corresponding to the entry for ane - 7 anal element for variable | in R. | Square of the term on the left is ¢ 1; Likewise the partia tis exactly what is found partial variance for ve di iat it Vi ce for that variable in the residual matrix.) 2 is found in the diagonal : agonal since in R, the diagonal terms are oyariances, it is easy to convert the satve aol ¥ pier ee Or able to diagonal terms are partial oa am: tpi ms : \ ; ae ea leo! artial correlations, For this purpos n dividing the elements i ow by the square-root of the diagonal cle: oy has - oo 4 ra pier: poate cle nal element for that row thet each column by iquare- 1¢ diagonal element for tha t for that column. 1 “After obtaining R,,one must reflect some of the variables sarables are given negative ignsin the sum (Thisis em mrcenine lterety thatsome ofthe Mould be to obtain a reflected matrix, R', which Fa in Forany variable which isso reflected thes Me teeter ee ee Be atastteare changed When this is dono, eect ae that column and row of the ee pe sre tenn in eticual way clveady explained inthe eee Bete fee ore varablezwhich were reloacd mu bogien neg rhe ful tol i ; : 3 fll eto ca Pee cretion k, feet (oy ® ‘Thus loadings on the second (v) For subsequent factors (C, D. etc.) the same process outlined above is repeated. After the second centroid factor is obtained, cross products are computed forming, matrix, his is then subtracted from R, (and not from 2’) resulting in R,, To obtain a third factor (©. vo chould operate on R, inthe samo way as on, Fits some of th arabes ule have res roflecte to maximize the sum of loadings, which would produce, ndings would be computed from R',asthey were from R), Ags. would be necessary to give negative signs to the loadings of variables svhigh were reflected which would result in third centroid factor (C). We may now illustrate this method Py #0 example, Example 15.1: Givenis the follow! diagonal spaces: Variables ing correlation matrix, relating eight variables ith unites in the iM 724 120 152 1,000 Variables er CHAPTER 15 Research Met @ sst and second centre Using the centroid method of factor analysis, work out the the above information Solution: Given correlation atrix, R, isa positive manifold and as such the weight ly, we calculate the first centroid factor (4) as under: 1.0. Accordi Table 15.4(a) 1.000 2] 0709 | 1.000 | .051 3 | 0204 | 051 | 1.000 Variabtes 4 | 0081 | .089_|_.671 3__| 0626} .581_| .123 6} e113 | 098 [089 | 798 | 7.000 801 | 7 {0.135 {083 | 582 | 613 | 201.801 | 1.000 | 152 | # [0774 o72 [ANN | 724 [120 | 152 [7.000] Column sums | 3.662 | 3.263 [3.392 | 3.385 [3.324 | 3.6601 387) 3605 71 | Sum ofthe column sums(7)=27.884 », VT 81 3 3302 3385 3324 3666 281" 5.281’ 5281’ 5281’ 5.281 ~ 693, 618, 642, 641, .629, .694, .679, 683 | Meet A eae SS 0.618 0.642 0.641 0.629 0.694 | 7c ea L 0.683 a nd ge) roid factor B, we first of all develop (as shown on the next P28 matrix of factor cross product Q: rel To obtain the second cent Analysis & First Matrix of Factor Cross Product (Q,) First centroid 0.693 0.618 0.642 0.641 0.629 0.694 0.679 0.683 factor A pe 0,693 0.618 0.642 ‘ 0.641 | 0.629 0.694 0.679 0.683, 0.480 0.428 0.445 0.444 0.436 0.481 0.471 0.473 0.428 0.382 0.397 0.396 0.389 0.429 0.420 0.422 0.445 0.397 0.412 0.412 0.404 0.446 0.436 0.438 0.444 0.396 0.412 0.411 0.403 0.445 0.435 0.438 0.436 0.389 0.404 0.403 0.396 0.437 0.427 0.430 0.481 0.429 0.446 0.445 0.437 0.482 0.471 0.474 0.471 0.420 0.436 0.435 0.427 0.471 0461 0.464 73 0.422 0.438 0.438 0.430 0.474 0.464 0.466 CHAPTER 15, Now we obtain first matrix of residual coefficient (R,) by subtracting Q, from R as shown i below: First Matrix of Residual Coefficient (R,) Variables 1 2 4 Sa -0.363 0.190 0.307 0.192 Variables ieee coefficient (R') 3 ain reflected matrix of re: filet oo Reflecting the variables 3. 4, 6 and 7, we obtain reflected mat ee on the d centroid factor (B) 0 Hex and then We San ext cof Resdal Contents) ‘and Extraction of 2nd Centroid Factor (2) Variables 0294 2558 ‘Sum ofcolumn sums (7)=20.987_* Second centroid factor B= 563 S77 ~ “These variables were reflected, 558 -.630 -S18 593 Research Met) iggy vgs as under atrix of factor Now we ean write the ma Tm Centroid Factor Poe nag 0.630 0518 0593 Example 15.2: Work out the communality and eigen values from the final results obtained in Exanpie 15.1, Also explain what they (along with the said two factors) indicate ven problem as under Solution: We work out the communality and eigen values for the g Table 15.2 Ce Mee p (6937 +(563¥ (618p+(STIP=.715 (.642F +(-539F (oy + 002) (6299 +558) (694 + (630) (e197 +(-SI8) (683) +53) 6121 2631 n 33 76) pcase G3") 10 Proportion of a3 1.00%) (43%) mal common variance (57%) Facto Eac row) v4 invarial variance attribute assess: usually ¢ Itha absolute loading i conventic results. In called “th common with posit tee taken a above Th Thus the tot: Proportion o techni : twos 8660 actors 4 ysis Fach communality in the above table represents the proportion of sovariableand is accounted forby the two factors 0 ; sintleone is accounted for by the centroid factor ariance in RO MECainseaesioe meer an e rema , ave in variable one scores is thought of as being ds Bree erate wet presented by variable one, anda portion due to ers of eat assessment of variable one (but there is no mention of these portions in the saa be = wally concentrate on common variance in factor analysis) im Jehas become customary in factor analysis literature for a loading of 0.33 to be the mi oat value t be interpreted. The portion of a variable’s variance accounted forby this aie trating is approximately 10%, This criterion, though arbitrary, is being used prrea ak aa jab jon and as such must be Kept in view when one reads and interprets the malivariatres ach wails Inourexample, factor has loading in excess of 0:33 onal variables; sucha factor is usually talled “the general factor” and is taken to represent whatever itis that all of the variables have in ‘immon, Wemight consider all the eight variables to be product of some unobserved variable (which canbenamed subjectively by the researcher considering the nature of his study). The factor name is Gesen in such a way that it conveys what itis that all variables that correlate w ith it (that “load on ie)havein common. Factor B in our example has all loadings in exe ss of 0.33, but half of them are ‘vith negative signs. Such a factor is called jor? and is taken to represent a single dimension with two poles. Each of these poles ter of variables—one pole by those ‘vith positive loadings and the other pole with nest! a We can give different names to the said wo eo rows at the bottom of the above table give us ‘actors in explai Jing the relations among the eigh | tthenas equal to the number of variables involved (0 Inthis present example, then V=8.0. The row label thenumerical value of that portion of the variance attributed to the fa a it. These are found by summing up the squared valuct of th res Thsthe otal value, 8.0, i partitioned Frio 3.490 as eigen vale for flo onl Thetoe Band he toil 612 1/ns the sum of elzen values fr these Ww 22 a Hpporton ofthe total variance, Oy arethovman the next row there WEGHTOTT ig al variance seated to these two factors, approximately Tt ed ie raining 23% of made up of portions unique 10 tA gues used to measure them. The last row shows that of the common va Sy ert ‘and the other 43% by factor B. Thus 1 con tors together “explain” the common variance, 4s interpret and name factor B. The further information about the usefulness of the two Jes, The total variance (7) in the analysis is sumption that variables are standardized) ed “Figen value” or “Common variane tor in the concemi ng factor load 631 asei can notice that ¢ to individual ¥ approximatel Juded that the 4 5.2.2 Principal Components Method observations of possibly dure to convert a set of ea ete a uncorrelated varia of original variables, anata Of principal components 18 Methanol HEE tr he valonmation is defined in such a way that the Ft Te al eonpon es "iit inthe data, and each sueceeding component 1 hashes | Bee eretar ces uneceteea ih ce preeeaig COPEL nlged to be independent under some conditions & Research \ ology PC tech very popularly used for factor analysis. PC method of facto» . x eerie f ‘i ad loadings of each factor extracted in turn, Accor: Pa maximize the sum of squared loz earned cto bility in the data than that is accounted for using rs account for the larger variability in th e ca method The aim of the principal components method is the construction out of a 2 i &t OF varany s(J=1,2 ew variables (p), called principal components which are linear ks X's (/= 1,2, ....#) of new variables (p,), called f "combine of the X X4+4,X,+.. +4, %, X, +a, X,+... +4, X, P= 4 X,+4,X,+.. +4, eX, The method is being applied mostly by using standardized variables, ie., z,=(X,— XJ The a's are called loadings and are worked out in such way that the extracted princi components satisfy wo conditions: (i) principal components are uncorrelated (orthogonal) and iva, frst prineipal component (p,) has the maximum variance, the second principal component (p) hs the next maximum variance and so on, lowing steps are usually involved in principal components method () Estimates of a,'s are obtained with which Xs are transformed into orthogonal vail ic. the principal components. A decision is also taken with regard to the question: hos Gi many of the components to retain into the analysis? (i) We then proceed with the regression of Yon these pr pal components ie Y= Sup, + Say ++ Sp Dy (m

You might also like