SPSS
Data view: where we define data
Variable view: where we define variable
Open variable: variable not to have space, rather use underscore
When questionnaire in hand, each of the question is your variable
In Data view
Response from one respondent: one whole row
Each column is your variable
Variable view:
Label area: can be used to give label to the variable, if you want you can write your question here
100 sum wala question:
Number of variables depend upon the number of attributes/characteristics in the question
Is case mein use the following code [question number_100sum_attributename] or more advisable
[attributeName_questionNumber_100sum]
All group members should have same codes
On a construct scale, one variable for each question
e.g service variety k liye every variable can be labelled as SQ_1, SQ_2 so on because in the end
Service Quality ka ek he number
it is advisable to start ordinal with 1 (likert and semantic differential k liye)
Semantic differential mein values:
Example bad --- good
Bad = 1; 2=2; 3=3; 4=4; Good = 5
22/3/2018
Download data set
Hypothesis testing:
If car in box 3
1st assertion: in box1: false
2nd assertion: in box2: true
3rd assertion: in box3: true
Since condition is that only one assertion is true, in this case two assertions are true
If car in box 2
1st assertion: false
2nd assertion: false
3rd assertion: true
Therefore, the answer is that car is in the box 2
Null and alternate hypothesis:
Null hypothesis: negates your evidence. Says there is no influence, effect, relationship..
e.g: if studying difference of productivity of employees of two plants; null hypothesis will be there is
no difference
but the researcher want the null hypothesis to be false therefore researchers will try to get an
evidence to prove it to be false. But if you don’t have enough evidence to phir null hypothesis will be
true..
“Null hypothesis accepted” wrong conclusion because it means you have not found evidence and
the null hypothesis is true.
But if you have enough evidence you say: “we fail to reject null hypothesis” rather than accept
because of above masla..
Null hypothesis for exploring differences in groups or categories
An ad agency is testing comparative impact of adverting A and B on Brand Attitude.
Ho: Impact of both A and B is the same
P-value testing
1% is more rigorous; it means you are saying that 99% result is not by chance : measure statistical
significance
In social sciences we keep it 5%
Correlation example interpretation of total optimism and total life satisfaction:
Correlation between two variables you see
1- Direction
2- Magnitude
3- Statistical significance
Incase of this example:
Correlation Co-efficient +tive (peasron)
Sig. (significance) = 0.000 < 0.05 therefore result is statistically significant
Now for magnitude:
Table of strength of correlation small .1 to .3, medium .3 to 0.5 large .5 to 0.8 (in the slides)
Type errors:
In the next class
In the dataset provided:
Op: optimism
Mast: mastery in life
Lifesat: life satisfaction
Reverse coded items:
When statement in such a way that is not in line with our basic proposition..
When measuring satisfaction you have more questions asking you are happy and one question
asking if you have depression.. now the depression wala measure mein answers on the scale will be
opposite to the happiness waley question.. if we take it as it is it will dampen our happiness stats
thus we have..
How do you know you have a reverse coded items?
Usually authors advice about the reverse coded ones..
Analyse -> dimension reduction -> factor
A box aa gya
Saarey opt load kero > click roation -> promax and Loading plots -> continue and okey.. table mein jo
opt waley with negative aa ry thy are reverse coded
Now checking reliability:
Analyse – scale- Reliability Analysis - selct opt variables – statistics – check scale if item deleted -okey
Cronbach’s alpha pata chalye ga
Cronbach Alpha ki value kia honi chaye for your construct to be reliable: should be greater than
+0.70
Here it is negative and .243 therefore it is not reliable and this might be because of the reverse
coding jo haemin kerni chaye thi or nae ki..
How to get back to a reliable condition:
We call it reverse coding through SPSS
Here in the data view we observe that the data scale is from 1-5 (apny waley mein to pata hoga scale
ka), for reverse coding we’ll have to swap the scale.. jo 1-5 hua usko 5-1 kerna hoga
Disclaimer: hmesha apny original data set ko save ker k baad mein copies py analysis karen. Even
reverse coded wali copy bhi alag sy save ekro, as you proceed into the steps save your files on each
step
For recoding:
Transform – recode into same variable, click on the reverse coded wali values take them to the box,
click on old and new value wala button
Check ager recode hui hy in the data set
Ab same analysis to check k reliability sae hui ya nae
26/03/18
Today we’ll do factor and intial analysis; not more
KMO measure of sample adequacy: tells adequacy of sample for factor analysis not the adequacy
as a whole.
Factor Rotation:
We did rotation last time: cromax rotation.. we’ll do it again today.
In descriptives.
Mainly it was related to rotation-> when a lot of data validity ko visualise kerna impossible, rotation
determines which items belong to where.. example: if people with certain nationality and we ask
them to specify that on form, when we run analysis all nationalities will appear together i.e.
Pakistanis in one group brits in other etc. but in certain cases where dual nationalities then we have
a confusion, then we ask factor analysis k isko kidher daalna hy us case mein factor analysis will put a
person either in pak or brit (brit-paki) based on statisitcs.. (this Is just example, spss mein inter-
correlation hogi)
KMO adquecy: if less than 0.5 factor analysis not useful
Bartlett’s test sphericity:
Should be less than 0.05 of the significance level
Factor analysis looks for clusters of related items, but wha happens when no relatin in the samples?
Tab use hota Barlett’s for qualifying test of sample..
For checking how many dimension we have?
Jin factors ki eigenvalue more than 1 hogi wo factors/ related questions lengy.. the questions jinki
eigen values bohot barri hongi un mein correlation ya relevancy zaada hogi..
The more the eigen value the more is the explanation
Two types of rotations:
Oblique and orthogonal
When correlation b/w two questions is greater than 0.90 is ka matlb dono questions same baat
pooch ry hain.. how ever, same construct k liye questions mein inter-correlation 0.70 sy zzada honi
chaye BUT ager 0.90 sy zaada hogi to wo overlap ho ry hongy
Now we’ll do factor analysis:
Analyse-dimension reduction – factor
Select mast1-mast7
Click on descriptives -> check KMO and Blabla test
Click on extraction -> scree plot
Click on rotation-> promac and click on plot wala option..
Score -> save as variable abhi click kero but baad mein uncheck hua wa hona chaye
Don’t compute regression variable (score) before dimension clearance..
Regression variables were added as FAC1_1 and Fac 2 at the end of data set..
Therefore jab dimension reduction ho jayegi tab score calculate karengy abhi nahi ki v thi, test sy 2
factors mil ry thy is liye regression variables bhi 2 calcualte huye wy..
Example, in this example hum ny last time mast2 ko remove ker diya tha for reducing two dimension
to one dimension.. abhi bhi ager analyse-bla bla mein mast2 nikal k rotate kero to sirf ek factor ka
cluster baney ga in the eigen value..
Kia kaise tha? Reverse coding ker k dekha tha..
Regression variable of one factor in the data view, why you need this? Cox baqi sab ordinal hain and
we need a scale quantity.
Scree plot ager elbow (L shape) hota to that would have been perfect meaning ek he variable saari
info cover ker rah y, ager shape change ho re ho (as in this one) tab matlb baqi variables bhi variance
dy ry hain..
Abhi we calculated regression variables
Now we’ll cacluate standardize variable
Sum tab kerty jab reliability condition or dimension condition satisfy ho jaati hy.. reliability last time
ki thi
For reliability
Analyse-scale-statistics- scale if deleted
Item total Statistics Table mein check the chrobach-aplha value for each variable, is case mein ager
kisi ek variable ka chronbach alpha if deleted is greater than chronbach-alpha overall to phir usko
rmove kerdo..
And don’t add it in the sum too..
For sum, mean reg:
Look for reverse coding
Look for dimensionality – single dimension honi chaye
Look for reliability –
For z-score
Analyse-descriptive-descriptive – check the sum variable and check the box at the bottom which says
standard something.
z-Lifestatisfaction sum and lifesatisfaction reg k darmiyan correlation check kero:
analyse-correlate-bivariate. HIGH CORRELARTION
write atleast four nominal variables you have asked
gender
age
spice level
occupation
frequency of visit
do you want to change the taste or not
Prefer garam masala or not.
For next class:
categorise all your variables in three categories (scale, nominal, ordinal) and make a statistical map
in your mind k konsa variable kis k saath mil k extra info dy sakta hy
29/3/2018
Get responses in numbers rather than strings before importing it to SPSS
Chi Square:
When do we do Chi square test?
Two factors make you merky in terms of your understading during analysis:
How variables are related? What analysis to bhi used?
What type of variables (nominal, ordinal, scale) we need to have to apply Chi square?
Is mostly used for nominal variables.
**Scale mein sab sy zada power(explanatory) in scale and us sy kam in ordinal and least in nominal
jiki wajja sy nominal ko ignore kerty log**
Used when nominal has two categories(males and females, online/offline) but can be used for more
two.
Median Split:
Small but useful, turns scale into nominal variables
Graph - > legacy dialogue -> 3d
Double click on graph, click on second last icon to invert transpose
Cick on Y: graph properties
Click in X: variable properties
Graph - > legacy dialogue -> line graph
Otherstatistic - > add a variable -> click no of means
Add another variable on x-axis
Nominal k liye no of cases use karengy,
Scale variable k liye mean use karengy
Now using scale variable
Chi Square:
Two categorical variables or two nominal variables
Tells the Statistical significance of the relationship
Phi and Cramers V test tells about the strength of the relationship between nominal variables
Analyse- descriptive – crosstabs
Add sex and source of stress in x and y placeholders
Click continue, then check “Display bar chart” option before clicking Ok
A beautifull graph and result will appear..
Chi value <0.05 means significant
Phi Cramers: closer to 1 means stronger
Now we will add layer
Layer up kerny sy individual Yes and No ki significance kam ho gae
To add even more layer or next layer go to ooper wali screen and click on next and then add child
Median Split:
We sum a scale variable, calculate median and then split the variable to make it nominal
Working:
In variable view a new row: create variable with the name “Stress”
Copy tpercieved stress and paste it in the stress column
02/04/2018
Test on 12/04/2018
Chi square; best for two variables where both variables have two categories
Median-Split: enables to convert the scale variables into two category nominal variable
We will do the median split now:
[Link] wala data set
Pehly check median.
Analyse – descriptive ..- frequency – select total perceivedstress – statistics – check median –
continue
Now we’ll transform in compute variables
Trasform – compute variable – type vaiable name (any of your choice in the target Variable) – write
1 in the formula – click if button at the bottom –
Same procedure for coding tspdtress >26 as 2
Based on your own data how can you use this type of analysis
In our ones:
Each construct k attributes ko accumulate kerna hy and then median split..
Example: snsory appeal
Opern personality color data set:
Total responses are 470 by adding up all freqs
Why are we weigthing cases? Because in one cell multiple frequencies are added isi liye
liye we are weighting on frequencies : data- weight – weight by freq
If weight not on:
Moving on Pair test, pair t-test..
Parametric vs non-parametric test
Parametric tests are calculated on the basis of the mean scores ici liye normal distribution is a pre-
condition for parametric test.
On the other hand: non-parametric based on frequencies.
Paired or non-paired comparisons repeated measures.
Difference among above:
If you want to check life sat between two groups and these groups are different. Gave same surveys
to both the responses we get will be call independent sample tests since both were separate groups.
This is called independent samples
Paired samples:
Whole class is one group; calculated lifesat of whole class will have got one value:
Then show them a documentary about people living under poverty and then the same lifesat scale
given to this class and then measure life sat. now we calculated two variables pre and post.. this is
called paired because same respondent but changed variable..
Hundred sum question is an example of paired test.
Non-Parametric to repeated measure ANOVA
Response 3: grading
Pulpi:
05/04/2018
Leader to submit:
Word file with title, brief summary, conceptual model, scenario based willadd a methodology basis,
model walon ny sirf model copy paste kerna, that will be the word doc, then padf of the
questionnaire, and then submit the data set, compile/combine send to shuja in Folder, folder ka
naam is the name of all the group mates
That has to be handed over to Shuja by today
Quiz based on electronic, you’ll get data set and instructions, follow instructions and submit three
files, one of the data set you get updated one, output file and the word file with interpretation of
the results. Some steps if you don’t do correctly, the rest is effected too.
From sir ki slides:
Legacy Dialogue: analyse-> nonparametric test -> legacy dialogue -> K related Samples
A friedman test was carried out to compare the total understanding scores
for….
New: Analyse -> nonparemtric tests -> related samples
Did New: Analyse -> nonparemtric tests -> related samples -> field tab -> choose the 100 sum waley
variables
Output:
Double click on the box and you’ll get a detailed analysis now
On the bottom screen click on view-> pairwise comparison
Hamesha statistical significance is not our priority,
Purity and trusted brand ki statistical difference nae aa ra..
For managerial purposes, we’ll see proximity of the factors.. pulpiness and trusted brand are close
by because we have accepted null hypothesis.
Purity and trusted brand ka significance level was 0.916 which means that we fail to reject null
hypothesis, therefore while making the anlaysis we need to see how close together are two samples
and the hypothesis we are testing..
Abhi tk we were doing nominal p test.
Scale py we did dimension reduction then reliabiltity analysis and then sum pehly in previous classes.
Research Design:
Scientific vs non scientific description:
In scientific experiemnts the result are probablisitic, they are not absolute thus should be described
in the same way.. instead of getting overwhelmed by the data.
t-test
Two types of tests: independent t-test and paired sample t-test
Independent t-test: we have two conditions which should be there in order to run the test
Condition 1: we need to have 1 scale variable and one nominal variable with two conditions,
independtant is liye kehty kiyun k two unrelated groups. Eg. Measuring brand loyalty among male
and female, brand loyalty is scale and male/female Is nominal two conditions thus is case mein we
will use independent t-test.
Paired test: respondent should be the same.. same observation from same person..
Pre-and-post test when test between two observations. Can be such that the same person is
evaluating two samples..
What if in the nominal you have more than two nominal categories e.g: brand loyalty among three
age groups.
You cannot run independent t-test when you have more than two categories in the nominal variable
In that case you run one-way ANOVA for one scale and more than two categories wala nominal
variables
Paired mein if you have more than two observation on the same sample you cannot use paired t-test
rather you’ll use “one way ANOVA repeated measure”
If we show an ad. To our respondent for a certain period attitude toward a brand may change..
Attitude towards a brand from respondents taken at T1 and then show audience the Ad. And then
ask same questions you asked earlier at T2 – paired t-test used
Now if you want to check the attitude towards brands after three weeks of watching the add so now
you have three observations (2 pichi wali ek ye wali) thus one way ANOVA repeated measure used.
Paired t-test: Significance level tells if there is a difference between two samples
Phir to check k ager effect hua ktna us k liye we get Effect size by
Ets Squared = t^2/(t^2+N-1) t is the value from paired t-test results
In the data set
Analyse > compare means > paired t test
Okay and you’ll get t values and N eta khud calculate ker lo repeat for dep 3 and dep2 eta shows k
dep kam ho ra
Now for independent t-test
[Link] wala dataset open
In this test we get two different significance score..
When significance of Levene is > 0.05 we take first row i.e ‘Equal variance assumed’ and thus usi
row wala t score and significance level
Otherwise we’ll take second row wali values
Eta square -= t^2 / (t^2 + N1+N2-2)
Define groups as defined in the data set;
09/04/2018
One way annova
Post Hoc analysis sy pata chalta hy following:
When pairwaise comparison of significance of more than two categories; we may find that pariwaise
comparison mein two particular pairs have grater significance than the overall categories
Ets-squared = sum of square between-groups/total sum of square
Scale variable: total optimism
Grouping variable: age with 3 cats..
Analyse – compare means – one way annove
Post hoc -> turkey
Options -> means plot
Annova wala table: over all significance
Ad hoc: shows pairwise
Which group has higher optimism?
Either graph and from the table mean
Coorelation :
If more than two dependent variables Regression is not recommendable
First we’ll run correlation:
Doesn’t tell causality waghera only tells if variables are related or not:
Two scale or continuous variables can run correlation
If you want to control a third variable to see whether it affects the other two ka relation; that is
called partial correlation: all three vairabes must be scale e.g: seeing correlation between total stress
and total life satisfaction while keeping optimism ka control
Delta R = change in correlation co-efficient; how do we know? When no control applied the
correlation was -.58 now it tells us when social desirability is control the co-efficient becomes -0.55
this reduced but no substantial difference
Thus delta R will be the difference between the two correlations
Three variables:
Tpercieved stress, t lifestatisfaction , total optimism
Analyse -> correlate -> bivariate
See the betufiul result
Analyse > corelate > partial >
Analyse > regression > linear
Output:
Model summary:
How much the dependant variable is explained by independent variable:
Rsquare: .326 means 32.6% is explained by these two independent variables, (it is a stand alone
condition without the significance ki interpretation)
Model fit: tells you k or variable chaye explanantion k liye nan ae, and over all significance of the
model
Higher Rsqaure, the more the observed data is explained by the straight line
Model ki significance dekhny k liye annova waley table ka last column, if value is less than 0.05 that
means result is statistically significant
Third table tells the regreassion line, whether the predictors are significant or not.. basically
equation mein variable x1 and x2 k cofficients:
Y = 0.497x1-0.394x2+21.905
Both variables are significant.
Which one is better predictor: total perceived stress cox higher abs. beta
Choose independent/dependent of your choice with highest R square value
Apney model py regression kin py chal sakti?
Perceived food quality + perceived service quality = customer satisfaction
Sensory appeal of Student biriyani + perceived servive quality + perceived value for money =
customer satisfaction
File split:
Data mein bhi bottom right py “split by “ aa jayega
Dta will remain splitted until you reset it
For different block in regression:
Click on next
And add another block
Regression models for all the layers added will be shown
For quiz:
Reverse coding
Dimension reduction
Reliability
Graphical condition
Chi-sqaure
Median split
Based on median split Chi-sqaure
100 sum analysis
Friedman test
Pair t-test
Independent test
One way ANOVA
Corelation
Regression
Split file
Block regression (2,3 different models of regression)