Introduction to SPSS
The abbreviation SPSS stands for Statistical Package for the Social Science. SPSS was first
released in 1968 after being developed by Norman H. Nie and C. Hadlai Hull. During a couple of
years, number of versions of SPSS has been released as SPSS 15.0.1 - November 2006; SPSS
16.0.2 - April 2008; SPSS Statistics 17.0.1 - December 2008; PASW (Predictive Analytics
Software) Statistics 17.0.3 - September 2009; PASW Statistics 18.0.1 - December 2009; PASW
Statistics 18.0.2 - April 2010; PASW Statistics 18.0.3 - September 2010; IBM SPSS Statistics
19.0 - August 2010; IBM SPSS Statistics 20.0 - August 2011; IBM SPSS Statistics 21.0 - August
2012. This package is available for both personal and mainframe (or multi-user) computers. It is
available for several operating systems such as Windows, Macintosh, LINUX and UNIX
Systems. SPSS is a Window based full-featured data analysis program that offers a variety of
applications such as statistical analysis, graphics, reporting, and data base management. It is one
of the most popular statistical packages which can perform highly complex data manipulation
and analysis with simple instruction.
SPSS (Statistical Package for Social Sciences) is a set of software programs that are combined
together in a single package. The basic application of this program is to analyse scientific data
related with the social sciences. With the help of the obtained statistical information, researchers
may draw inferences, describe the characteristics of a population or come to a conclusion
regarding their findings. SPSS first stores and organises the provided data and then compiles the
data set to produce a suitable output. SPSS is designed in such a way that it can handle a large set
of variable data formats.
The software was acquired by IBM in 2009. IBM made significant changes in the programming
of SPSS because of which the software can now be used for many types of research tasks in
various fields. Thus, the use of this software is extended to many industries and organisations
such as marketing, healthcare, education, etc.
SPSS Windows:
SPSS makes statistical analysis accessible for the casual user and convenient for the experienced
user. The data editor offers a simple and efficient spreadsheet-like facility for entering data and
browsing the working data file. To invoke SPSS in the windows environment, select the
appropriate SPSS icon. There are a number of different types of windows in SPSS:
SPSS Data Editor: When we start an SPSS session, the Data Editor window (otherwise a
Viewer window) is opened. The Data Editor displays the contents of the working data
file. There are two related windows in the data editor window: Data View displays the
data in a spreadsheet format with variable names listed for column headings, and
Variable View displays information about the variables of data set. With the Data Editor,
one can modify data values in the Data View or Spread Sheet in many ways viz. change
data values; copy, cut and paste data values; add and delete cases; add and delete
variables; change the order of variables. The data entered in the Data View can be saved
for later use. By default, the data files are saved with .sav as extension. The files can also
be saved as “Comma Separate value” *.csv, “EXCEL Spreadsheet” *.xls, “Tab
delimited” *.dat or “Fixed ASCII” [Link] the Variable View onecan change the format
of a variable, add format and variable labels, etc.
SPSS Viewer/Output: Statistical results and graphs are displayed in the Viewer window.
The Viewer window is divided into two panes. The right-hand pane contains the all the
output and the left-hand pane contains a tree-structure of the results. One can use the left-
hand pane for navigating through, editing and printing of results. The output files are
saved with .spo or *.spv extension depending upon the version available. A Viewer
window opens automatically the first time when the procedure generates output.
Pivot Table Editor: Output is displayed in pivot tables that can be modified in many ways
with this editor. One can edit text, swap data in rows and columns, create
multidimensional tables, and selectively hide and show results.
Chart Editor: The chart editor is used to edit graphs. When we double-click on figure or
graph, it will reappear in a chart editor window. One can change the colours, select
different type of fonts and sizes etc.
SPSS Syntax Editor: The Syntax Editor is used to create SPSS command syntax for using
the SPSS production facility. Usually, one will be using the point and click facilities of
SPSS, and hence, there is no need to use the Syntax Editor. For advanced features of
SPSS (for which the click facility is not available), one has to use the Syntax Editor.
More information about the Syntax Editor and using the SPSS syntax is given in the
SPSS Help Tutorials under Working with Syntax. One of the ways to get Syntax Editor,
click on Paste in the dialogue box. The syntax files are saved with .sps as extension.
Important Functions of SPSS:
Working on large quantities of data is a complex and time-consuming process, but with the use
of SPSS this task becomes much simpler. Specific statistical functions that SPSS performs are:
Descriptive Statistics:
This submenu provides techniques for summarizing data with statistics, charts, and reports. The
various sub-sub menus under this are as follows:
Frequencies provide information about the relative frequency of the occurrence of each category
of a variable. This can be used it to obtain summary statistics that describe the typical value and
the spread of the observations. Descriptives is used to calculate statistics that summarize the
values of a variable like the measures of central tendency, measures of dispersion, skewness,
kurtosis etc. Explore produces and displays summary statistics for all cases or group-wise cases.
Boxplots, stem-and leaf plots, histograms, tests of normality, robust estimates of location,
frequency tables and other statistics and plots can also be obtained. Crosstabs is used to count
the number of cases that have different combinations of values of two or more variables, and to
calculate summary statistics and tests. P-P (Proportion-proportion) plots observed cumulative
proportion is plotted against the expected cumulative proportion when the data were a sample
from a specified distribution. Q-Q (Quantile-quantile) plots quantiles of the observed values are
plotted against the quantiles of the specified distribution.
Compare Means:
This submenu provides techniques for testing; one sample, two samples and more than two
samples have been drawn from the populations having the same mean.
Means computes summary statistics for a variable when the cases are subdivided into groups
based on their values for other variables. Independent Sample t test is used to test whether two
independent samples have been drawn from two populations having the same mean. It is most
powerful parametric test for testing the equality of two populations when the sample size is small
and the distribution of population is normal. For more than two independent groups, the One-
way ANOVA option could be used. Paired Sample t test is used to compare the means of the
same subjects in two conditions or at two points in time i.e., to compare subjects who had been
matched to be similar in certain respects and then to test if two related samples come from
populations with the same mean. For example, if one is interested to test the equality of milk
yield of the animals for the two lactations. Here the observations have been taken on the same
animals at two time points and thus the observations are related. One-Way ANOVA is used to
test whether several independent groups come from populations with the same mean. To
examine which groups are significantly different from each other, multiple comparison
procedures can be used through Post Hoc Multiple Comparison option which consist of the
options like Least-significant difference, Duncan’s multiple range test, Tukey etc. The contrast
analysis can also be performed in order to compare the different groups or treatments by using
the Contrast option.
General Linear Model:
This submenu uses the general linear model procedure for analyzing the univariate and
multivariate Analysis-of-Variance models including repeated measures model. The Univariate
submenu could be used to analyze the experimental designs like CRD, RCB design, Latin square
design, Designs for factorial experiments etc. The covariance analysis can also be performed and
alternate methods for partitioning sums of squares can be selected.
Multivariate analyses analysis-of-variance and analysis-of-covariance designs when there are
two or more than two dependent variables. It is used to test hypotheses about the relationship
between a set of interrelated dependent variables and one or more factor or grouping variables.
For example, one can test whether the three varieties are having the same mean effect for grain
yield and straw yield. Repeated Measures refer to the situation in which multiple measurements
of the response variables are obtained, over several time periods on each experimental unit, such
as an animal. In it the interest of analysis is to test between-subject effects such as GROUP;
within-subject effects such as TIME and interactions between the two types of effects such as
GROUP*TIME.
Correlate:
This submenu provides measures of association for two or more variables measured at the
ratio/interval scale.
Bivariate calculates matrices of Pearson product-moment correlations, and of Kendall and
Spearman nonparametric correlations, with significance levels and optional statistics. Pearson
correlation coefficient is used when the data are measured at the interval/ ratio scale. Spearman
and Kendall correlation coefficients are nonparametric measures which are particularly useful
when the data is in ordinal scale or the distribution of the variables is not normal. Both the
Spearman and Kendall coefficients are based on assigning ranks to the variables. Partial
correlation coefficient computes partial correlation coefficients that describe the linear
relationship between two variables while controlling for the effects of one or more additional
variables. Nominal variables should not be used in the partial correlation procedure. Correlations
are measures of linear association. Two variables can be perfectly related, but if the relationship
is not linear, a correlation coefficient is not an appropriate statistic for measuring their
association. If the value of dependent variable is to be predicted from a set of independent
variables then the Linear Regression procedure is commonly used.
Regression:
This submenu provides a variety of regression techniques, including linear, logistic, nonlinear,
ordinal, probit, weighted, and two-stage least-squares regression.
Linear is used to examine the relationship between a dependent variable and a set of
independent variables. If the dependent variable is dichotomous, then the logistic regression
procedure should be used. If the dependent variable is censored, such as survival time after
surgery, use the Life Tables, Kaplan-Meier, or proportional hazards procedure. Logistic
estimates regression models in which the dependent variable is dichotomous. If the dependent
variable has more than two categories, use the Discriminant procedure to identify variables
which are useful for assigning the cases to the various groups. If the dependent variable is
continuous, use the Linear Regression procedure to predict the values of the dependent variable
from a set of independent variables. Probit performs the analysis which is used to measure the
relationship between a response proportion and the strength of a stimulus. The analysis is used to
estimate the parameters (mean) and 2 (variance) of the distribution of tolerances that is
generally based upon the probit transformation of the experimental results. In probit analysis, the
response is dichotomous eg. alive/dead, disesed/not-diseased--and several groups of subjects are
exposed to different levels of some stimulus . For each stimulus level, the data must contain
counts of the totals exposed and the totals responding. If the response variable is dichotomous
but you do not have groups of subjects with the same values for the independent variables you
should use the Logistic Regression procedure. Nonlinear estimates the parameters of nonlinear
regression models. A „nonlinear model‟ is one in which at least one of the parameters appears
nonlinearly. The commonly used non-linear models are Logistic and Gompertz models. In non-
linear model, the parameter estimates are obtained iteratively and the initial values of the
parameters are to be given in the beginning. If the non-linear function can be transformed to a
linear function, then the Linear Regression procedure should be used on the transformed data.
Non-Parametric Tests:
This submenu provides a number of nonparametric tests for one sample, or for two and more
paired or independent samples. These are Chi-Square, Binomial, Runs 1-Sample Kolmogorov-
Smirnov; 2-Independent Samples; KIndependent Samples; 2 Related Samples; K Related
Samples.
Graphs
This menu generates a number of graphs and plots. These are Bar; 3-D Bar; Line; Axis; Pie;
High-Low; Box plot; Error plot; Population pyramid; Scatter/ Dot; Histogram.
Saving Data and Output
SPSS data can be saved as a variety of file formats, including
MS Excel;
plain text (.txt or .csv);
Stata;
SAS.
The options for output are even more elaborate: charts are often copy-pasted as images in .png
format. For tables, rich text format is often used because it retains the tables’ layout, fonts and
borders.
Besides copy-pasting individual output items, all output items can be exported in one go to .pdf,
HTML, MS Word and many other file formats. A terrific strategy for writing a report is creating
an SPSS output file with nicely styled tables and chart. Then export the entire document to Word
and insert explanatory text and titles between the output items.
Advantages and Disadvantages of SPSS:
Advantages:
The advantages of using SPSS as a software package compared to other are:
• SPSS is a comprehensive statistical software.
• Many complex statistical tests are available as a built in feature.
• Interpretation of results is relatively easy.
• Easily and quickly displays data tables.
• Can be expanded.
Limitations:
• SPSS can be expensive to purchase for students.
• Usually involves added training to completely exploit all the available features.
• The graph features are not as simple as of Microsoft Excel.