Contents
Acknowledgements ............................................................................................................... 2
Introduction .......................................................................................................................... 3
Pie-chart ............................................................................................................................... 4
Bar Graphs ............................................................................................................................ 5
Mean (Arithmetic Mean) ....................................................................................................... 7
Standard Deviation (SD) ........................................................................................................ 8
1. Find mean, median, mode, p25, p50, p75 from the following data: ........................................8
[Link] mean, median, mode, p25, p50, p75 from the following data: .......................................10
3. Find mean, median, mode, SD, and percentiles .................................................................... 11
4. Enter the value in SPSS and find mean, median, mode, p25, p50, p75, SD, variance, Range,
minimum, maximum from the following data ..........................................................................13
Correlation .......................................................................................................................... 15
Regression .......................................................................................................................... 15
1. Calculate Karl Pearsons correlation coefficient, coefficient of determination ......................16
2. Calculate the correlations Coefficients from the following data: ..........................................17
3. From the following data find the regression equation y on x ................................................18
4. Fit the Poisson distribution and find the expected frequencies .............................................19
5. Fit the binominal distribution to the data given below ..........................................................20
6. Fit the Poisson distribution and find expected frequency ......................................................22
7. Fit the Poisson distribution and find expected frequency ......................................................23
ANOVA ................................................................................................................................ 25
1. The yield of treatments in different plots are as shown in the following plots. Carry .........25
out analysis and create ANOVA table .......................................................................................25
2. Create ANOVA table: ............................................................................................................26
1
Acknowledgements
I would like to express my sincere gratitude to my Instructor, Rajkumar Ghimire
for their invaluable guidance, support, and encouragement throughout the
completion of this work. Their clear explanations and dedication to teaching have
greatly enhanced my understanding of the subject.
I am also deeply thankful to my college, Kabhre Multiple Campus for providing
a supportive learning environment and the necessary resources to carry out this
study successfully. The facilities and academic atmosphere have played an
important role in broadening my knowledge and skills.
Lastly, I would like to acknowledge everyone who directly or indirectly contributed
to the completion of this work.
2
Introduction
SPSS, which stands for "Statistical Package for the Social Sciences," is a powerful
software program used for statistical analysis and data management. It is widely
employed in various fields, including social sciences, psychology, business, and
healthcare, to analyze data and make informed decisions based on research
findings. Here's an introduction to SPSS and its five main features:
1. Data Management: SPSS provides robust tools for data entry, manipulation,
and management. You can easily input data, import data from various file formats,
clean and transform data, and handle missing values. This feature streamlines the
data preparation process, ensuring your data is ready for analysis.
2. Statistical Analysis: SPSS offers a wide range of statistical techniques to
explore and analyze data. These include descriptive statistics (mean, median,
mode), inferential statistics (t-tests, ANOVA, regression), non-parametric tests,
factor analysis, and more. It allows you to perform both basic and advanced
statistical analyses to uncover patterns, relationships, and insights within your
data.
3. Data Visualization: Visualization is a crucial aspect of data analysis, and
SPSS offers various tools to create meaningful graphs and charts. You can create
bar charts, histograms, scatterplots, and more to visually represent your data.
Effective visualization helps you communicate your findings and make data-
driven decisions.
4. Reporting and Output: SPSS generates comprehensive and customizable
output reports that include tables, charts, and statistical summaries. These reports
can be exported to various formats, such as PDF, Word, Excel, or HTML, making
it easy to share your results with colleagues or stakeholders. SPSS also supports
syntax, allowing you to automate and reproduce analyses.
5. Data Mining and Predictive Analytics: SPSS includes advanced features for
data mining and predictive analytics. You can use techniques like decision trees,
clustering, and logistic regression to identify patterns and make predictions based
on historical data. This is particularly valuable in fields like marketing, where
predictive modeling can inform future strategies.
In addition to these five main features, SPSS offers a user-friendly interface that
makes it accessible to individuals with varying levels of statistical expertise. It's a
versatile tool that supports both basic data analysis tasks and complex research
projects, making it a popular choice for researchers and data analysts across
different industries.
3
Pie-chart
A pie chart is a circular graph used to represent data as parts of a whole. The
circle is divided into slices, where each slice shows the proportion or percentage of
a category. It helps in easily comparing different parts of data at a glance.
Example: Construct the pie-chart of given data in SPSS.
Class v vii vii viii ix x
No. of 55 60 40 80 70 50
student
Solution:
Statistics
No_of_students
N Valid 6
Missing 0
No_of_students
Frequency Percent Valid Percent Cumulative Percent
4
Bar Graphs
A bar graph is a chart used to compare different categories of data using
rectangular bars. The length or height of each bar represents the value of each
category. It makes it easy to see differences and compare quantities visually.
Example: Construct the bar graphs of given data in SPSS.
Class v vii vii viii ix x
No. of 55 60 40 80 70 50
student
Solution:
Statistics
No_of_Students
N Valid 6
Missing 0
5
No_of_Students
Frequency Percent Valid Percent Cumulative Percent
Valid 40.00 1 16.7 16.7 16.7
50.00 1 16.7 16.7 33.3
55.00 1 16.7 16.7 50.0
60.00 1 16.7 16.7 66.7
70.00 1 16.7 16.7 83.3
80.00 1 16.7 16.7 100.0
Total 6 100.0 100.0
6
Mean (Arithmetic Mean)
The average of all values.
Formula:
∑𝑛𝑖=1 𝑥𝑖
Mean =
𝑛
Where:
• 𝑥𝑖= each value
• 𝑛= total number of values
Median
The middle value when data is arranged in ascending order.
Formula:
• If 𝑛is odd:
Median = 𝑥 𝑛+1
( )
2
• If 𝑛is even:
𝑥(𝑛2) + 𝑥(𝑛2+1)
Median =
2
Percentiles (P25, P50, P75)
Percentiles divide data into 100 equal parts after sorting.
General Percentile Formula:
𝑃𝑘 = 𝑥( 𝑘 (𝑛+1))
100
Where:
• 𝑘= percentile (e.g., 25, 50, 75)
• 𝑛= number of observations
7
Standard Deviation (SD)
Standard deviation measures how spread out the data values are around the mean. A small SD
means values are close to the mean; a large SD means they are more spread out.
1. Population Standard Deviation (σ) Formula:
Where:
• 𝑥𝑖= each value
• 𝜇= population mean
• 𝑁= total number of values 2. Sample Standard Deviation (s) Formula:
Where:
• 𝑥𝑖= each value
• 𝑥ˉ= sample mean
• 𝑛= sample size
1. Find mean, median, mode, p25, p50, p75 from the following data:
Data: 10 20 30 40 50 60 70 80 90 30 25 31
Solution:
SPSS SYNTAX:
FREQUENCIES VARIABLES=x
/NTILES=4
STATISTICS=STDDEV VARIANCE RANGE MINIMUM MAXIMUM MEAN MEDIAN
/ORDER=ANALYSIS.
Solution using SPSS:
8
Statistics
data
N Valid 12
Missing 0
Mean 44.6667
Median 35.5000
Mode 30.00
Std. Deviation 25.30660
Variance 640.424
Skewness 0.566
Std. Error of Skewness 0.637
Kurtosis -0.828
Std. Error of Kurtosis 1.232
Range 80.00
Sum 536.00
Percentiles 25 26.2500
50 35.5000
75 67.5000
9
From the above table, we found that Mean = 44.6667, Median = 35.5000, Mode = 30, p25 =
26.2500, p50 = 35.5000, p75 = 67.5000.
[Link] mean, median, mode, p25, p50, p75 from the following data:
Data: 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25,26, 27, 28, 29, 30, 31, 32, 33,
34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 50, 50, 50, 50, 51, 52, 53, 54,
55, 56, 57, 58, 59, 60, 60, 60, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77,
78, 79, 80
SPSS SYNTAX:
FREQUENCIES VARIABLES=x
/NTILES=4
STATISTICS=STDDEV VARIANCE RANGE MINIMUM MAXIMUM MEAN MEDIAN
/ORDER=ANALYSIS.
Solution using SPSS:
Statistics
Data
N Valid 79
Missing 0
Mean 45.8861
Std. Error of Mean 2.22690
Median 49.0000
Mode 50.00
Std. Deviation 19.79315
Variance 391.769
Skewness -0.120
10
Std. Error of Skewness 0.271
Kurtosis -1.062
Std. Error of Kurtosis 0.535
Range 70.00
Minimum 10.00
Maximum 80.00
Sum 3625.00
Percentiles 25 29.0000
50 49.0000
75 61.0000
From the above table, we found that Mean = 45.8861, Median = 49.0000, Mode = 50.00, p25 =
29.0000, p50 = 49.0000, p75 = 61.0000
3. Find mean, median, mode, SD, and percentiles.
weight mid value frequency
20-30 25 4
30-40 35 6
40-50 45 7
50-60 55 21
11
60-70 65 23
70-80 75 2
80-90 85 3
Solution using SPSS:
Statistics
Weight Mid Frequency
Value
N Valid 7 7 7
Missing 0 0 0
Mean 55.00 9.43
Median 55.00 6.00
Mode 25a 2a
Std. Deviation 21.602 8.772
Variance 466.667 76.952
Range 60 21
Minimum 25 2
Maximum 85 23
Percentiles 25 35.00 3.00
12
50 55.00 6.00
75 75.00 21.00
a. Multiple modes exist. The smallest value is
shown
From the above table, we found Mean = 9.43, Median = 6.00, Mode = 2a, SD = 8.772, Variance
= 76.952, Range = 21, Minimum = 2, Maximum = 23, p25 = 3, p50 = 6, p75 = 21.
4. Enter the value in SPSS and find mean, median, mode, p25, p50, p75, SD,
variance, Range, minimum, maximum from the following data.
Weight midvalue frequency
20-30 25 4
30-40 35 6
40-50 45 7
50-60 55 21
60-70 65 23
70-80 75 2
Solution using SPSS:
Statistics
13
Weight Mid Value Frequency
N Valid 6 6 6
Missing 0 0 0
Mean 50.00 10.50
Median 50.00 6.50
Mode 25a 2a
Std. Deviation 18.708 9.094
Variance 350.000 82.700
Range 50 21
Minimum 25 2
Maximum 75 23
Percentiles 25 32.50 3.50
50 50.00 6.50
75 67.50 21.50
From the above table, we found that Mean Mid value = 50, Mean Frequency = 10.50,
Median Mid value = 50, Median Frequency = 6.50, Mode Mid Value = 25a, Mode Frequency =
2a, Std. Deviation Mid Value = 17.708, Std. Deviation Frequency = 9.094, Variance Mid Value =
350, Variance Frequency = 82.7 and so on as shown in the above table.
14
Correlation
Correlation measures the strength and direction of the linear relationship between two variables.
Karl Pearson’s Correlation Coefficient (r):
Alternative (computational) formula:
𝑛∑𝑥𝑦 − (∑𝑥)(∑𝑦)
𝑟
Where:
• 𝑥𝑖, 𝑦𝑖= individual values
• 𝑥ˉ, 𝑦ˉ= means
• 𝑛= number of observations Range:
−1 ≤ 𝑟 ≤ 1
Regression
Regression shows the relationship between variables and helps predict one variable from
another.
(a) Linear Regression Equation (Line of best fit)
𝑦 = 𝑎 + 𝑏𝑥
Where:
• 𝑦= dependent variable
• 𝑥= independent variable
• 𝑎= intercept
• 𝑏= slope (regression coefficient)
(b) Slope (b) Formula:
Alternative form:
15
𝑛∑𝑥𝑦 − (∑𝑥)(∑𝑦)
𝑏= 𝑛∑𝑥2 − (∑𝑥)2
(c) Intercept (a):
𝑎 = 𝑦ˉ − 𝑏𝑥ˉ
Regression Lines:
• Regression of 𝑦on 𝑥:
𝑦 − 𝑦ˉ = 𝑏𝑦𝑥(𝑥 − 𝑥ˉ)
• Regression of 𝑥on 𝑦:
𝑥 − 𝑥ˉ = 𝑏𝑥𝑦(𝑦 − 𝑦ˉ)
1. Calculate Karl Pearsons correlation coefficient, coefficient of
determination.
12 10 26 6 10 19 23 17 13.9 3 30 16 9 6 11 10 8.4
9.5 9 11.8 8 7 20 24 21 10.7 4 12 12 12 9 8.3 9 4.7
Solution using SPSS:
Correlations
Child Nutrition
Mortality
Child Mortality Pearson 1 .614**
Correlation
Sig. (2tailed) 0.009
16
N 17 17
Nutrition Pearson .614** 1
Correlation
Sig. (2tailed) 0.009
N 17 17
**. Correlation is significant at the 0.01 level (2tailed).
From the above table, we found that Karl Pearsons correlation coefficient is .614**.
2. Calculate the correlations Coefficients from the following data:
Age 56 42 36 47 49 42 60 72 63 55
Blood 147 125 118 128 145 140 155 160 149 150
Pressure
Solution using SPSS:
Correlations
Age Blood
Pressure
Age Pearson 1 .892**
Correlation
Sig. (2-tailed) 0.001
N 10 10
17
Blood Pressure Pearson .892** 1
Correlation
Sig. (2-tailed) 0.001
N 10 10
**. Correlation is significant at the 0.01 level (2-tailed).
From the above table, we found that the correlations Coefficients is .892**.
3. From the following data find the regression equation y on x
1
x 2 3 4 5 6 7
y 6 7 5 4 3 1 2
Solution form SPSS:
ANOVAa
Model Sum of df Mean Square F Sig.
Squares
1 Regression 24.143 1 24.143 31.296 .003b
Residual 3.857 5 0.771
Total 28.000 6
18
a.
Dependent Variable: y
b.
Predictors: (Constant), x
4. Fit the Poisson distribution and find the expected frequencies.
f 0 1 2 3 4 5 6 7
x 71 112 117 57 27 11 3 1
Solution using SPSS:
x f fx p(x)
Expected Rounded
frequency Expected
NXp(x) frequency
0 71 0 0.170858423 68.17251085 68
1 112 112 0.301893165 120.4553729 120
2 117 234 0.266710536 106.4175037 106
3 57 171 0.157085393 62.67707189 63
4 27 108 0.069389331 27.68634297 28
19
5 11 55 0.024521079 9.783910623 10
6 3 18 0.007221131 2.881231226 3
7 1 7 0.001822737 0.727272154 1
total 399 705 0.999501795 398.8012163 399
Form the above table we get the following expected frequency:
x 0 1 2 3 4 5 6 7
f 71 112 117 57 27 11 3 1
Expected Frequency 68 120 106 63 28 10 3 1
5. Fit the binominal distribution to the data given below
f 0 1 2 3 4
x 28 62 46 10 4
Solution using SPSS:
x f fx p(x) Expected Rounded
frequency NXp(x) Expected frequency
0 28 0 0.197530864 29.62962963 30
1 62 62 0.395061728 59.25925926 59
20
2 46 92 0.296296296 44.44444444 44
3 10 30 0.098765432 14.81481481 15
4 4 16 0.012345679 1.851851852 2
total 150 200 1 150 150
Cases and Values:
case Symbol value
[Link] Object n 4
Mean= 1.33 (Fx sum/ f sum) np 1.333333333
[Link] Success (np/n) p 0.333333333
Prob of failuer (1-p) q 0.666666667
Total Frequency N 150
From the above table, we get the following expected frequency:
x 0 1 2 3 4
f 28 62 46 10 4
Expected Frequency 30 59 44 15 2
21
6. Fit the Poisson distribution and find expected frequency.
x 0 1 2 3 4 5 6 7
0 71 122 117 57 27 11 3 1
Solution using SPSS
x f fx p(x) Expected Rounded Expected frequency
frequency NXp(x)
0 71 0 0.170858423 68.17251085 68
1 112 112 0.301893165 120.4553729 120
2 117 234 0.266710536 106.4175037 106
3 57 171 0.157085393 62.67707189 63
4 27 108 0.069389331 27.68634297 28
5 11 55 0.024521079 9.783910623 10
6 3 18 0.007221131 2.881231226 3
7 1 7 0.001822737 0.727272154 1
total 399 705 0.999501795 398.8012163 399
From the above table, we get the following expected frequency:
22
x 0 1 2 3 4 5 6 7
f 71 112 117 57 27 11 3 1
Expected Frequency 68 120 106 63 28 10 3 1
7. Fit the Poisson distribution and find expected frequency.
MPP 0 1 2 3 4 5
NOP 142 156 69 27 5 1
Solution Using SPSS
MPP NOP fx p(x) Expected Rounded
frequency Expected
Nxp(x) frequency
0 142 0 0.367879441 147.151776 147
5
1 156 156 0.367879441 147.151776 147
5
2 69 138 0.183939721 73.5758882 74
3
3 27 81 0.06131324 24.5252960 25
8
4 5 20 0.01532831 6.13132402 6
5 1 5 0.003065662 1.22626480 1
4
23
total 400 400 0.999405815 399.762326 400
1
From the above table, we get the following expected frequency:
MPP 0 1 2 3 4 5
NOP 142 156 69 27 5 1
Rounded Expected 147 147 74 25 6 1
frequency
24
ANOVA
ANOVA (Analysis of Variance) is a statistical method used to compare the means of three or
more groups to determine whether there is a significant difference among them.
Formula of ANOVA
The F-ratio is calculated as:
F=\frac{\text{Variance Between Groups}}{\text{Variance Within Groups}} Or,
F=\frac{MS_B}{MS_W}
Where:
• (MS_B) = Mean Square Between groups
• (MS_W) = Mean Square Within groups
Decision Rule
• If calculated (F) value > table (F) value → Reject Null Hypothesis ((H_0))
• Otherwise, accept (H_0)
1. The yield of treatments in different plots are as shown in the following plots.
Carry out analysis and create ANOVA table.
t1 2537 2069 1797 2104
t2 2211 3366 2591 2544
t3 2536 2459 2827 2385 2460
t4 1401 1170 1516 2104 1077
Solution from SPSS
ANOVA
Value
df F Sig.
Sum of Mean
Squares Square
25
Between 4265689.961 3 1421896.654 11.253 0.001
Groups
Within 1768941.150 14 126352.939
Groups
Total 6034631.111 17
Solution from Excel:
Source of SS df MS F P-value F crit
Variation
Between Groups 4265689.961 3 1421896.65 11.2533722 0.0005037 3.34388867
8
Within 1768941.15 14 126352.939
Groups
Total 6034631.111 17
2. Create ANOVA table:
A B C
10 5 15
20 6 11
15 10 22
16 12 18
26
Solution From Excel:
ANOVA
Source of SS df MS F P-value F crit
Variation
Between 158.1666667 2 79.0833333 4.79292929 0.0382622 4.256494729
Groups
Within Groups 148.5 9 16.5
Total 306.6666667 11
Solution from SPSS:
ANOVA
Value
Sum of df Mean F Sig.
Squares Square
Between 158.1666667 2 79.0833333 4.79292929 0.001
Groups
Within 148.5 9 16.5
Groups
Total 306.6666667 11
27