0% found this document useful (0 votes)
3 views27 pages

Math Labsheet

The document provides a comprehensive overview of statistical concepts and techniques using SPSS, including data management, statistical analysis, data visualization, reporting, and predictive analytics. It covers various statistical measures such as mean, median, mode, standard deviation, percentiles, correlation, and regression, along with practical examples and SPSS syntax for analysis. Additionally, it includes sections on creating pie charts and bar graphs to visualize data effectively.

Uploaded by

avishekdahal1431
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views27 pages

Math Labsheet

The document provides a comprehensive overview of statistical concepts and techniques using SPSS, including data management, statistical analysis, data visualization, reporting, and predictive analytics. It covers various statistical measures such as mean, median, mode, standard deviation, percentiles, correlation, and regression, along with practical examples and SPSS syntax for analysis. Additionally, it includes sections on creating pie charts and bar graphs to visualize data effectively.

Uploaded by

avishekdahal1431
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Contents

Acknowledgements ............................................................................................................... 2
Introduction .......................................................................................................................... 3
Pie-chart ............................................................................................................................... 4
Bar Graphs ............................................................................................................................ 5
Mean (Arithmetic Mean) ....................................................................................................... 7
Standard Deviation (SD) ........................................................................................................ 8
1. Find mean, median, mode, p25, p50, p75 from the following data: ........................................8
[Link] mean, median, mode, p25, p50, p75 from the following data: .......................................10
3. Find mean, median, mode, SD, and percentiles .................................................................... 11
4. Enter the value in SPSS and find mean, median, mode, p25, p50, p75, SD, variance, Range,
minimum, maximum from the following data ..........................................................................13
Correlation .......................................................................................................................... 15
Regression .......................................................................................................................... 15
1. Calculate Karl Pearsons correlation coefficient, coefficient of determination ......................16
2. Calculate the correlations Coefficients from the following data: ..........................................17
3. From the following data find the regression equation y on x ................................................18
4. Fit the Poisson distribution and find the expected frequencies .............................................19
5. Fit the binominal distribution to the data given below ..........................................................20
6. Fit the Poisson distribution and find expected frequency ......................................................22
7. Fit the Poisson distribution and find expected frequency ......................................................23
ANOVA ................................................................................................................................ 25
1. The yield of treatments in different plots are as shown in the following plots. Carry .........25
out analysis and create ANOVA table .......................................................................................25
2. Create ANOVA table: ............................................................................................................26

1
Acknowledgements
I would like to express my sincere gratitude to my Instructor, Rajkumar Ghimire
for their invaluable guidance, support, and encouragement throughout the
completion of this work. Their clear explanations and dedication to teaching have
greatly enhanced my understanding of the subject.
I am also deeply thankful to my college, Kabhre Multiple Campus for providing
a supportive learning environment and the necessary resources to carry out this
study successfully. The facilities and academic atmosphere have played an
important role in broadening my knowledge and skills.
Lastly, I would like to acknowledge everyone who directly or indirectly contributed
to the completion of this work.

2
Introduction
SPSS, which stands for "Statistical Package for the Social Sciences," is a powerful
software program used for statistical analysis and data management. It is widely
employed in various fields, including social sciences, psychology, business, and
healthcare, to analyze data and make informed decisions based on research
findings. Here's an introduction to SPSS and its five main features:
1. Data Management: SPSS provides robust tools for data entry, manipulation,
and management. You can easily input data, import data from various file formats,
clean and transform data, and handle missing values. This feature streamlines the
data preparation process, ensuring your data is ready for analysis.
2. Statistical Analysis: SPSS offers a wide range of statistical techniques to
explore and analyze data. These include descriptive statistics (mean, median,
mode), inferential statistics (t-tests, ANOVA, regression), non-parametric tests,
factor analysis, and more. It allows you to perform both basic and advanced
statistical analyses to uncover patterns, relationships, and insights within your
data.
3. Data Visualization: Visualization is a crucial aspect of data analysis, and
SPSS offers various tools to create meaningful graphs and charts. You can create
bar charts, histograms, scatterplots, and more to visually represent your data.
Effective visualization helps you communicate your findings and make data-
driven decisions.
4. Reporting and Output: SPSS generates comprehensive and customizable
output reports that include tables, charts, and statistical summaries. These reports
can be exported to various formats, such as PDF, Word, Excel, or HTML, making
it easy to share your results with colleagues or stakeholders. SPSS also supports
syntax, allowing you to automate and reproduce analyses.
5. Data Mining and Predictive Analytics: SPSS includes advanced features for
data mining and predictive analytics. You can use techniques like decision trees,
clustering, and logistic regression to identify patterns and make predictions based
on historical data. This is particularly valuable in fields like marketing, where
predictive modeling can inform future strategies.
In addition to these five main features, SPSS offers a user-friendly interface that
makes it accessible to individuals with varying levels of statistical expertise. It's a
versatile tool that supports both basic data analysis tasks and complex research
projects, making it a popular choice for researchers and data analysts across
different industries.
3
Pie-chart
A pie chart is a circular graph used to represent data as parts of a whole. The
circle is divided into slices, where each slice shows the proportion or percentage of
a category. It helps in easily comparing different parts of data at a glance.
Example: Construct the pie-chart of given data in SPSS.
Class v vii vii viii ix x
No. of 55 60 40 80 70 50
student
Solution:
Statistics

No_of_students
N Valid 6

Missing 0

No_of_students

Frequency Percent Valid Percent Cumulative Percent

4
Bar Graphs
A bar graph is a chart used to compare different categories of data using
rectangular bars. The length or height of each bar represents the value of each
category. It makes it easy to see differences and compare quantities visually.
Example: Construct the bar graphs of given data in SPSS.
Class v vii vii viii ix x
No. of 55 60 40 80 70 50
student
Solution:
Statistics
No_of_Students

N Valid 6

Missing 0

5
No_of_Students

Frequency Percent Valid Percent Cumulative Percent


Valid 40.00 1 16.7 16.7 16.7

50.00 1 16.7 16.7 33.3

55.00 1 16.7 16.7 50.0

60.00 1 16.7 16.7 66.7

70.00 1 16.7 16.7 83.3

80.00 1 16.7 16.7 100.0

Total 6 100.0 100.0

6
Mean (Arithmetic Mean)
The average of all values.
Formula:

∑𝑛𝑖=1 𝑥𝑖
Mean =
𝑛

Where:
• 𝑥𝑖= each value

• 𝑛= total number of values


Median
The middle value when data is arranged in ascending order.
Formula:
• If 𝑛is odd:

Median = 𝑥 𝑛+1
( )
2

• If 𝑛is even:

𝑥(𝑛2) + 𝑥(𝑛2+1)

Median =
2

Percentiles (P25, P50, P75)


Percentiles divide data into 100 equal parts after sorting.
General Percentile Formula:

𝑃𝑘 = 𝑥( 𝑘 (𝑛+1))
100
Where:
• 𝑘= percentile (e.g., 25, 50, 75)
• 𝑛= number of observations

7
Standard Deviation (SD)
Standard deviation measures how spread out the data values are around the mean. A small SD
means values are close to the mean; a large SD means they are more spread out.
1. Population Standard Deviation (σ) Formula:

Where:
• 𝑥𝑖= each value

• 𝜇= population mean
• 𝑁= total number of values 2. Sample Standard Deviation (s) Formula:

Where:
• 𝑥𝑖= each value

• 𝑥ˉ= sample mean


• 𝑛= sample size

1. Find mean, median, mode, p25, p50, p75 from the following data:
Data: 10 20 30 40 50 60 70 80 90 30 25 31
Solution:
SPSS SYNTAX:
FREQUENCIES VARIABLES=x
/NTILES=4
STATISTICS=STDDEV VARIANCE RANGE MINIMUM MAXIMUM MEAN MEDIAN
/ORDER=ANALYSIS.
Solution using SPSS:

8
Statistics

data

N Valid 12

Missing 0

Mean 44.6667

Median 35.5000

Mode 30.00

Std. Deviation 25.30660

Variance 640.424

Skewness 0.566

Std. Error of Skewness 0.637

Kurtosis -0.828

Std. Error of Kurtosis 1.232

Range 80.00

Sum 536.00

Percentiles 25 26.2500

50 35.5000

75 67.5000

9
From the above table, we found that Mean = 44.6667, Median = 35.5000, Mode = 30, p25 =
26.2500, p50 = 35.5000, p75 = 67.5000.

[Link] mean, median, mode, p25, p50, p75 from the following data:
Data: 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25,26, 27, 28, 29, 30, 31, 32, 33,
34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 50, 50, 50, 50, 50, 51, 52, 53, 54,
55, 56, 57, 58, 59, 60, 60, 60, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77,
78, 79, 80
SPSS SYNTAX:
FREQUENCIES VARIABLES=x
/NTILES=4
STATISTICS=STDDEV VARIANCE RANGE MINIMUM MAXIMUM MEAN MEDIAN
/ORDER=ANALYSIS.
Solution using SPSS:

Statistics

Data

N Valid 79

Missing 0

Mean 45.8861

Std. Error of Mean 2.22690

Median 49.0000

Mode 50.00

Std. Deviation 19.79315

Variance 391.769

Skewness -0.120

10
Std. Error of Skewness 0.271

Kurtosis -1.062

Std. Error of Kurtosis 0.535

Range 70.00

Minimum 10.00

Maximum 80.00

Sum 3625.00

Percentiles 25 29.0000

50 49.0000

75 61.0000

From the above table, we found that Mean = 45.8861, Median = 49.0000, Mode = 50.00, p25 =
29.0000, p50 = 49.0000, p75 = 61.0000

3. Find mean, median, mode, SD, and percentiles.


weight mid value frequency

20-30 25 4

30-40 35 6

40-50 45 7

50-60 55 21

11
60-70 65 23

70-80 75 2

80-90 85 3

Solution using SPSS:

Statistics

Weight Mid Frequency


Value

N Valid 7 7 7

Missing 0 0 0

Mean 55.00 9.43

Median 55.00 6.00

Mode 25a 2a

Std. Deviation 21.602 8.772

Variance 466.667 76.952

Range 60 21

Minimum 25 2

Maximum 85 23

Percentiles 25 35.00 3.00

12
50 55.00 6.00

75 75.00 21.00

a. Multiple modes exist. The smallest value is


shown

From the above table, we found Mean = 9.43, Median = 6.00, Mode = 2a, SD = 8.772, Variance
= 76.952, Range = 21, Minimum = 2, Maximum = 23, p25 = 3, p50 = 6, p75 = 21.

4. Enter the value in SPSS and find mean, median, mode, p25, p50, p75, SD,
variance, Range, minimum, maximum from the following data.
Weight midvalue frequency

20-30 25 4

30-40 35 6

40-50 45 7

50-60 55 21

60-70 65 23

70-80 75 2

Solution using SPSS:


Statistics

13
Weight Mid Value Frequency

N Valid 6 6 6

Missing 0 0 0

Mean 50.00 10.50

Median 50.00 6.50

Mode 25a 2a

Std. Deviation 18.708 9.094

Variance 350.000 82.700

Range 50 21

Minimum 25 2

Maximum 75 23

Percentiles 25 32.50 3.50

50 50.00 6.50

75 67.50 21.50

From the above table, we found that Mean Mid value = 50, Mean Frequency = 10.50,
Median Mid value = 50, Median Frequency = 6.50, Mode Mid Value = 25a, Mode Frequency =
2a, Std. Deviation Mid Value = 17.708, Std. Deviation Frequency = 9.094, Variance Mid Value =
350, Variance Frequency = 82.7 and so on as shown in the above table.

14
Correlation
Correlation measures the strength and direction of the linear relationship between two variables.
Karl Pearson’s Correlation Coefficient (r):

Alternative (computational) formula:


𝑛∑𝑥𝑦 − (∑𝑥)(∑𝑦)
𝑟

Where:

• 𝑥𝑖, 𝑦𝑖= individual values

• 𝑥ˉ, 𝑦ˉ= means

• 𝑛= number of observations Range:


−1 ≤ 𝑟 ≤ 1

Regression
Regression shows the relationship between variables and helps predict one variable from
another.
(a) Linear Regression Equation (Line of best fit)
𝑦 = 𝑎 + 𝑏𝑥
Where:
• 𝑦= dependent variable
• 𝑥= independent variable
• 𝑎= intercept
• 𝑏= slope (regression coefficient)
(b) Slope (b) Formula:

Alternative form:
15
𝑛∑𝑥𝑦 − (∑𝑥)(∑𝑦)

𝑏= 𝑛∑𝑥2 − (∑𝑥)2

(c) Intercept (a):


𝑎 = 𝑦ˉ − 𝑏𝑥ˉ

Regression Lines:
• Regression of 𝑦on 𝑥:
𝑦 − 𝑦ˉ = 𝑏𝑦𝑥(𝑥 − 𝑥ˉ)

• Regression of 𝑥on 𝑦:
𝑥 − 𝑥ˉ = 𝑏𝑥𝑦(𝑦 − 𝑦ˉ)

1. Calculate Karl Pearsons correlation coefficient, coefficient of


determination.
12 10 26 6 10 19 23 17 13.9 3 30 16 9 6 11 10 8.4

9.5 9 11.8 8 7 20 24 21 10.7 4 12 12 12 9 8.3 9 4.7

Solution using SPSS:

Correlations

Child Nutrition
Mortality

Child Mortality Pearson 1 .614**


Correlation

Sig. (2tailed) 0.009

16
N 17 17

Nutrition Pearson .614** 1


Correlation

Sig. (2tailed) 0.009

N 17 17

**. Correlation is significant at the 0.01 level (2tailed).

From the above table, we found that Karl Pearsons correlation coefficient is .614**.

2. Calculate the correlations Coefficients from the following data:


Age 56 42 36 47 49 42 60 72 63 55
Blood 147 125 118 128 145 140 155 160 149 150
Pressure

Solution using SPSS:

Correlations

Age Blood
Pressure

Age Pearson 1 .892**


Correlation

Sig. (2-tailed) 0.001

N 10 10

17
Blood Pressure Pearson .892** 1

Correlation

Sig. (2-tailed) 0.001

N 10 10

**. Correlation is significant at the 0.01 level (2-tailed).

From the above table, we found that the correlations Coefficients is .892**.

3. From the following data find the regression equation y on x


1
x 2 3 4 5 6 7

y 6 7 5 4 3 1 2

Solution form SPSS:


ANOVAa

Model Sum of df Mean Square F Sig.


Squares

1 Regression 24.143 1 24.143 31.296 .003b

Residual 3.857 5 0.771

Total 28.000 6

18
a.
Dependent Variable: y

b.
Predictors: (Constant), x

4. Fit the Poisson distribution and find the expected frequencies.


f 0 1 2 3 4 5 6 7

x 71 112 117 57 27 11 3 1

Solution using SPSS:

x f fx p(x)
Expected Rounded
frequency Expected
NXp(x) frequency

0 71 0 0.170858423 68.17251085 68

1 112 112 0.301893165 120.4553729 120

2 117 234 0.266710536 106.4175037 106

3 57 171 0.157085393 62.67707189 63

4 27 108 0.069389331 27.68634297 28

19
5 11 55 0.024521079 9.783910623 10

6 3 18 0.007221131 2.881231226 3

7 1 7 0.001822737 0.727272154 1

total 399 705 0.999501795 398.8012163 399

Form the above table we get the following expected frequency:


x 0 1 2 3 4 5 6 7

f 71 112 117 57 27 11 3 1

Expected Frequency 68 120 106 63 28 10 3 1

5. Fit the binominal distribution to the data given below


f 0 1 2 3 4

x 28 62 46 10 4

Solution using SPSS:

x f fx p(x) Expected Rounded


frequency NXp(x) Expected frequency

0 28 0 0.197530864 29.62962963 30

1 62 62 0.395061728 59.25925926 59

20
2 46 92 0.296296296 44.44444444 44

3 10 30 0.098765432 14.81481481 15

4 4 16 0.012345679 1.851851852 2

total 150 200 1 150 150

Cases and Values:


case Symbol value

[Link] Object n 4

Mean= 1.33 (Fx sum/ f sum) np 1.333333333

[Link] Success (np/n) p 0.333333333

Prob of failuer (1-p) q 0.666666667

Total Frequency N 150

From the above table, we get the following expected frequency:


x 0 1 2 3 4

f 28 62 46 10 4

Expected Frequency 30 59 44 15 2

21
6. Fit the Poisson distribution and find expected frequency.

x 0 1 2 3 4 5 6 7
0 71 122 117 57 27 11 3 1

Solution using SPSS

x f fx p(x) Expected Rounded Expected frequency


frequency NXp(x)

0 71 0 0.170858423 68.17251085 68

1 112 112 0.301893165 120.4553729 120

2 117 234 0.266710536 106.4175037 106

3 57 171 0.157085393 62.67707189 63

4 27 108 0.069389331 27.68634297 28

5 11 55 0.024521079 9.783910623 10

6 3 18 0.007221131 2.881231226 3

7 1 7 0.001822737 0.727272154 1

total 399 705 0.999501795 398.8012163 399

From the above table, we get the following expected frequency:

22
x 0 1 2 3 4 5 6 7

f 71 112 117 57 27 11 3 1

Expected Frequency 68 120 106 63 28 10 3 1

7. Fit the Poisson distribution and find expected frequency.


MPP 0 1 2 3 4 5
NOP 142 156 69 27 5 1
Solution Using SPSS

MPP NOP fx p(x) Expected Rounded


frequency Expected
Nxp(x) frequency

0 142 0 0.367879441 147.151776 147


5

1 156 156 0.367879441 147.151776 147


5

2 69 138 0.183939721 73.5758882 74


3

3 27 81 0.06131324 24.5252960 25
8

4 5 20 0.01532831 6.13132402 6

5 1 5 0.003065662 1.22626480 1
4

23
total 400 400 0.999405815 399.762326 400
1

From the above table, we get the following expected frequency:

MPP 0 1 2 3 4 5

NOP 142 156 69 27 5 1

Rounded Expected 147 147 74 25 6 1


frequency

24
ANOVA
ANOVA (Analysis of Variance) is a statistical method used to compare the means of three or
more groups to determine whether there is a significant difference among them.

Formula of ANOVA

The F-ratio is calculated as:

F=\frac{\text{Variance Between Groups}}{\text{Variance Within Groups}} Or,

F=\frac{MS_B}{MS_W}

Where:

• (MS_B) = Mean Square Between groups

• (MS_W) = Mean Square Within groups

Decision Rule

• If calculated (F) value > table (F) value → Reject Null Hypothesis ((H_0))

• Otherwise, accept (H_0)

1. The yield of treatments in different plots are as shown in the following plots.
Carry out analysis and create ANOVA table.
t1 2537 2069 1797 2104

t2 2211 3366 2591 2544

t3 2536 2459 2827 2385 2460

t4 1401 1170 1516 2104 1077

Solution from SPSS


ANOVA

Value

df F Sig.
Sum of Mean
Squares Square

25
Between 4265689.961 3 1421896.654 11.253 0.001
Groups

Within 1768941.150 14 126352.939


Groups

Total 6034631.111 17

Solution from Excel:


Source of SS df MS F P-value F crit
Variation

Between Groups 4265689.961 3 1421896.65 11.2533722 0.0005037 3.34388867


8

Within 1768941.15 14 126352.939


Groups

Total 6034631.111 17

2. Create ANOVA table:


A B C

10 5 15

20 6 11

15 10 22

16 12 18

26
Solution From Excel:
ANOVA

Source of SS df MS F P-value F crit


Variation

Between 158.1666667 2 79.0833333 4.79292929 0.0382622 4.256494729


Groups

Within Groups 148.5 9 16.5

Total 306.6666667 11

Solution from SPSS:

ANOVA

Value

Sum of df Mean F Sig.


Squares Square

Between 158.1666667 2 79.0833333 4.79292929 0.001

Groups

Within 148.5 9 16.5


Groups

Total 306.6666667 11

27

You might also like