Quantitative Study on Mathematical Thinking
Quantitative Study on Mathematical Thinking
ASSIGNMENT 2:
Mathematical perceptions
Full Name:
1. Introduction
SPSS (Statistic Package for the Social Sciences) provides a powerful statistical analysis and data
management system in a graphical environment. With the assistance of SPSS, researchers can
perform most of the regular data management and statistical analysis functionalities such as
descriptive statistics, commonly used tests, linear regression, ANOVA, etc. In this study, the data
Six scales including Generalization, Induction, Deduction, Use of symbols, Logical relations,
and Mathematical proof were used to access students’ mathematical thinking. Each scale was
measured by five test items. Each item in the mathematical thinking test was scored out of three.
Students who did not answer a question were taken as not being able to do the question and were
given a score of zero for that item. The total score for each scale was calculated by adding up the
2. Participants
The sample consisted of 576 eleventh-grade students attending 9 schools. Of these, 176 students
(30.6%) were from two schools in an urban area, 111 students (19.3%) were from two schools in
a suburban area, and 289 students (50.2%) from five schools in a rural area. There were 279
males (48.4%) and 297 females (51.6%) in the sample. Table 1 shows student numbers by school
1
School location Total
Urban Suburban Rural
1 0 0 54 54
2 0 67 0 67
3 0 0 58 58
4 0 0 53 53
School
5 0 44 0 44
Number
6 71 0 0 71
7 105 0 0 105
8 0 0 77 77
9 0 0 47 47
Total 176 111 289 576
(30.6%) (19.3%) (50.2%) (100%)
Five of the six scales of mathematical thinking (Generalization, Induction, Deduction, Use of
symbols, Logical relations) had been constructed. It was required to construct the sixth scale
score (Mathematical proof). This was done using the COMPUTE VARIABLE option within
TRANSFORM. Then “Totalproof” was inserted as the TARGET VARIABLE. From the
FUNCTIONS on the right hand side of the Compute Variable box, SUM was selected using the
up arrow. The five Mathematical proof scale items were inserted within the parentheses after
SUM, separated by commas. Then, OK was clicked to complete the compute. The new
2
Totalproof variable which was labeled Total Mathematical proof was now added to the end of
It is important to determine the number of missing data before proceeding in the development of
a new variable. However, in this circumstance, the problem of missing data was not of concern
because, as mentioned above, students who omitted a question were considered not able to
Table 3 shows the summary statistics for the Mathematical proof scale. Table 4 in Appendix 1
shows the frequency distribution for the Mathematical proof scale for all students.
Valid 576
N
Missing 0
Mean 4.84
Std. Deviation 3.66
Skewness .53
Std. Error of Skewness .10
Kurtosis -.39
Std. Error of Kurtosis .20
Minimum .00
Maximum 15.00
As can be seen in Table 4 in Appendix 1, only 4 students (0.7%) could answer all the
Mathematical proof test items. At the other extreme, 83 students (14.4%) could not answer any
test items. The mean score for the Mathematical proof scale is 4.84, and the Std. Deviation is
3
3.66 (See Table 3). The histogram of the distribution of the Mathematical proof scale scores
below shows the scores are non-normally distributed (See Figure 1). This variable has a positive
skewness of .53 (Std. Error of Skewness = .10) and kurtosis of -.39 (Std. Error of Kurtosis = .20)
Means, standard deviations and reliabilities for the six sub-scales of mathematical thinking and
for the scale of total mathematical thinking score scale are presented in Table 5.
4
Table 5 Means, standard deviations, and internal consistency reliabilities for the six sub-scales
The internal consistency reliability for each scale was calculated separately using the items in
that scale as the scale variables. For example, to calculate the Cronbach’s Alpha for
Gen1, Gen2, Gen3, Gen4, Gen5 were selected for the ITEMS box. Then click the STATISTICS,
select the SCALE IF ITEM DELETED in Descriptives for box, then click CONTINUE and OK.
It should be noted that to calculate the alpha reliability for the Total Mathematical Thinking
scale, all 30 items from the six sub-scales were selected for the ITEMS box.
As can be seen in Table 5, the scale reliability for Total Mathematical Thinking was satisfactory
with a Cronbach alpha coefficient of 0.83. Considering the reliability statistics for the six sub-
Mathematical proof), the reliability for Logical relations was low, 0.55. The reliability for the
other sub-scales ranged from 0.61 to 0.68. These reliabilities are not particularly strong and raise
5
In terms of the standard deviation for each of the six sub-scales, the Deduction scale had the
highest standard deviation of 4.30 while the Logical relations scale had the lowest standard
deviation of 2.77. This suggests that for the Deduction scale, students’ answers were widely
distributed while for Logical relations, students’ answers clustered more closely around the
mean score.
As shown in Table 5, the sub-scale of Mathematical proof had the lowest mean score (4.84)
indicating that this is students’ weakest mathematical area. Students did best at Induction (M=
8.50) and Generalization (M= 8.02), and not as well at Deduction (M= 7.32) or Use of symbols
(M= 7.45), or Logical relations (M= 7.33). However, given that the mean score for each sub-
scale would be 7.50, students were performing under the mean for four of the six scales.
5. Gender differences
To determine if there were any differences between male and female students in any of the
mathematical thinking scales or in mathematics achievement, t-tests for independent samples for
each of the six sub-scales, for the total mathematical thinking scale, and for the mathematics
Table 6 in Appendix 2 presents the results of T-tests for independent samples for the
mathematical thinking scales and the mathematics achievement test. For the variables Induction,
Deduction, Use of symbols, Mathematical proof, Total mathematical thinking, and Mathematics
Achievement, Levene’s tests of equal variance were not significant (p>0.05), so the t-value, df
6
and two-tail significance for the equal variance estimates were used to determine whether gender
differences exist.
As can be seen in Table 6 and Table 7, although the male students had a higher mean score for
Induction than the female students, there was no significant difference between male and females
students (t= 1.10, df= 558, p= .274). Likewise, there was no statistically significant gender
difference in Use of symbols (t=0.24, df= 558, p= .814). Although the female students’ score was
higher than that of the male students for Deduction, the difference was not statistically
significant (t= .04, df= 558, p= .972). However, for the Total mathematical thinking, female
students had a statistically significantly higher mean score than male students (t=1.98, df= 558,
p= .048).
Statistically significant differences were found in Mathematical proof between male and female
students. Females had a statistically significantly higher mean score than males for Mathematical
proof (t= 3.53, df= 574, p< 0.001). Similarly, female students had a higher mean score than male
students for Mathematical achievement, and the difference was statistically significant (t= 4.05,
In relation to Generalization and Logical relations, given that Levene’s tests have probabilities
less than .05, as shown in Table 6, the unequal variance estimates are interpreted. For the scale of
Generalization, although females had a higher mean score than males, the difference was not
statistically significant (t= 1.11, df= 555.78, p= .267). In contrast, there were statistically
significant gender differences found in Logical relations (t= 5.65, df= 530.95, p< .001).
7
Table 7 Gender differences in the Mathematical thinking scales and the Mathematics achievement test
6. School differences
To determine if there were differences among schools in any of the six aspects of mathematical
Table 8 in Appendix 3 shows the results of One-way ANOVA for school differences in the
Generalization:
8
As shown in Table 8, there was a statistically significant difference in Generalization among
schools (F (8,551) = 6.99, p< .001). A Scheffe post-hoc test (See Table 9) revealed that school 1
with a mean of 10.31 differed statistically significantly from schools 7, 3, and 9 with means of
6.50, 7.20, and 7.36 respectively. These differences are illustrated in Figure 2.
Table 9 Scheffe results: Mean scores for the Generalization scale by school
9
Induction
As can be seen in Table 8, there was also a significant difference in Induction among schools (F
(8, 551) = 10.13, p< .001). A Scheffe post-hoc test revealed that school 1 with a mean of 10.63
differed statistically significantly from schools 2, 7, 6 with means of 6.77, 6.94, and 7.96
respectively. Both schools 2 and 7 differed significantly from schools 4, 3, 5 with means of 9.56,
9.69, and 10.35 respectively. Table 10 and Figure 3 below show these differences graphically.
Table 10 Scheffe results: Mean scores for the Induction scale by school
10
5 43 10.35 10.35
1 52 10.63
Sig. .175 .102 .112
Means for groups in homogeneous subsets are displayed.
a. Uses Harmonic Mean Sample Size = 58.084.
b. The group sizes are unequal. The harmonic mean of the group sizes is used. Type I error levels are not guaranteed.
Use of symbols
As shown in Table 8, Use of symbols differed significantly among nine schools (F (8, 551) =
8.59, p < .001). As can be seen in Table 11, a Scheffe post-hoc test of comparisons of the nine
groups indicates that while students at school 8 had a significantly lower mean score than
students at schools 2, 7, and students at school 1 had a significantly higher mean than students at
11
Table 11 Scheffe results: Mean scores for the Use of symbols scale by school
12
Logical relations
Table 8 shows that there was also a significant difference in Logical relations among schools (F
(8, 551) = 7.54, p< .001). A Scheffe post-hoc test revealed that school 5 with a mean of 9.14
differed statistically significantly from schools 2, 7, 3 with means of 6.00, 6.50, and 7.07
respectively. Both schools 2 and 7 differed significantly from schools 1, 9 with means of 8.19,
Table 12 Scheffe results: Mean scores for the Logical relations scale by school
13
8 74 7.36 7.36 7.36
4 51 7.89 7.89 7.89
1 52 8.19 8.19
9 45 8.23 8.23
5 43 9.14
Sig. .066 .139 .065
Means for groups in homogeneous subsets are displayed.
a. Uses Harmonic Mean Sample Size = 58.084.
b. The group sizes are unequal. The harmonic mean of the group sizes is used. Type I error levels are not guaranteed.
Mathematical proof
There was a significant difference in Mathematical proof among schools (F (8, 567) = 3.17, p =
.002) (See Table 8). A Scheffe post-hoc test revealed that school 4 with a mean of 6.65 differed
significantly from schools 6 with a mean of 3.98. Table 13 and Figure 6 show these differences.
14
Table 13 Scheffe results: Mean scores for the Mathematical proof scale by school
15
Mathematical thinking overall
As shown in Table 8, Mathematical thinking overall differed significantly among the schools (F
(8, 551) = 8.55, p < .001). As can be seen in Table 14, a Scheffe post-hoc test indicates that
while students at school 7 had a significantly lower mean score of Mathematical thinking overall
than students at schools 1, 4, 5, students at schools 1, 5 had a significantly higher mean than
Table 14 Scheffe results: Mean scores for the Total Mathematical thinking scale by school
Figure 7 Mean scores for the Total Mathematical thinking scale by school
16
Mathematics achievement
schools (F (8, 534) = 6.89, p < .001). A Scheffe post-hoc test revealed that school 7 with a mean
of 22.67 differed significantly from schools 5, 6, 1 with means of 29.85, 29.93, and 30.77
17
Sig. .087 .118
Means for groups in homogeneous subsets are displayed.
a. Uses Harmonic Mean Sample Size = 56.860.
b. The group sizes are unequal. The harmonic mean of the group sizes is used. Type I error levels are not guaranteed.
Deduction
A Scheffe post-hoc test revealed that there were no significant differences in Deduction among
Table 16 Scheffe results: Mean scores for the Deduction scale by school
To look at the separate and joint effects of school location and student gender on Generalization,
a two-way ANOVA was conducted. Gender and school location were considered as independent
As shown in Table 17 and Table 18, a main effect of gender was found (F (1, 554) = 4.01, p =
.046), indicating that the Generalization mean was significantly higher for female students (M =
8.18, SD = 3.67) than male students (M = 7.85, SD = 3.30). There was also a main effect of
school location (F (2, 554) =6.96, p= .001). This effect showed that students in the suburban
area (M = 8.53, SD = 3.89) and rural area (M = 8.34, SD = 3.43) had higher means for
19
Source Type III Sum of Df Mean Square F Sig.
Squares
Corrected Model 275.40a 5 55.08 4.66 .000
Intercept 30891.34 1 30891.34 2610.87 .000
gender 47.41 1 47.41 4.01 .046
location 164.65 2 82.33 6.96 .001
gender * location 71.24 2 35.62 3.01 .050
Error 6554.84 554 11.83
Total 42862.50 560
Corrected Total 6830.24 559
a. R Squared = .040 (Adjusted R Squared = .032)
Table 18 Means for the Generalization scale by Gender and School location
As can be seen in Table 17, there was a significant interaction between the effects of gender and
school location on Generalization, F (2, 554) = 3.01, P = .05. It should be noted that the
significance of the interaction here is 0.05. As such, it may not reach statistical significance if the
usual level of significance is less than 0.05. These interaction effects are shown graphically in
Figure 9.
20
Figure 9 Mean scores for the Generalization scale by Gender and School location
The figure shows that the female students in suburban and rural areas had the higher scores than
male students in these areas, but this was not the case for students in the urban area.
Similarly, a two-way ANOVA was conducted to examine the effect of gender and school
location on Induction. Gender and school location were considered as independent variables with
21
Source Type III Sum Df Mean Square F Sig.
of Squares
Corrected Model 599.41a 5 119.88 9.16 .000
Intercept 33092.82 1 33092.82 2528.89 .000
gender 36.24 1 36.24 2.77 .097
location 425.14 2 212.57 16.24 .000
gender * location 154.27 2 77.13 5.89 .003
Error 7249.59 554 13.09
Total 48343.00 560
Corrected Total 7848.99 559
a. R Squared = .076 (Adjusted R Squared = .068)
Table 20 Means for the Induction scale by Gender and School location
As shown in Table 19, there was a main effect of school location (F (2, 554) =16.24, p<.001).
This effect showed that students in the rural area (M = 9.35, SD = 3.70) and suburban area (M =
8.18, SD = 3.64) had higher mean for Induction than those in the urban area (M = 7.35, SD =
22
3.56) (See Table 20). However, as can be seen in Table 19, the main effect of gender was non-
As can be seen in Table 19, there was a significant interaction between the effects of gender and
school location on Induction, F (2, 554) = 5.89, P = .003. These interaction effects are shown
graphically in Figure 10. The figure shows that the male students in the urban and suburban areas
performed better on the Induction test than female students in these areas. However, in the rural
Figure 10 Mean scores for the Induction scale by Gender and School location
23
8. Correlational relationships: Mathematics achievement with Generalization;
relations
As part of the assignment, I am required to describe, investigate, and report one interesting
aspect of the data. I have chosen to investigate the correlational relationships between
and Generalization and Logical relations using the Pearson’s product-moment correlation
coefficients.
24
Table 21 Correlations between the Mathematics achievement, Generalization scale and
Logical relations scale
Mathematics Generalization Logical
Achievement relations
Pearson
1 .63** .55**
Mathematics Correlation
Achievement Sig. (2-tailed) .000 .000
N 543 527 527
Pearson
.63** 1 .40**
Correlation
Generalization
Sig. (2-tailed) .000 .000
N 527 560 560
**. Correlation is significant at the 0.01 level (2-tailed).
As can be seen in Table 21, the correlation of Generalization with Mathematical achievement
was quite high (r= 0.63), and was also statistically significant (p <001). Likewise, the correlation
of Logical relations with Mathematical achievement was moderately high (r= 0.55), and was
statistically significant (p < .001). The correlation of Logical relations with Generalization was
moderately low (r= 0.40), and was also statistically significant (p < .001).
It is interesting to see that both Generalization and Logical relations were strongly correlated to
Mathematics achievement. Because none of the items in the scales were provided as part of the
assignment, it is difficult to interpret these correlations. The high correlation between the
Mathematics achievement test and the Generalization scale (r= .63) suggests one of two things:
the Mathematics achievement test contains items that require students to make generalizations,
or students who do well on Mathematics achievement tests also have the ability to draw
generalizations. Similarly, for the correlation between the Mathematics achievement test and the
25
The correlation between the Generalization and Logical relations (r= .40) may suggest that these
two measures of Generalization and Logical relations are not entirely independent measures. But
it is difficult to say more than this without access to the items that comprise the scales.
26
APPENDIX 1
27
APPENDIX 2
Table 6 T-tests for independent samples for the mathematical thinking scales and the mathematics achievement test
28
APPENDIX 3
Table 8 One- way ANOVA for school differences in the Mathematical thinking scales and the
Mathematics achievement
29