Otto 2020
Otto 2020
To cite this article: Siegmar Otto, Franziska Körner, Beatrice A. Marschke, Martin J. Merten,
Steffen Brandt, Sofoklis Sotiriou & Franz X. Bogner (2020): Deeper learning as integrated
knowledge and fascination for Science, International Journal of Science Education, DOI:
10.1080/09500693.2020.1730476
Introduction
Assessments in STEM (science, technology, engineering, and mathematics) education
have been under discussion for decades. The assessment of high-quality knowledge
content has been undertaken with tools such as the knowledge questionnaires used by
PISA (OECD, 2012) or TIMSS (Martin, Mullis, Foy, & Hooper, 2016). However, in
order to foster interest in and a fascination with science to motivate more students to
acquire a mastery of STEM and to eventually follow a related career, more than high-
quality content has to be promoted and assessed. An educational concept that goes
beyond memorising facts and simple information is deeper learning (in contrast to simple
factual learning). However, the effectiveness of deeper learning approaches has not been
tested empirically, probably due to a lack of assessment tools that are able to measure
the essence of deeper learning.
Deeper learning is conceptualised in different ways in the literature, and a com-
monly acknowledged definition does not yet exist. For instance, the U.S. National
Research Council concluded that twenty-first century skills or competencies cannot
be defined precisely because not enough research is available to support such
definitions (Pellegrino & Hilton, 2013). Some researchers conceptualise deeper learn-
ing as a deeper understanding of a specific matter (i.e. computer science) that involves
not just a simple understanding of a topic but also the ability to combine existing
knowledge and to develop a more mature understanding of the matter (Grover, Pea,
& Cooper, 2015). Nevertheless, a widely shared conceptualisation of deeper learning
as ‘skills for the twenty-first century’ seems to rely on three broad domains of compe-
tence, which are the cognitive, intrapersonal, and interpersonal domains (e.g. Dede,
2014; Hilton, 2008; Pellegrino & Hilton, 2013). In this line of conceptualisation,
teaching for deeper learning should (a) use multiple and varied representations of con-
cepts and tasks, (b) encourage elaboration, questioning, and explanation, (c) engage
learners in challenging tasks, (d) teach with examples and cases, (d) prime student
motivation, and (e) use formative assessments (Pellegrino & Hilton, 2013). Concep-
tually, such teaching addresses knowledge that is integrated in terms of knowing,
applying, and reasoning (i.e. intellectual abilities) as well as a motivational component
(i.e. motivational ability). We refer to this motivational ability as fascination because
this reflects a teacher’s task in the classroom with respect to children, that is, to
ignite their fascination with a subject.
education through an increasing fascination with science. The project included and used
innovative and meaningful digital technologies (e.g. advanced interfaces, learning ana-
lytics, visualisation dashboards, and Augmented/Virtual reality applications) to build a
storytelling platform where students developed and presented stories about a Mars
mission. Thus, our assessment tools were developed for sixth-grade students located
throughout our project’s regional scope (i.e. Europe) with the aim of covering a broad
range of individual knowledge and fascination across a diverse set of schools participating
in our project (i.e. private schools, public schools, and schools for socially disadvantaged
students).
Integrated knowledge
To assess the intellectual ability of the deeper learning paradigm in STEM (i.e. integrated
knowledge), we needed to develop a fixed set of items for sixth graders with a suitable wide
range of item difficulties related to the scientific content taught in sixth grade. Whereas
physical science and earth science are common elements of science education, life
science is too. However, many school curricula begin to differentiate life science into
biology and chemistry in the sixth or seventh grades. In order to avoid problems with
different curricula that introduce biology or chemistry earlier or later over the course of
sixth or seventh grade, we decided to stick with the domain of life science.
A suitable basis for our item set is the Trends in International Mathematics and Science
Study (TIMSS; Martin et al., 2016). TIMMS assesses integrated knowledge by differentiat-
ing between knowledge, its application, and reasoning, and it covers the content domains
earth science, physical science, and life science. TIMSS is an international comparison
study that has been assessing the STEM knowledge of fourth and eighth graders every 4
years since 1995 with an extensive set of nearly 800 items. The development of TIMSS
is continuous, and the developers publish sets of items and replace them with new
items in every other study. For the published items, extensive data on the difficulty
across countries are available online at [Link]. The questionnaires’ tasks
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 5
are categorised into the three cognitive domains Knowing, Applying, and Reasoning. The
first domain, Knowing, covers the facts, concepts, and procedures students need to know,
whereas the second, Applying, focuses on students’ ability to apply knowledge and con-
ceptual understanding to solve problems or answer questions. The third domain, Reason-
ing, goes beyond the solution of routine problems to encompass unfamiliar situations,
complex contexts, and multistep problems (Martin et al., 2016).
However, the TIMSS items were developed for fourth and eighth graders, and they are a
large collection of items that are not ready to use as predefined scales because these items
have to be compiled into suitable scales for an assessment by experts on the basis of their
properties (e.g. item difficulties). Thus, the TIMSS items are not ready to use as a scale, nor
were they developed for sixth graders. We used the available data on item difficulties to
identify the TIMSS items that were rather difficult for fourth graders but relatively easy
for eighth graders because our target population consisted of sixth graders who have a
knowledge level that lies between fourth and eighth graders. Even though TIMSS
focuses on science and mathematics, the TIMSS science items include many practical
issues that are related to technology and engineering (e.g. Items 33A.P, 30A.P, 9 K.P),
and thus, we decided that the TIMSS science items would be a good basis for assessing
the core of science, technology, and engineering in the sixth grade.
In line with our goal, we selected items from the content domains earth science, phys-
ical science, and life science, which include not just science but also elements of engineer-
ing and technology (e.g. Items 33A.P, 30A.P, 9 K.P). We adjusted these items so that they
would be appropriate for assessing sixth graders. In TIMMS assessments, items are
unevenly distributed across the three content domains to match the content of the
science that is taught in each grade. Because our goal was to test students as economically
as possible across the three content domains and to be able to compare the results, we
chose an even distribution (i.e. 12 items with a suitable distribution of difficulty levels
in each content domain).
We chose 36 items, four for each combination of the three cognitive domains with the
three content domains. Some items were adapted to yield a multiple-choice format with
four response options. Twenty-four of the 36 items had one correct answer, eight ques-
tions had two, and four had three correct answers. We used a partial-credit model with
the students’ response patterns to develop a scoring schema for the individual items.
The patterns were defined by the number of correct options chosen or the number of
incorrect options not chosen. For instance, Item 40A.L (Applying/Life Science) had
three correct (A, B, and C) and one false (D) response option. Thus, if a student chose
A, B, and C but not D, the answer pattern was 1, 1, 1, 1—with four correct responses.
Three correct responses resulted when a student chose A, B, not C, and not D (1, 1, 0,
1) but also if a student choose A, not B, C, and not D (1, 0, 1, 1) and so forth.
The point-biserial correlations of the response scores of one item with the students’
abilities should be strictly ordered, that is, higher scores should be more highly correlated
with students’ ability. Thus, using the point-biserial correlations between the response
scores with participants’ overall ability score, we merged the items’ response scores. If
two response scores (e.g. two and three correct responses) were nearly equally correlated
with the overall ability score (a difference smaller than .1), they were merged into the same
scoring category. In addition, if a response score had a lower correlation with students’
overall ability than a lower response score on the same item, we merged these response
6 S. OTTO ET AL.
scores into one category. All items were translated into the national languages of the
partner countries involved (Greek, German, French, Portuguese, Finnish) via parallel-
or back-translation (Brislin, 1980; Hambleton, 1996) by us and our project partners.
For an overview of all 36 knowledge items, see Appendix A.
items that were evenly distributed across these seven parts, resulting in a final set of 21 cells
(the three components times the seven parts of science). To validate the items as relevant
to the topic, we asked four experts (i.e. trained teachers from Germany, Greece, France,
and Portugal) to rate how much fascination with science they thought a child must
have to agree with the proposed statements and behaviours, and we then reduced the
item pool to four items per cell in such a way that a broad spectrum of fascination was
covered with respect to topics and difficulty.
The resulting 84 items were administered using a 5-point Likert scale with strongly dis-
agree, disagree, partially agree, agree, and strongly agree as response categories for the
affective and cognitive component, and never, seldom, sometimes, often, and very often
for the behavioural component. The scales were translated into Greek, German, French,
Portuguese, and Finnish via parallel translation by us and our project partners.
To test whether the wording and the general design of the scale were suitable for sixth
graders, we conducted a prestudy with 40 Greek sixth graders.
Analysis
Our assessment instruments were based on Item Response Theory (IRT), including
measures with only dichotomous items based on the Rasch model (Rasch, 1980) and com-
petencies measured with items that included more than two answer categories based on
the partial-credit model (Masters, 1982). The Rasch model assumes that the probability
of a certain response to an item depends on the item difficulty (item parameter) and
the underlying latent trait of the person who is answering (person parameter). Assuming
that a test or scale includes only dichotomous items, the log-odds form of the equation that
describes this characteristic is given by
pki1
log = uk − di , (1)
pki0
where pki1 is the probability that person k will provide a correct answer to item i, pki0 is the
corresponding probability of an incorrect answer, θk is the ability of person k, and δi is the
difficulty of item i. The partial-credit model, which was used to evaluate answers to items
with more than two answer categories (e.g. those used in Likert items), is an extension of
the Rasch model. The log-odds form of the corresponding equation is given by
pkij
log = uk − dij , (2)
pkij−1
where pkij is person k’s probability of answering with category number j of item i, pkij−1 is
person k’s probability of answering with category number j-1 of item i, θk is person k’s
ability, and δij is the difficulty of category j of item i.
Due to this mathematical formalisation, people can be discriminated with respect to
their abilities and items by how demanding they are to agree with (Brügger, Kaiser, &
Roczen, 2011). For example, if a person’s ability equals the difficulty of a dichotomous
item, the person’s probability of answering this item correctly (or positively) is .5.
Further, the higher the positive discrepancy between the person’s ability and the item’s
difficulty, the more likely the person will be to answer the item correctly.
8 S. OTTO ET AL.
A feature of the Rasch model and its extension, the partial-credit model, is their
specific objectivity, which means that the resulting estimate of a person’s ability is inde-
pendent of the questions that are asked, and the estimate of a question’s difficulty is inde-
pendent of the individuals who are surveyed. Unlike in classical test theory, scales based
on IRT can also differentiate the low and high ends of the ability range if items repre-
senting a broad range of difficulties are used (Bond & Fox, 2007), which makes them
especially suitable for measuring students with unknown ability distributions such as
in the STORIES project or other projects. To confirm that the instruments we developed
complied with the assumptions of the Rasch model, we evaluated how well the models fit
the items and persons. We estimated the fit of the models with the R package TAM using
the Marginal Maximum Likelihood method with Quasi-Monte-Carlo Integration
(Robitzsch, Kiefer, & Wu, 2018).
Sample
Overall, N = 1,261 students took the fascination with science questionnaire, and N = 1,246
students took the knowledge test (see Table 1). To calibrate the measurement instruments,
we used answers from N = 711 and N = 688 students, respectively. The answer data were
selected on the basis of the following strict criteria:
(a) data from students who answered more than 50% of each questionnaire,
(b) data from students who took a plausible amount of time to answer the fascination
with science scale (more than 6 min),
(c) data from students who did not provide more than two contradictory answers on the
collaborative problem-solving scale (i.e. they did not give the same response to the
negatively and positively worded questions).
The different sample sizes for the two scales were due to dropout and/or loss of motiv-
ation to answer the questionnaires attentively during the assessment.
Results
Integrated knowledge
For the analysis of the knowledge test, we used data from N = 688 students. Some of the 36
items (for an overview, see Appendix A) had more than one correct answer. We analysed
the point-biserial correlation coefficients of the response numbers (i.e. number of correct
choices) with the students’ estimated overall person abilities as described above. On the
basis of these analyses, we developed a specific scoring scheme for the items: All items
with only one correct answer were scored 1 when the correct answer was chosen and
no incorrect answer was chosen. Of the eight items with two correct answers, we scored
two items 1 when three correct choices were made and 2 when all four choices were
correct; two items were scored 1 when all four choices were correct; and four items
were scored 1 when three or four choices were correct. Of the four items with three
correct answers, two items were scored 2 when at least three correct choices were made
and 1 when two correct choices were made (and 0 otherwise); and two items were
scored 1 when all four choices were correct (and 0 otherwise). Thus, we estimated four
items with a two-step partial-credit Rasch model and the remaining 32 items with one
step.
These rescored data were then calibrated using the partial-credit Rasch model. Three
items (one reasoning-earth, one knowing-life, and one applying-life) exceeded reasonable
MS values < 1.11 calculated in accordance with Wu and Adams (2013). Overall, the item
fit ranged from 0.83–1.14 with a mean MS value of 1.00 and a standard deviation of 0.07
(see Appendix C).
The mean of the Guttman errors on the knowledge test was 11%. Only one person’s
Guttman error was above 40%.
The final calibration of the test data resulted in the following test characteristics: The sep-
aration reliability WLE was r = .85. Item difficulties ranged from −2.35 logits to 2.80 logits
around M = −0.02 with SD = 1.22 and covered almost the full range of observed student
abilities for the area of science knowledge. Figure 1 shows the specific difficulties per ques-
tion and the distribution of the students’ abilities. The mean of the students’ abilities was
fixed to zero for scale identification, and the observed standard deviation was 1.03.
The correlations between the cognitive and content domains are shown in Tables 2 and 3.
1.30 (Bond & Fox, 2007). For the scale with 72 items, the separation reliability was also very
high (r = .94). Item difficulties ranged from δ = −1.99 to δ = 2.44 around M = −0.07 with
SD = 1.10. Figure 2 shows the specific difficulties per question. The average Guttman
error was 24% with 73 students (10%) exceeding the cut-off of 40%. For an overview of
all items with difficulties and fit values, see Appendix B.
As expected, the behavioural items were more difficult than the affective and cognitive
items (M = 1.23 vs. M = −0.53/−0.91). The item difficulties in general were clustered
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 11
Figure 2. Item Person Map for the Fascination for Science Scale.
slightly differently than students’ abilities (see Figure 2). Abilities were distributed
approximately normally with M = 0.03 (SD = 1.51) as shown in Figure 2. Five students
reached the maximum score, and two students had scores of 0.
Table 4 shows that the correlational analysis between the assessed affective, behavioural,
and cognitive fascination components indicated a close relationship between the attitude
domains. Because of the length, the very high reliability, and a few items that did not meet
Wu and Adams (2013) strict fit criteria, we developed a shorter version with 36 items for
further use. We reduced the number of items evenly across the two dimensions (cognitive
and affective) from four to two by keeping the items with the best fit and a constant per-
formance across countries (low differential item functioning). The calibration of the
resulting scale with the same sample also had a high separation reliability of .90. The
item difficulties ranged from δ = −2.08 to δ = 2.49 around M = −0.05 with SD = 1.12.
The distribution of the abilities looked very similar to the 72-item version reported
above (M = 0.01, SD = 1.64). The fit of the items ranged from MS = 0.87 to MS = 1.16
with a mean MS of 0.99 and a standard deviation of 0.07.
Discussion
The deeper learning paradigm incorporates the idea that a range of competences and their
skilful application lead to a sustainable mastery of STEM. In order to empirically assess the
overall success of approaches based on the deeper learning paradigm, we monitored the
consequences of the two core abilities addressed by deeper learning: integrated knowledge
and fascination with science. The objectives of our study were to develop a practical ready-
to-use assessment tool to test the effectiveness of deeper learning approaches holistically.
The knowledge test we developed can be used to reliably assess students’ integrative knowl-
edge levels. The 36 items from the test were scored with a partial-credit model. Item fit ranged
from 0.83–1.14. Three items only marginally exceeded the reasonable fit range of < 1.11. Con-
sequently, the test offers a good separation reliability of r = .85, and the item difficulty range
matches the range of observed student abilities from a very diverse sample of European stu-
dents (see Figure 1). Furthermore, Tables 2 and 3 reveal a lack of evidence of discriminant val-
idity as the correlations between the measures of the three cognitive and the three content
domains matched or exceeded the respective measures’ reliabilities (Campbell & Fiske,
1959). In other words, all these measures seemed to reflect the same psychological attribute
irrespective of the differences in the item sets that were collated, and thus, for sixth graders,
the cognitive and science facets cannot yet be differentiated. Finally, our knowledge test is a
ready-to-use test with a difficulty range that is suitable for sixth graders.
Also, our fascination test can be used to reliably assess students’ fascination with science
—the second of two deeper learning outcomes. With a separation reliability of r = .94, item
difficulties ranged from δ = −1.99 to δ = 2.44, which is rather narrow in comparison with
students’ ability range (see Figure 2). This moderate mismatch can also be seen in the
finding that five students achieved the maximum score and two students scored
0. Thus, in further applications, items that are potentially very difficult and very easy
should be added and tested.
Even though strongly embraced by educators and policy makers (e.g. Pellegrino & Hilton,
2013), to date, empirically based theory building for deeper learning is rare, probably due to
a lack of a common conceptualisation and definition. However, in the discussion of deeper
learning as ‘skills for the twenty-first century,’ modern teaching approaches do not just have
to foster knowledge acquisition, but they must also address a motivational component that
helps to increase students’ fascination with the subject (e.g. Dede, 2014; Hilton, 2008; Pelle-
grino & Hilton, 2013). Thus, the outcome or consequences of deeper learning teaching can
be defined as knowledge that is sustainably integrated in terms of knowing, applying, and
reasoning as well as a motivational component (i.e. fascination). Although we could not
identify an existing and dedicated holistic empirical approach that could be used to assess
the consequences of deeper learning, we built conceptually on the evidence-based environ-
mental competence model (Roczen, Kaiser, Bogner, & Wilson, 2014). Thus, our measure-
ment tools can be regarded as a first step toward empirically assessing the consequences
of deeper learning and toward helping to provide an empirical knowledge base to foster
the appropriate development and valid assessment of deeper learning approaches.
To our knowledge, the tripartite model has not been used to assess fascination with
science so far. In contrast to most other models, the tripartite model describes the connec-
tion between a latent propensity and manifest expressions thereof (i.e. affect, cognition,
and behaviour). Our measurement approach (i.e. the Campbell Paradigm), which was
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 13
Acknowledgments
This project has received funding from the European Union’s Horizon 2020 research and inno-
vation program under grant agreement No 731872 as part of the joint research project Stories of
Tomorrow. Any opinions, findings, conclusions, or recommendations expressed in this material
are those of the authors and do not necessarily reflect the position of the funding institution.
We wish to thank Florian G. Kaiser for critically discussing the conceptual background and
measurement approach. Furthermore, we thank Jane Zagorsky for her language support and two
anonymous reviewers for their substantial efforts in improving the manuscript.
Disclosure statement
No potential conflict of interest was reported by the authors.
Funding
This work was supported by European Union: [Grant Number 731872].
ORCID
Franz X. Bogner [Link]
References
American Institutes for Research. (2014). Does deeper learning improve student outcomes?
Washington: American Institutes for Research.
Ark, V. D., & Schneider, C. (2014). Deeper learning for every student every day. GettingSmart.
[Link]
14 S. OTTO ET AL.
Bond, T. G., & Fox, C. M. (2007). Applying the Rasch model: Fundamental measurement in the
human sciences (2nd ed.). Mahwah, NJ: Lawrence Erlbaum.
Brislin, R. W. (1980). Translation and content analysis of oral and written material. In H. C.
Triandis, & J. W. Berry (Eds.), Handbook of cross-cultural psychology (pp. 389–444). Boston:
Allyn and Bacon.
Brügger, A., Kaiser, F. G., & Roczen, N. (2011). One for all? Connectedness to nature, inclusion of nature,
environmental identity, and implicit association with nature. European Psychologist, 16, 324–333.
Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-
multimethod matrix. Psychological Bulletin, 56, 81–105.
Davier, M. V., & Davier, A. A. V. (2007). A unified approach to irt scale linking and scale transform-
ations. Methodology, 3(3), 115–124.
Dede, C. (2014). The role of digital technologies in deeper learning. Students at the center: Deeper
learning research series. Harvard: Harvard University.
DeFleur, M. L., & Westie, F. R. (1958). Verbal attitudes and overt acts: An experiment on the sal-
ience of attitudes. American Sociological Review, 23, 667–673.
DeWitt, J., Archer, L., Osborne, J., Dillon, J., Willis, B., & Wong, B. (2011). High aspirations but low
progression: The science aspirations–careers paradox amongst minority ethnic students.
International Journal of Science and Mathematics Education, 9, 243–271.
Dummer, T. J. B., Cook, I. G., Parker, S. L., Barrett, G. A., & Hull, A. P. (2008). Promoting and asses-
sing ‘deep learning’ in geography fieldwork: An evaluation of reflective field diaries. Journal of
Geography in Higher Education, 32(3), 459–479.
Eagly, A. H., & Chaiken, S. (1993). The psychology of attitudes. Fort Worth, TX: Harcourt Brace
Jovanovich College.
Fullan, M., & Langworthy, M. (2014). A rich seam. How new pedagogies find deep learning. London:
Pearson.
Glynn, S. M., Brickman, P., Armstrong, N., & Taasoobshirazi, G. (2011). Science motivation ques-
tionnaire II: Validation with science majors and nonscience majors. Journal of Research in
Science Teaching, 48, 1159–1176.
Grover, S., Pea, R., & Cooper, S. (2015). Designing for deeper learning in a blended computer
science course for middle school students. Computer Science Education, 25, 199–237.
Hambleton, R. K. (1996). Guidelines for adapting educational and psychological tests. European
Journal of Psychological Assessment, 10, 229–244.
Hilton, M. (2008). Research on future skill demands: A workshop summary. Washington, DC:
National Academies Press.
Kaiser, F. G., & Byrka, K. (2015). The Campbell paradigm as a conceptual alternative to the expectation
of hypocrisy in contemporary attitude research. The Journal of Social Psychology, 155, 12–29.
Kaiser, F. G., Byrka, K., & Hartig, T. (2010). Reviving Campbell’s paradigm for attitude research.
Personality and Social Psychology Review, 14, 351–367.
Kaiser, F. G., & Wilson, M. (2019). The Campbell paradigm as a behavior-predictive reinterpreta-
tion of the classical tripartite model of attitudes. European Psychologist, 24, 359–374.
Kolen, M. J., & Brennan, R. L. (2014). Test equating, scaling, and linking: Methods and practices (3rd
ed.). New York: Springer.
Martin, M. O., Mullis, I. V. S., Foy, P., & Hooper, M. (2016). TIMSS 2015 international results in
science. Fourth grade in science. Boston: International Association for the Evaluation of
Educational Achievement.
Masters, G. N. (1982). A rasch model for partial credit scoring. Psychometrika, 47, 149–174.
Meijer, R. R. (1994). The number of guttman errors as a simple and powerful person-fit statistic.
Applied Psychological Measurement, 18, 311–314.
Moore, R. W., & Foy, R. L. H. (1997). The scientific attitude inventory: A revision(SAI II). Journal of
Research in Science Teaching, 34, 327–336.
OECD. (2012). Pisa 2012 assessment and analytical framework. Mathematics, reading, science,
problem solving and financial literacy. Paris: OECD Publishing.
Otto, S., Kröhne, U., & Richter, D. (2018). The dominance of introspective measures and what this
implies: The example of environmental attitude. PLOS ONE, 13(2), e0192907.
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 15
Pellegrino, J. W., & Hilton, M. L. (2013). Education for life and work: Developing transferable knowl-
edge and skills in the 21st century. Washington, DC: National Academies Press.
Podsakoff, P. M., MacKenzie, S. B., Lee, J.-Y., & Podsakoff, N. P. (2003). Common method biases in
behavioral research: A critical review of the literature and recommended remedies. Journal of
Applied Psychology, 88, 879–903.
Rasch, G. (1980). Probabilistic models for some intelligence and attainment tests. Chicago, IL:
University of Chicago.
Robitzsch, A., Kiefer, T., & Wu, M. (Producer). (2018). TAM: Test analysis modules. R package
version 2.12-18. Retrieved from [Link]
Roczen, N., Kaiser, F. G., Bogner, F. X., & Wilson, M. (2014). A Competence Model for
Environmental Education. Environment and Behavior, 46, 972–992.
Rosenberg, M. J., & Hovland, C. I. (1960). Cognitive, affective, and behavioral components of atti-
tudes. In M. J. Rosenberg & C. I. Hovland (Eds.), Attitude organization and change: An analysis
of consistency among attitude components (pp. 1–14). New Haven, CT: Yale University.
Schumm, M. F., & Bogner, F. X. (2016). Measuring adolescent science motivation. International
Journal of Science Education, 38, 434–449.
Sjøberg, S., & Schreiner, C. (2010). The rose project. An overview and key findings. Oslo: University
of Oslo.
Wilson, M. (2013). Using the concept of a measurement system to characterize measurement
models used in psychometrics. Measurement, 46, 3766–3774.
Wu, M., & Adams, R. J. (2013). Properties of Rasch residual fit statistics. Journal of Applied
Measurement, 14, 339–355.
Zeiser, K. L., Taylor, J., Rickles, J., Garet, M. S., & Segeritz, M. (2014). Evidence of deeper learning
outcomes. New York: American Institutes for Research.
Appendices
Appendix A
Items of the Science Knowledge Test
How much do you know about Science?
The following questions are about Earth, life, and science. Please read the questions carefully and
think thoroughly about the answers. There might be more than one correct answer, so pay atten-
tion. To answer a question, click on the box of the response option(s) you think is/are correct. It’s
OK if you don’t know the answers to every question. If you don’t know an answer, please move on
to the next question. Please answer the questions on your own.
Good luck!
77A.E (Applying)
The picture shows the three main layers of the Earth. Where is it the hottest?
16 S. OTTO ET AL.
A: layer A
B: layer B
C: layer C
D: All three layers are the same temperature.
A: It disappears.
B: It turns into clouds.
C: It rains down.
D: It changes to snow.
A: underground water
B: sandy soil
C: fossils of fish
D: salty lakes
A: Part A
B: Part B
C: Part C
D: Part D
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 17
A: feathers
B: hair
C: internal skeleton
D: wings
A: to protect seeds
B: to produce food for seeds
C: to disperse seeds via animals
D: to store water for seed germination
A: It changes colour.
B: It becomes heavier.
C: It changes into water vapour.
D: It starts bubbling.
18 S. OTTO ET AL.
Gerry connects a battery, a light bulb, and some wire as shown below. Will the light bulb light and
why?
A: No, because the ‘+’ side of the battery is pointed in the wrong direction.
B: No, because it is not a full circuit.
C: Yes, because it is connected.
D: No, because the light bulb is broken.
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 19
A: Mars
B: the Sun
C: the Moon
D: all other planets
A: dead trees
B: water
C: oil wells
D: rocks
A: grains of sand
B: lumps of clay
C: layers of gravel
D: decaying plants and animals
A: hot water
B: solar power
C: electricity
D: drinking water
A: fish
B: stone
C: water lily
D: trees
A: air
B: soil
C: water
D: sunlight
A: air
B: wood
C: sand
D: gasoline
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 21
A: magnetism
B: gravity
C: air resistance
D: the push from your hand
A: gravity
B: magnetism
C: electricity
D: none of the above
Terry tested four rocks to see how hard they are. He rubbed each of them against some hard steel for
one minute. He drew pictures of what they looked like before and after he rubbed them.
Which of Terry`s rocks is the hardest?
22 S. OTTO ET AL.
A: rock A
B: rock B
C: rock C
D: rock D
A: location A
B: location B
C: location C
D: location D
89R.E
The table below shows some weather information for four different towns during a 24-hour period.
In which town did it most likely snow?
Clouds in the Sky Lowest Temperature Highest Temperature
Town A No 10°C 25°C
Town B Yes 20°C 30°C
Town C No −10°C −1°C
Town D Yes −15°C 5°C
A: Town A
B: Town B
C: Town C
D: Town D
After some time, they compared the plants and saw that there was a large difference in their growth,
as shown in the picture below.
In which way might Carl have treated his plant differently from the way Jan treated hers?
Appendix B
Item Difficulties and Infit MS for the Science Fascination Items
calibration calibration
without with technic
Item technic Items Items
Attitude Science Infit Infit
Code Component Field δ MS δ MS
FB05 Behavior Astr. Watch documentaries about the universe. 0.78 0.90 0.73 0.89
FB07 Behavior Astr. Try to find the star constellations in the sky. 0.87 1.03 0.82 1.02
FB01 Behavior Astr. Look for the polar star. 1.52 0.96 1.46 0.95
FB06 Behavior Astr. Read about astronomy in books or magazines. 1.56 0.88 1.50 0.87
FB12 Behavior Biol. Watch animals. −0.39 1.10 −0.40 1.08
FB16 Behavior Biol. Plant seeds and watch them grow. 0.87 0.95 0.83 0.95
FB09 Behavior Biol. Collect small water animals (e.g. tadpole. fish). 1.42 1.08 1.37 1.06
FB10 Behavior Biol. Examine small animals and plants under the 1.83 0.92 1.77 0.91
microscope.
FB21 Behavior Chem. Bake bread, pastry, cake, etc. 0.09 1.23 1.96 0.94
(Continued )
26 S. OTTO ET AL.
Continued.
calibration calibration
without with technic
Item technic Items Items
Attitude Science Infit Infit
Code Component Field δ MS δ MS
FB05 Behavior Astr. Watch documentaries about the universe. 0.78 0.90 0.73 0.89
FB18 Behavior Chem. I experimented with ice. 1.18 1.03 0.07 1.19
FB20 Behavior Chem. Use a science kit for chemistry 2.01 0.96 2.05 1.08
FB19 Behavior Chem. Dyed clothes. 2.11 1.10 1.14 1.00
FB55 Behavior Gen. Played educational video games on scientific 1.32 0.99 1.27 0.96
topics.
FB53 Behavior Gen. Watch science shows. 1.35 0.89 1.29 0.89
FB52 Behavior Gen. Read about science in books or magazines. 1.56 0.87 1.50 0.88
FB54 Behavior Gen. Visit science exhibitions. 1.84 0.92 1.77 0.91
FB30 Behavior Geo. Collect different stones. 0.51 1.00 0.47 0.99
FB26 Behavior Geo. Watch a thunderstorm or storm. 0.55 1.04 0.51 1.02
FB25 Behavior Geo. Use an atlas. 0.87 1.05 0.83 1.03
FB29 Behavior Geo. Read about geography in books or magazines. 1.12 0.96 1.08 0.95
FB34 Behavior Phys. Measure the temperature with a thermometer. 0.54 1.06 0.51 1.02
FB33 Behavior Phys. Read about physics in books or magazines. 1.18 0.91 1.14 0.90
FB38 Behavior Phys. Used a rope and pulley for lifting heavy things. 2.32 1.00 2.25 0.98
FB36 Behavior Phys. I constructed a windmill, watermill, or 2.44 0.94 2.38 0.92
waterwheel.
FB41 Behavior Tech. Install programmes or apps. – – −0.82 1.35
FB47 Behavior Tech. Opened a device (radio, watch, computer, – – 0.42 1.14
telephone, etc.) to find out how it works.
FB42 Behavior Tech. Modify system settings on a computer or phone. – – 0.54 1.14
FB50 Behavior Tech. Use tools like a saw, screwdriver or hammer. – – 0.88 1.11
FC07 Cognition Astr. The discovery of new planets and stars is −1.21 0.88 −1.21 0.87
important.
FC01 Cognition Astr. To learn about space and stars in school is −0.63 0.85 −0.64 0.86
important.
FC02 Cognition Astr. Astronomy is important. −0.48 0.95 −0.50 0.94
FC06 Cognition Astr. Everybody should have some basic knowledge −0.06 0.89 −0.09 0.89
about the stars.
FC15 Cognition Biol. It is necessary for humans to understand nature. −1.38 1.00 −1.37 0.98
FC14 Cognition Biol. Everybody should have basic knowledge about −1.12 1.02 −1.11 0.99
the functioning of the human body.
FC10 Cognition Biol. Biology is important. −0.75 0.96 −0.75 0.94
FC09 Cognition Biol. To learn about biology in school is important. −0.65 0.95 −0.65 0.94
FC17 Cognition Chem. To learn about chemistry in school is important. −0.81 0.91 −0.81 0.90
FC22 Cognition Chem. Everybody should be able to cook or bake −0.74 1.28 −0.74 1.22
something.
FC23 Cognition Chem. It is important for people to know something −0.73 0.92 −0.73 0.91
about chemistry.
FC19 Cognition Chem. Every student should experiment with 0.17 1.07 0.14 1.05
chemicals.
FC53 Cognition Gen. Scientific discoveries are necessary. −1.34 1.00 −1.33 0.96
FC52 Cognition Gen. To learn about science in school is important. −1.21 0.95 −1.20 0.93
FC50 Cognition Gen. Science is important. −1.08 0.93 −1.08 0.92
FC51 Cognition Gen. Everybody should know something about −0.98 1.01 −0.98 0.98
science.
FC31 Cognition Geo. Everybody should know something about the −1.99 0.91 −1.97 0.89
earth.
FC25 Cognition Geo. To learn about the earth in school is important. −1.88 0.94 −1.86 0.92
FC27 Cognition Geo. Everybody should be able to read a map. −1.11 1.03 −1.11 0.99
FC30 Cognition Geo. It is necessary that humans understand the −0.09 0.95 −0.12 0.93
greenhouse effect.
FC33 Cognition Phys. To learn about physics in school is important. −1.14 0.91 −1.12 0.89
FC34 Cognition Phys. Physics is important. −1.11 0.86 −1.10 0.86
FC36 Cognition Phys. The discovery of gravitation was one of the most −1.10 1.11 −1.10 1.06
important discoveries.
(Continued )
INTERNATIONAL JOURNAL OF SCIENCE EDUCATION 27
Continued.
calibration calibration
without with technic
Item technic Items Items
Attitude Science Infit Infit
Code Component Field δ MS δ MS
FB05 Behavior Astr. Watch documentaries about the universe. 0.78 0.90 0.73 0.89
FC39 Cognition Phys. Everybody should understand some basic −0.44 0.90 −0.45 0.88
physics.
FC43 Cognition Tech. It is important to know how to work with a – – −1.44 1.11
computer.
FC48 Cognition Tech. It is necessary to develop new technologies. – – −1.15 1.03
FC44 Cognition Tech. To learn about technical tasks in school is – – −0.99 1.11
important.
FC42 Cognition Tech. It is necessary to develop intelligent computers. – – −0.87 1.14
FA08 Emotion Astr. The possibility of life outside earth is fascinating. −1.69 1.07 −1.67 1.03
FA04 Emotion Astr. I like to learn about stars, planets, and the −1.45 0.86 −1.44 0.87
universe.
FA06 Emotion Astr. I wonder about unresolved mysteries in outer −0.93 1.02 −0.93 0.99
space.
FA02 Emotion Astr. The Moon interests me. −0.86 1.04 −0.87 1.02
FA09 Emotion Biol. Animals are exciting. −1.34 1.15 −1.32 1.12
FA14 Emotion Biol. The origin and evolution of life on earth −1.04 0.96 −1.03 0.95
fascinates me.
FA15 Emotion Biol. I would like to know about heredity, and how −0.19 1.04 −0.21 1.00
genes influence how we develop.
FA11 Emotion Biol. I like to learn about plants. −0.07 1.06 −0.10 1.05
FA21 Emotion Chem. I am thrilled to learn about explosive chemicals. −1.10 1.12 −1.09 1.07
FA24 Emotion Chem. I would like to know about how crude oil is −0.22 1.03 −0.23 1.00
converted to other materials, such as plastics
and textiles.
FA23 Emotion Chem. I am curious about the influence of temperature −0.17 1.00 −0.18 0.98
and pressure conditions on chemical
reactions.
FA19 Emotion Chem. I am interested in detergents, soap and how 0.96 1.11 0.92 1.10
they work.
FA52 Emotion Gen. Science is fascinating. −0.73 0.91 −0.74 0.90
FA51 Emotion Gen. New knowledge about science excites me. −0.57 0.87 −0.58 0.87
FA49 Emotion Gen. I love Science. −0.48 0.93 −0.49 0.93
FA50 Emotion Gen. I am curious about scientific topics. −0.44 0.89 −0.05 0.87
FA25 Emotion Geo. I want to know, what can be done to ensure safe −0.97 1.08 −0.97 1.05
drinking water.
FA27 Emotion Geo. It is interesting how mountains, rivers and −0.60 1.10 −0.61 1.07
oceans develop and change.
FA28 Emotion Geo. Clouds, rain and the weather impress me. 0.13 1.18 0.10 1.15
FA32 Emotion Geo. I would like to know more about the 0.13 0.99 0.10 0.98
greenhouse effect.
FA34 Emotion Phys. I want to learn about different sources of −0.45 0.99 −0.46 0.97
energy.
FA40 Emotion Phys. The use of satellites for communication and −0.39 1.00 −0.40 0.96
other purposes fascinates me.
FA36 Emotion Phys. It is exciting how different musical instruments −0.14 1.17 −0.15 1.14
produce different sounds.
FA37 Emotion Phys. I am eager to learn about atoms and molecules. −0.05 1.02 −0.07 0.99
FA43 Emotion Tech. How computers work impresses me. – – −0.85 1.12
FA41 Emotion Tech. I am overwhelmed by space rockets. – – −0.46 1.04
FA46 Emotion Tech. Lasers for technical purposes (CD-players, bar- – – −0.03 1.04
code readers, etc.) attract my attention.
FA48 Emotion Tech. I am curious about how a nuclear power plant – – −0.03 1.02
functions.
Note. Items are ordered by attitude component. δ = Item difficulty estimate, Astr. = Astronomy,
Biol. = Biology, Chem. = Chemistry, Geo. = Geoscience, Phys. = Physics, Tech. = Technology,
Gen. = General.
28 S. OTTO ET AL.
Appendix C
Item Difficulties and Infit MS for the Science Knowledge Items
Code Cognitive Domain Content Domain Difficulty δ per step Infit MS
077A.E Applying Earth Science −1.27 0.97
073A.E Applying Earth Science −1.21 1.04
021A.E Applying Earth Science 1.03 1.00
075A.E Applying Earth Science 1.51 1.07
066A.L Applying Life Science −1.95 0.94
040A.L Applying Life Science −1.04 1.03
1.09 1.05
037A.L Applying Life Science −0.62 0.98
034A.L Applying Life Science 0.11 1.12
069A.L Applying Physical Science −2.35 0.87
2.22 0.99
033A.P Applying Physical Science −0.12 1.06
032A.P Applying Physical Science 0.12 0.92
030A.P Applying Physical Science 0.35 1.00
103 K.E Knowing Earth Science −1.72 1.08
105 K.E Knowing Earth Science −0.21 0.98
001 K.E Knowing Earth Science −0.08 1.01
004 K.E Knowing Earth Science 1.49 1.01
015 K.L Knowing Life Science −0.75 0.93
016 K.L Knowing Life Science 1.03 0.95
093 K.L Knowing Life Science 1.09 1.12
018 K.L Knowing Life Science 2.80 1.01
009 K.P Knowing Physical Science −2.27 0.86
0.56 1.05
012 K.P Knowing Physical Science −1.18 1.01
011 K.P Knowing Physical Science −0.69 0.86
101 K.P Knowing Physical Science −0.59 0.98
112R.E Reasoning Earth Science −0.85 0.95
088R.E Reasoning Earth Science −0.59 0.83
090R.E Reasoning Earth Science −0.12 1.14
089R.E Reasoning Earth Science 0.48 0.99
059R.L Reasoning Life Science −2.04 0.98
1.55 1.02
057R.L Reasoning Life Science 0.81 0.98
062R.L Reasoning Life Science 1.10 0.97
080R.L Reasoning Life Science 1.23 1.03
047R.P Reasoning Physical Science −0.48 1.11
083R.P Reasoning Physical Science −0.17 0.97
048R.P Reasoning Physical Science 0.38 1.00
045R.P Reasoning Physical Science 0.71 1.04
Note. Items are ordered by cognitive domain. For the items 009K.P, 040A.L, 059R.L, and 069A.L
two steps were calculated, the first line contains the difficulty and fit estimates for the lower step,
the second line the estimates of the higher step.