See discussions, stats, and author profiles for this publication at: [Link]
net/publication/390922030
Ten Guidelines for Scoring Psychological Assessments Using Artificial
Intelligence
Article in European Journal of Psychological Assessment · May 2025
DOI: 10.1027/1015-5759/a000904
CITATION READS
1 390
3 authors:
Denis Dumas Samuel Greiff
University of Georgia Technical University of Munich
117 PUBLICATIONS 2,567 CITATIONS 407 PUBLICATIONS 10,656 CITATIONS
SEE PROFILE SEE PROFILE
Eunike Wetzel
Rheinland-Pfälzische Technische Universität Kaiserslautern-Landau
69 PUBLICATIONS 1,911 CITATIONS
SEE PROFILE
All content following this page was uploaded by Denis Dumas on 18 April 2025.
The user has requested enhancement of the downloaded file.
GUIDELINES FOR AI TEST-SCORING 1
Please cite this post-print as:
Dumas, D., Greiff, S., & Wetzel, E., (2025). Ten guidelines for scoring psychological
assessments using artificial intelligence. European Journal of Psychological Assessment.
TEN GUIDELINES FOR SCORING PSYCHOLOGICAL ASSESSMENTS
USING ARTIFICIAL INTELLIGENCE
Denis Dumas
Department of Educational Psychology, University of Georgia
Samuel Greiff
School of Social Sciences & Technology & Centre of International Student Assessment,
Technical University Munich
Eunike Wetzel
University of Kaiserslautern-Landau
Author Note
Correspondence concerning this paper should be addressed to Denis Dumas, Department of
Educational Psychology, University of Georgia: 624 Aderhold Hall, 110 Carlton St., Athens,
GA, 30602. Email: [Link]@[Link].
GUIDELINES FOR AI TEST-SCORING 2
Ten Guidelines for Scoring Psychological Assessments Using Artificial Intelligence
Man makes machines to man the machines that make the machines
-Pete Townshend (1989)
As Pete Townshend (of the legendary rock band The Who) mused in the epigraph above,
modern history is largely defined by a continual increase in the mechanized automation of
human work (e.g., Autor et al., 2003). This process of automation has greatly accelerated and
expanded because of the rise of generative Artificial Intelligence (AI) systems such as Large-
Language Models and Vision Transformer Models. Through the large-scale statistical modeling
of input-data from humans (e.g., text, classifications, quantities, images), these AI models can be
trained to approximate or mimic those humans, thereby producing AI-generated outputs that
have similar properties to their human-generated training inputs (Sejnowksi, 2024). Because of
this, AI now has the capacity to automate tasks that, up until very recently, were the auspices
solely of humans.
One area of novel automation in which AI has already shown great promise is the scoring
of psychological assessments (e.g., Kjell et al., 2024; Organisciak et al., 2023; Yavuz et al.,
2024). For decades, psychologists have heavily relied upon some assessment-types (e.g., Likert
style questionnaires in mental health, or multiple-choice tests in education) that, because of their
closed-ended and well-structured format, could be scored rapidly in a straightforward and
automatic way. In contrast, assessments with a more open-ended or ill-structured item format
(e.g., essays, narratives, drawings) produce data that is not initially numeric, and therefore have
historically relied on human judgments to quantify the responses. Because these open-ended
assessments could not be scored as rapidly and easily as their closed-ended competitors, they
were more difficult and expensive to administer, and therefore much less utilized. However, if
GUIDELINES FOR AI TEST-SCORING 3
datasets of human-judged responses exist for a given open-ended assessment, then AI systems
can now be trained to mimic those human-judges, thereby automating the process of scoring.
We are in the middle of an immensely interesting revolution in psychological assessment.
If participants’ open-ended writing or drawing can be reliably judged and quantified instantly
using AI, this measurement process may one day eclipse close-ended Likert or multiple-choice
assessments in popularity. This shift is relevant across areas of assessment spanning from
education through mental health, as well as both summative and formative assessment purposes.
But just because the AI-based scoring of psychological assessments is now technically possible
and seems to confer advantages vis-à-vis the efficiency of current assessment practices, does not
mean it is by-definition a good thing. On the contrary, irresponsible or unethical applications of
AI to psychological assessments may have the potential to do untold harm if the quantities they
produce are unreliable or invalid.
For this reason, some in the field of psychological assessment (e.g., Burstein et al., 2024;
Dorsey & Michaels, 2022; Ho, 2024; Tay et al., 2022) have begun to consider norms,
recommendations, or standards for how the field can successfully import AI methodology to
automate the scoring of open-ended assessments, while carefully maintaining the integrity of our
field, and especially protecting the rights of the participants we serve. In the current editorial, we
build on this past work, as well as our own research and experience with AI-based assessment, to
forward ten key guidelines that we see as critically important as the field moves toward an
increasingly AI-automated assessment practice.
<Insert Table 1 about here>
GUIDELINES FOR AI TEST-SCORING 4
Do: Disclose the Use of AI to Participants Before They Are Assessed
Before a participant responds to an assessment, ethical practice is to provide them with
key pieces of information about how their responses will be used, including some explanation of
the process of quantification and scoring. Disclosing the usage of AI to quantify open-ended
responses is analogous to the desirable (although not necessarily common) practice of providing
participants with a rubric or checklist that will be used to rate their responses prior to responding;
it allows all participants to be on the same page about how their responses are being judged, and
why. Of course, supervised AI models that are trained to mimic or re-create past human
judgments might not be able to explain precisely how they come up with their ratings (Joyce et
al., 2023), which potentially undermines the transparency of the assessment scoring process. In
such cases, best practice might be to provide participants with an example of a response that was
rated highly on the attribute being assessed by the AI, so participants can see for themselves
what is expected of them.
Do Not: Train AI-Scoring Models on AI-Generated Material
In any field of application, AI models need to be trained on the highest quality data
available. In the ever-burgeoning literature on AI, much has already been written about the
“garbage in, garbage out” phenomenon (e.g., Vidgen & Derczynski, 2020). When supervised AI
models are trained on datasets that have disadvantageous or undesirable properties, they
inevitably reproduce those same problems. In our view, one perhaps so-far overlooked aspect of
this phenomenon is that AI models trained to judge responses to a psychological assessment
must be trained on human judgments, and not prior AI-judgments. Imagine a situation where a
psychology laboratory is collecting human judgments of test responses, with the aim of training
an AI-based automatic scoring system. If one or more of those human judges (perhaps a student
GUIDELINES FOR AI TEST-SCORING 5
research assistant or crowdsourced worker) secretly uses a commercially available AI (e.g.,
ChatGPT) to do their ratings for them, then a circular scenario would be created whereby the AI
model might seem internally valid by available metrics (i.e., it will correlate well with the
judgments it was trained on), when actually it is simply agreeing with another AI.
Do: Account for Uncertainty in AI Response-Ratings
When humans rate responses to open-ended psychological assessments, inevitably there
are some responses for which they are more confident in their ratings, and other responses that
represent edge-cases and for which they have lower confidence in their ratings. AI scoring
models trained on those human ratings will show the same phenomenon when they rate new
responses that they were not trained on. For some responses that are well-represented in their
training dataset, they will be strongly confident, but for other responses they will be less so. One
way to address this issue is to ensure that AI models are trained on large datasets of responses
with several human judgments for each response that meet inter-rater reliability standards
(typically agreement above .80), and hence they are representative and consistent enough to
provide the AI with a reasonably authoritative consensus to mimic (Dumas, 2023a). Another
method is to flag responses for which the AI scoring model has low confidence in its ratings and
bring in an additional human judge (or two) to provide ratings for them. In cases where AI
ratings are highly uncertain, they should not be used as part of the calculation of a participants’
test score, and human ratings should be used instead.
Do Not: Use AI to Score AI-Generated or AI-Assisted Responses
Clearly, when we think about psychological assessment, we tend to assume that a human
is responding to the assessment, and that the intention is to measure some psychological attribute
of that human. But the wide availability of commercial AI models like ChatGPT subverts this
GUIDELINES FOR AI TEST-SCORING 6
assumption. Today, whenever a test is administered away from a watchful proctor, it is a real
possibility that participants might choose to utilize such an AI system to assist them in
formulating their responses as a form of cheating or lazy responding (Susnak & McIntosh, 2024).
Some preliminary evidence suggests that AI scoring systems might prefer, and score more
highly, responses written by AI even in cases where human judges do not (e.g., Tang et al.,
2024). This finding makes sense, given both the AI model that is generating the work and the AI
model that is scoring it might have some training data in common, and may use similar methods
to express (or conversely, observe) psychological attributes in writing or drawing. This finding
also implies that AI scoring might exacerbate the temptation for participants to cheat using AI.
Thus, if AI is scoring an assessment, special care should be taken to ensure that participants do
not use AI to generate their responses.
Do: Predict Human-Scored Criteria to Build a Validity Argument
In order to build the case for the validity of an AI-scored psychological assessment,
evidence must be gathered that its scores predict human-scored criteria. As already discussed,
AI-scoring of psychological assessments opens the door for multiple possible circularly invalid
conclusions that could involve AI-generated material in the training dataset for the scoring
model, or AI-generated material in the responses being scored. Both of these circular situations
would mean that validity evidence internal to the assessment would likely be inflated, because it
would appear that the AI scoring model is capturing the target construct very well, while in
reality it is simply agreeing with another AI. Another similar and equally problematic circular
validation process would be if the correlation between two AI-scored assessments were used as
validity evidence: AI-scored assessments must covary with human-rated assessments in order to
show that their judgments are not insular or idiosyncratic to AI.
GUIDELINES FOR AI TEST-SCORING 7
Do Not: Assume AI-Scoring Models Generalize Across Cultures
Because AI scoring models are capable of quantifying open-ended assessment responses
rapidly and automatically, they open the door for a major scaling-up of those assessment efforts.
New and interesting open-ended assessments with a fully international scope, scored via AI, may
soon be developed. This possibility is exciting, but it also creates the need to ensure the fairness
of the AI-scoring process across cultural groups. As with more typical assessment practices,
which become less and less likely to yield invariant and comparable scores across participants as
the cultural distance between them increases (Dong & Dumas, 2020), the assumption of “one
size fits all” seems unlikely to hold for AI scoring models. Especially if participants come from a
culture or language background that was not represented by the human judgments used to train
the AI, it seems unlikely that those AI-scores could be valid for them. One way to help avoid
these issues in the future is to utilize very wide-ranging and culturally diverse human judgments
to train AI models, in an attempt to produce a cosmopolitan scoring model that can genuinely
handle responses in multiple languages, or from multiple cultural perspectives. Even so,
psychometric methods such as measurement invariance (Meredith, 1993), differential item
functioning (Holland & Wainer, 2012), or regression-based predictive fairness metrics (Dumas et
al., 2023b) should be utilized to address the empirical question of AI-score fairness.
Do: Establish A Clear Protocol for Human Oversight and Auditing
Even the most well-designed and well-validated psychological assessment can never be
said to be completely finished. Over time, test scores can decrease in quality as cultural shifts
make the items less relevant to the target construct, participants become aware of the assessment
structure and how to game it without meaningfully responding, or the target population shifts or
expands to include new participants for whom the test was not originally designed. Any of these
GUIDELINES FOR AI TEST-SCORING 8
changes, and many more, can result in downstream shifts in the psychometric parameters of the
scoring model, decreases in reliability, larger standard errors around scores, and weaker validity
predictions (referred to as drift). When AI is used to score a test, it adds an additional layer of
complexity and possible problems that accrue over time: drift no longer refers only to the
parameters of a psychometric scoring model, but to the parameters of the AI-model used to rate
the responses as well. AI models need to be continually updated, trained, and re-validated on the
best-available human-rated data, and the effect of these changes on any psychometric scoring
model used to aggregate the AI-ratings (and hence the comparability of the scores) should also
be studied.
Do Not: Directly Compare Participant Scores Across AI Models or Versions
Currently, AI models are changing rapidly, with new advances arriving nearly monthly.
This means that proprietary and commercially available AI models that are owned by technology
companies (e.g., OpenAI owns ChatGPT and a variety of other models) are being continually
updated and changed in ways that users of those models cannot control, and completely new
models are coming out regularly. We suspect that, within the area of psychological assessment, it
will be rare for researchers or even testing companies to build their own generative AI model
from the ground up, and instead it is likely that psychological researchers will fine-tune existing
models. Every time the base model is changed by the technology company that controls it, those
changes will inevitably flow downstream and affect any researcher fine-tuned AI model built on
top of it, as well as any psychometric model used to aggregate AI-ratings of participant
responses. Each AI model update will result in some drift in the parameters of the scoring model
for the test and therefore also changes in the distribution of the scores. This means that, after an
AI update, individual participant scores become non-comparable with those from earlier
GUIDELINES FOR AI TEST-SCORING 9
versions. The degree to which those scores have changed due to the AI model update and
resulting psychometric parameter drift should always be an empirical question.
Do: Explain the Use of AI to Participants in the Final Score Report
As psychometricians, it can sometimes be woefully easy to dismiss the final score report
given to test stake-holders (e.g., clinicians, educators, participants, or their parents) as an
appendix on the assessment process. But the score report is the main way in which assessment
stake-holders interpret a score and understand what it might mean for them. For this reason, all
score reports should be designed and written with careful attention paid to stake-holder
comprehension of the testing process and score meaning. Score reports for AI-scored
assessments are certainly no exception to this, and for this reason special care should be taken to
explain the way AI was used to calculate scores in the final report. The emphasis here is to
provide the information that allows the participant to understand that the AI model has been
carefully developed and that it provides the best-available judgments of the psychological
construct being assessed, so that they feel that their score is trustworthy and meaningful to them.
Do Not: Expect Stake-holders to Trust a Black-Box Without Justification
As AI models become more advanced, the processes through which they make decisions
and generate output also become increasingly more opaque, resembling the classic metaphor of
the “black-box” (Bearman & Ajjawi, 2023). At the same time, as AI becomes more and more
capable of general tasks, individuals across disciplines might be tempted simply to trust its
output, without rigorous human justification for its validity. Especially in higher-stakes
psychological or educational contexts, stake-holders should never be expected simply to take AI-
based test scores at face value, without specific evidence forwarded to support the validity of the
scoring process. AI is a very useful tool to make the process of assessment scoring more
GUIDELINES FOR AI TEST-SCORING 10
efficient, but it should never be a replacement for careful thought on the part of a human
psychologist or psychometrician.
Conclusion
We are cautiously optimistic about the potential future role of AI in scoring
psychological assessments. Clearly, the field’s reliance on highly efficient closed-ended items
(e.g., multiple choice items) has produced some limitations in our capacity to measure the
complex and nuanced psychological constructs which we, as psychologists, are interested in.
But, without a sharp (and unlikely) increase in the resources available for psychological science,
the time- and cost-efficiency of our assessment practices is likely to be an important aspect of
our work going forward. For this reason, we see the use of AI as one potentially important
methodology for expanding the scope of psychological assessment. At EJPA, we welcome both
methodological and empirical submissions addressing this topic. If AI is applied to the scoring of
psychological assessments in a careful, responsible, and ethical way, it may be that it will help
with what we see as the ultimate goal of psychology: to contribute to human flourishing through
a better understanding of the human mind.
GUIDELINES FOR AI TEST-SCORING 11
References
Autor, D. H., Levy, F., & Murnane, R. J. (2003). The skill content of recent technological change: An
empirical exploration. The Quarterly Journal of Economics, 118(4), 1279–1333.
[Link]
Bearman, M., & Ajjawi, R. (2023). Learning to work with the black box: Pedagogy for a world with
artificial intelligence. British Journal of Educational Technology, 54(5), 1160–1173.
[Link]
Burstein, J., LaFlair, G. T., Yancey, K., Von Davier, A. A., & Dotan, R. (2024). Responsible AI for
Test Equity and Quality: The Duolingo English Test as a Case Study (arXiv:2409.07476). arXiv.
[Link]
Dong, Y., & Dumas, D. (2020). Are personality measures valid for different populations? A
systematic review of measurement invariance across cultures, gender, and age. Personality and
Individual Differences, 160, 109956. [Link]
Dorsey, D. W., & Michaels, H. R. (2022). Validity arguments meet artificial intelligence in innovative
educational assessment. Journal of Educational Measurement, 59(3), 267–271.
[Link]
Dumas, D., Acar, S., Berthiaume, K., Organisciak, P., Eby, D., Grajzel, K., Vlaamster, T., Newman,
M., & Carrera, M. (2023a). What makes children’s responses to creativity assessments difficult
to judge reliably? The Journal of Creative Behavior, 57(3), 419–438.
[Link]
Dumas, D., Dong, Y., & McNeish, D. (2023b). How fair is my test? A ratio coefficient to help
represent consequential validity. European Journal of Psychological Assessment, 39(6), 416–
423. [Link]
GUIDELINES FOR AI TEST-SCORING 12
Ho, A. D. (2024). Artificial intelligence and educational measurement: Opportunities and threats.
Journal of Educational and Behavioral Statistics, 49(5), 715–722.
[Link]
Holland, P. W., & Wainer, H. (2012). Differential item functioning. Routledge.
Joyce, D. W., Kormilitzin, A., Smith, K. A., & Cipriani, A. (2023). Explainable artificial intelligence
for mental health through transparency and interpretability for understandability. Npj Digital
Medicine, 6(1), 1–7. [Link]
Kjell, O. N. E., Kjell, K., & Schwartz, H. A. (2024). Beyond rating scales: With targeted evaluation,
large language models are poised for psychological assessment. Psychiatry Research, 333,
115667. [Link]
Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance.
Psychometrika, 58(4), 525–543. [Link]
Organisciak, P., Acar, S., Dumas, D., & Berthiaume, K. (2023). Beyond semantic distance:
Automated scoring of divergent thinking greatly improves with large language models. Thinking
Skills and Creativity, 49, 101356.
Sejnowski, T. J. (2024). ChatGPT and the future of AI: The deep language revolution. The MIT
Press.
Susnjak, T., & McIntosh, T. R. (2024). ChatGPT: The end of online exam integrity? Education
Sciences, 14(6), Article 6. [Link]
Tang, M., Hofreiter, S., Werner, C. H., Zielinska, A., & Karwowski, M. (2024). “Who” is the best
creative thinking partner? An experimental investigation of human-human, human-internet, and
human-AI co-creation. Journal of Creative Behavior. [Link]
GUIDELINES FOR AI TEST-SCORING 13
Tay, L., Woo, S. E., Hickman, L., Booth, B. M., & D’Mello, S. (2022). A conceptual framework for
investigating and mitigating machine-learning measurement bias (MLMB) in psychological
assessment. Advances in Methods and Practices in Psychological Science, 5(1),
25152459211061337. [Link]
Vidgen, B., & Derczynski, L. (2020). Directions in abusive language training data, a systematic
review: Garbage in, garbage out. PLOS ONE, 15(12), e0243300.
[Link]
Yavuz, F., Çelik, Ö., & Yavaş Çelik, G. (2025). Utilizing large language models for EFL essay
grading: An examination of reliability and validity in rubric-based assessments. British Journal
of Educational Technology, 56(1), 150–166. [Link]
GUIDELINES FOR AI TEST-SCORING 14
Table 1.
Ten Guidelines for AI Scoring of Psychological Assessments
Do: Do Not:
1. Disclose the Use of AI to Participants 2. Train AI-Scoring Models on AI-
Before They Are Assessed Generated Material
3. Account for Uncertainty in AI Response- 4. Use AI to Score AI-Generated or AI-
Ratings Assisted Responses
5. Predict Human-Scored Criteria to Build a 6. Assume AI-Scoring Models Generalize
Validity Argument Across Cultures
7. Establish A Clear Protocol for Human 8. Directly Compare Participant Scores
Oversight and Auditing Across AI Models or Versions
9. Explain the Use of AI to Participants in 10. Expect Stake-holders to Trust a Black-
the Final Score Report Box Without Justification
View publication stats