0% found this document useful (0 votes)
20 views15 pages

AI Guidelines for Psychological Assessment

The article outlines ten guidelines for scoring psychological assessments using artificial intelligence (AI), emphasizing the importance of ethical practices, transparency, and human oversight. It discusses the potential benefits and risks of AI in psychological assessment, particularly in automating the scoring of open-ended responses. The guidelines aim to ensure the integrity of assessments and protect participant rights while leveraging AI technology effectively.

Uploaded by

Fayiz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
20 views15 pages

AI Guidelines for Psychological Assessment

The article outlines ten guidelines for scoring psychological assessments using artificial intelligence (AI), emphasizing the importance of ethical practices, transparency, and human oversight. It discusses the potential benefits and risks of AI in psychological assessment, particularly in automating the scoring of open-ended responses. The guidelines aim to ensure the integrity of assessments and protect participant rights while leveraging AI technology effectively.

Uploaded by

Fayiz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

See discussions, stats, and author profiles for this publication at: [Link]

net/publication/390922030

Ten Guidelines for Scoring Psychological Assessments Using Artificial


Intelligence

Article in European Journal of Psychological Assessment · May 2025


DOI: 10.1027/1015-5759/a000904

CITATION READS
1 390

3 authors:

Denis Dumas Samuel Greiff


University of Georgia Technical University of Munich
117 PUBLICATIONS 2,567 CITATIONS 407 PUBLICATIONS 10,656 CITATIONS

SEE PROFILE SEE PROFILE

Eunike Wetzel
Rheinland-Pfälzische Technische Universität Kaiserslautern-Landau
69 PUBLICATIONS 1,911 CITATIONS

SEE PROFILE

All content following this page was uploaded by Denis Dumas on 18 April 2025.

The user has requested enhancement of the downloaded file.


GUIDELINES FOR AI TEST-SCORING 1

Please cite this post-print as:

Dumas, D., Greiff, S., & Wetzel, E., (2025). Ten guidelines for scoring psychological
assessments using artificial intelligence. European Journal of Psychological Assessment.

TEN GUIDELINES FOR SCORING PSYCHOLOGICAL ASSESSMENTS

USING ARTIFICIAL INTELLIGENCE

Denis Dumas

Department of Educational Psychology, University of Georgia

Samuel Greiff

School of Social Sciences & Technology & Centre of International Student Assessment,

Technical University Munich

Eunike Wetzel

University of Kaiserslautern-Landau

Author Note

Correspondence concerning this paper should be addressed to Denis Dumas, Department of

Educational Psychology, University of Georgia: 624 Aderhold Hall, 110 Carlton St., Athens,

GA, 30602. Email: [Link]@[Link].


GUIDELINES FOR AI TEST-SCORING 2

Ten Guidelines for Scoring Psychological Assessments Using Artificial Intelligence

Man makes machines to man the machines that make the machines

-Pete Townshend (1989)

As Pete Townshend (of the legendary rock band The Who) mused in the epigraph above,

modern history is largely defined by a continual increase in the mechanized automation of

human work (e.g., Autor et al., 2003). This process of automation has greatly accelerated and

expanded because of the rise of generative Artificial Intelligence (AI) systems such as Large-

Language Models and Vision Transformer Models. Through the large-scale statistical modeling

of input-data from humans (e.g., text, classifications, quantities, images), these AI models can be

trained to approximate or mimic those humans, thereby producing AI-generated outputs that

have similar properties to their human-generated training inputs (Sejnowksi, 2024). Because of

this, AI now has the capacity to automate tasks that, up until very recently, were the auspices

solely of humans.

One area of novel automation in which AI has already shown great promise is the scoring

of psychological assessments (e.g., Kjell et al., 2024; Organisciak et al., 2023; Yavuz et al.,

2024). For decades, psychologists have heavily relied upon some assessment-types (e.g., Likert

style questionnaires in mental health, or multiple-choice tests in education) that, because of their

closed-ended and well-structured format, could be scored rapidly in a straightforward and

automatic way. In contrast, assessments with a more open-ended or ill-structured item format

(e.g., essays, narratives, drawings) produce data that is not initially numeric, and therefore have

historically relied on human judgments to quantify the responses. Because these open-ended

assessments could not be scored as rapidly and easily as their closed-ended competitors, they

were more difficult and expensive to administer, and therefore much less utilized. However, if
GUIDELINES FOR AI TEST-SCORING 3

datasets of human-judged responses exist for a given open-ended assessment, then AI systems

can now be trained to mimic those human-judges, thereby automating the process of scoring.

We are in the middle of an immensely interesting revolution in psychological assessment.

If participants’ open-ended writing or drawing can be reliably judged and quantified instantly

using AI, this measurement process may one day eclipse close-ended Likert or multiple-choice

assessments in popularity. This shift is relevant across areas of assessment spanning from

education through mental health, as well as both summative and formative assessment purposes.

But just because the AI-based scoring of psychological assessments is now technically possible

and seems to confer advantages vis-à-vis the efficiency of current assessment practices, does not

mean it is by-definition a good thing. On the contrary, irresponsible or unethical applications of

AI to psychological assessments may have the potential to do untold harm if the quantities they

produce are unreliable or invalid.

For this reason, some in the field of psychological assessment (e.g., Burstein et al., 2024;

Dorsey & Michaels, 2022; Ho, 2024; Tay et al., 2022) have begun to consider norms,

recommendations, or standards for how the field can successfully import AI methodology to

automate the scoring of open-ended assessments, while carefully maintaining the integrity of our

field, and especially protecting the rights of the participants we serve. In the current editorial, we

build on this past work, as well as our own research and experience with AI-based assessment, to

forward ten key guidelines that we see as critically important as the field moves toward an

increasingly AI-automated assessment practice.

<Insert Table 1 about here>


GUIDELINES FOR AI TEST-SCORING 4

Do: Disclose the Use of AI to Participants Before They Are Assessed

Before a participant responds to an assessment, ethical practice is to provide them with

key pieces of information about how their responses will be used, including some explanation of

the process of quantification and scoring. Disclosing the usage of AI to quantify open-ended

responses is analogous to the desirable (although not necessarily common) practice of providing

participants with a rubric or checklist that will be used to rate their responses prior to responding;

it allows all participants to be on the same page about how their responses are being judged, and

why. Of course, supervised AI models that are trained to mimic or re-create past human

judgments might not be able to explain precisely how they come up with their ratings (Joyce et

al., 2023), which potentially undermines the transparency of the assessment scoring process. In

such cases, best practice might be to provide participants with an example of a response that was

rated highly on the attribute being assessed by the AI, so participants can see for themselves

what is expected of them.

Do Not: Train AI-Scoring Models on AI-Generated Material

In any field of application, AI models need to be trained on the highest quality data

available. In the ever-burgeoning literature on AI, much has already been written about the

“garbage in, garbage out” phenomenon (e.g., Vidgen & Derczynski, 2020). When supervised AI

models are trained on datasets that have disadvantageous or undesirable properties, they

inevitably reproduce those same problems. In our view, one perhaps so-far overlooked aspect of

this phenomenon is that AI models trained to judge responses to a psychological assessment

must be trained on human judgments, and not prior AI-judgments. Imagine a situation where a

psychology laboratory is collecting human judgments of test responses, with the aim of training

an AI-based automatic scoring system. If one or more of those human judges (perhaps a student
GUIDELINES FOR AI TEST-SCORING 5

research assistant or crowdsourced worker) secretly uses a commercially available AI (e.g.,

ChatGPT) to do their ratings for them, then a circular scenario would be created whereby the AI

model might seem internally valid by available metrics (i.e., it will correlate well with the

judgments it was trained on), when actually it is simply agreeing with another AI.

Do: Account for Uncertainty in AI Response-Ratings

When humans rate responses to open-ended psychological assessments, inevitably there

are some responses for which they are more confident in their ratings, and other responses that

represent edge-cases and for which they have lower confidence in their ratings. AI scoring

models trained on those human ratings will show the same phenomenon when they rate new

responses that they were not trained on. For some responses that are well-represented in their

training dataset, they will be strongly confident, but for other responses they will be less so. One

way to address this issue is to ensure that AI models are trained on large datasets of responses

with several human judgments for each response that meet inter-rater reliability standards

(typically agreement above .80), and hence they are representative and consistent enough to

provide the AI with a reasonably authoritative consensus to mimic (Dumas, 2023a). Another

method is to flag responses for which the AI scoring model has low confidence in its ratings and

bring in an additional human judge (or two) to provide ratings for them. In cases where AI

ratings are highly uncertain, they should not be used as part of the calculation of a participants’

test score, and human ratings should be used instead.

Do Not: Use AI to Score AI-Generated or AI-Assisted Responses

Clearly, when we think about psychological assessment, we tend to assume that a human

is responding to the assessment, and that the intention is to measure some psychological attribute

of that human. But the wide availability of commercial AI models like ChatGPT subverts this
GUIDELINES FOR AI TEST-SCORING 6

assumption. Today, whenever a test is administered away from a watchful proctor, it is a real

possibility that participants might choose to utilize such an AI system to assist them in

formulating their responses as a form of cheating or lazy responding (Susnak & McIntosh, 2024).

Some preliminary evidence suggests that AI scoring systems might prefer, and score more

highly, responses written by AI even in cases where human judges do not (e.g., Tang et al.,

2024). This finding makes sense, given both the AI model that is generating the work and the AI

model that is scoring it might have some training data in common, and may use similar methods

to express (or conversely, observe) psychological attributes in writing or drawing. This finding

also implies that AI scoring might exacerbate the temptation for participants to cheat using AI.

Thus, if AI is scoring an assessment, special care should be taken to ensure that participants do

not use AI to generate their responses.

Do: Predict Human-Scored Criteria to Build a Validity Argument

In order to build the case for the validity of an AI-scored psychological assessment,

evidence must be gathered that its scores predict human-scored criteria. As already discussed,

AI-scoring of psychological assessments opens the door for multiple possible circularly invalid

conclusions that could involve AI-generated material in the training dataset for the scoring

model, or AI-generated material in the responses being scored. Both of these circular situations

would mean that validity evidence internal to the assessment would likely be inflated, because it

would appear that the AI scoring model is capturing the target construct very well, while in

reality it is simply agreeing with another AI. Another similar and equally problematic circular

validation process would be if the correlation between two AI-scored assessments were used as

validity evidence: AI-scored assessments must covary with human-rated assessments in order to

show that their judgments are not insular or idiosyncratic to AI.


GUIDELINES FOR AI TEST-SCORING 7

Do Not: Assume AI-Scoring Models Generalize Across Cultures

Because AI scoring models are capable of quantifying open-ended assessment responses

rapidly and automatically, they open the door for a major scaling-up of those assessment efforts.

New and interesting open-ended assessments with a fully international scope, scored via AI, may

soon be developed. This possibility is exciting, but it also creates the need to ensure the fairness

of the AI-scoring process across cultural groups. As with more typical assessment practices,

which become less and less likely to yield invariant and comparable scores across participants as

the cultural distance between them increases (Dong & Dumas, 2020), the assumption of “one

size fits all” seems unlikely to hold for AI scoring models. Especially if participants come from a

culture or language background that was not represented by the human judgments used to train

the AI, it seems unlikely that those AI-scores could be valid for them. One way to help avoid

these issues in the future is to utilize very wide-ranging and culturally diverse human judgments

to train AI models, in an attempt to produce a cosmopolitan scoring model that can genuinely

handle responses in multiple languages, or from multiple cultural perspectives. Even so,

psychometric methods such as measurement invariance (Meredith, 1993), differential item

functioning (Holland & Wainer, 2012), or regression-based predictive fairness metrics (Dumas et

al., 2023b) should be utilized to address the empirical question of AI-score fairness.

Do: Establish A Clear Protocol for Human Oversight and Auditing

Even the most well-designed and well-validated psychological assessment can never be

said to be completely finished. Over time, test scores can decrease in quality as cultural shifts

make the items less relevant to the target construct, participants become aware of the assessment

structure and how to game it without meaningfully responding, or the target population shifts or

expands to include new participants for whom the test was not originally designed. Any of these
GUIDELINES FOR AI TEST-SCORING 8

changes, and many more, can result in downstream shifts in the psychometric parameters of the

scoring model, decreases in reliability, larger standard errors around scores, and weaker validity

predictions (referred to as drift). When AI is used to score a test, it adds an additional layer of

complexity and possible problems that accrue over time: drift no longer refers only to the

parameters of a psychometric scoring model, but to the parameters of the AI-model used to rate

the responses as well. AI models need to be continually updated, trained, and re-validated on the

best-available human-rated data, and the effect of these changes on any psychometric scoring

model used to aggregate the AI-ratings (and hence the comparability of the scores) should also

be studied.

Do Not: Directly Compare Participant Scores Across AI Models or Versions

Currently, AI models are changing rapidly, with new advances arriving nearly monthly.

This means that proprietary and commercially available AI models that are owned by technology

companies (e.g., OpenAI owns ChatGPT and a variety of other models) are being continually

updated and changed in ways that users of those models cannot control, and completely new

models are coming out regularly. We suspect that, within the area of psychological assessment, it

will be rare for researchers or even testing companies to build their own generative AI model

from the ground up, and instead it is likely that psychological researchers will fine-tune existing

models. Every time the base model is changed by the technology company that controls it, those

changes will inevitably flow downstream and affect any researcher fine-tuned AI model built on

top of it, as well as any psychometric model used to aggregate AI-ratings of participant

responses. Each AI model update will result in some drift in the parameters of the scoring model

for the test and therefore also changes in the distribution of the scores. This means that, after an

AI update, individual participant scores become non-comparable with those from earlier
GUIDELINES FOR AI TEST-SCORING 9

versions. The degree to which those scores have changed due to the AI model update and

resulting psychometric parameter drift should always be an empirical question.

Do: Explain the Use of AI to Participants in the Final Score Report

As psychometricians, it can sometimes be woefully easy to dismiss the final score report

given to test stake-holders (e.g., clinicians, educators, participants, or their parents) as an

appendix on the assessment process. But the score report is the main way in which assessment

stake-holders interpret a score and understand what it might mean for them. For this reason, all

score reports should be designed and written with careful attention paid to stake-holder

comprehension of the testing process and score meaning. Score reports for AI-scored

assessments are certainly no exception to this, and for this reason special care should be taken to

explain the way AI was used to calculate scores in the final report. The emphasis here is to

provide the information that allows the participant to understand that the AI model has been

carefully developed and that it provides the best-available judgments of the psychological

construct being assessed, so that they feel that their score is trustworthy and meaningful to them.

Do Not: Expect Stake-holders to Trust a Black-Box Without Justification

As AI models become more advanced, the processes through which they make decisions

and generate output also become increasingly more opaque, resembling the classic metaphor of

the “black-box” (Bearman & Ajjawi, 2023). At the same time, as AI becomes more and more

capable of general tasks, individuals across disciplines might be tempted simply to trust its

output, without rigorous human justification for its validity. Especially in higher-stakes

psychological or educational contexts, stake-holders should never be expected simply to take AI-

based test scores at face value, without specific evidence forwarded to support the validity of the

scoring process. AI is a very useful tool to make the process of assessment scoring more
GUIDELINES FOR AI TEST-SCORING 10

efficient, but it should never be a replacement for careful thought on the part of a human

psychologist or psychometrician.

Conclusion

We are cautiously optimistic about the potential future role of AI in scoring

psychological assessments. Clearly, the field’s reliance on highly efficient closed-ended items

(e.g., multiple choice items) has produced some limitations in our capacity to measure the

complex and nuanced psychological constructs which we, as psychologists, are interested in.

But, without a sharp (and unlikely) increase in the resources available for psychological science,

the time- and cost-efficiency of our assessment practices is likely to be an important aspect of

our work going forward. For this reason, we see the use of AI as one potentially important

methodology for expanding the scope of psychological assessment. At EJPA, we welcome both

methodological and empirical submissions addressing this topic. If AI is applied to the scoring of

psychological assessments in a careful, responsible, and ethical way, it may be that it will help

with what we see as the ultimate goal of psychology: to contribute to human flourishing through

a better understanding of the human mind.


GUIDELINES FOR AI TEST-SCORING 11

References

Autor, D. H., Levy, F., & Murnane, R. J. (2003). The skill content of recent technological change: An

empirical exploration. The Quarterly Journal of Economics, 118(4), 1279–1333.

[Link]

Bearman, M., & Ajjawi, R. (2023). Learning to work with the black box: Pedagogy for a world with

artificial intelligence. British Journal of Educational Technology, 54(5), 1160–1173.

[Link]

Burstein, J., LaFlair, G. T., Yancey, K., Von Davier, A. A., & Dotan, R. (2024). Responsible AI for

Test Equity and Quality: The Duolingo English Test as a Case Study (arXiv:2409.07476). arXiv.

[Link]

Dong, Y., & Dumas, D. (2020). Are personality measures valid for different populations? A

systematic review of measurement invariance across cultures, gender, and age. Personality and

Individual Differences, 160, 109956. [Link]

Dorsey, D. W., & Michaels, H. R. (2022). Validity arguments meet artificial intelligence in innovative

educational assessment. Journal of Educational Measurement, 59(3), 267–271.

[Link]

Dumas, D., Acar, S., Berthiaume, K., Organisciak, P., Eby, D., Grajzel, K., Vlaamster, T., Newman,

M., & Carrera, M. (2023a). What makes children’s responses to creativity assessments difficult

to judge reliably? The Journal of Creative Behavior, 57(3), 419–438.

[Link]

Dumas, D., Dong, Y., & McNeish, D. (2023b). How fair is my test? A ratio coefficient to help

represent consequential validity. European Journal of Psychological Assessment, 39(6), 416–

423. [Link]
GUIDELINES FOR AI TEST-SCORING 12

Ho, A. D. (2024). Artificial intelligence and educational measurement: Opportunities and threats.

Journal of Educational and Behavioral Statistics, 49(5), 715–722.

[Link]

Holland, P. W., & Wainer, H. (2012). Differential item functioning. Routledge.

Joyce, D. W., Kormilitzin, A., Smith, K. A., & Cipriani, A. (2023). Explainable artificial intelligence

for mental health through transparency and interpretability for understandability. Npj Digital

Medicine, 6(1), 1–7. [Link]

Kjell, O. N. E., Kjell, K., & Schwartz, H. A. (2024). Beyond rating scales: With targeted evaluation,

large language models are poised for psychological assessment. Psychiatry Research, 333,

115667. [Link]

Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance.

Psychometrika, 58(4), 525–543. [Link]

Organisciak, P., Acar, S., Dumas, D., & Berthiaume, K. (2023). Beyond semantic distance:

Automated scoring of divergent thinking greatly improves with large language models. Thinking

Skills and Creativity, 49, 101356.

Sejnowski, T. J. (2024). ChatGPT and the future of AI: The deep language revolution. The MIT

Press.

Susnjak, T., & McIntosh, T. R. (2024). ChatGPT: The end of online exam integrity? Education

Sciences, 14(6), Article 6. [Link]

Tang, M., Hofreiter, S., Werner, C. H., Zielinska, A., & Karwowski, M. (2024). “Who” is the best

creative thinking partner? An experimental investigation of human-human, human-internet, and

human-AI co-creation. Journal of Creative Behavior. [Link]


GUIDELINES FOR AI TEST-SCORING 13

Tay, L., Woo, S. E., Hickman, L., Booth, B. M., & D’Mello, S. (2022). A conceptual framework for

investigating and mitigating machine-learning measurement bias (MLMB) in psychological

assessment. Advances in Methods and Practices in Psychological Science, 5(1),

25152459211061337. [Link]

Vidgen, B., & Derczynski, L. (2020). Directions in abusive language training data, a systematic

review: Garbage in, garbage out. PLOS ONE, 15(12), e0243300.

[Link]

Yavuz, F., Çelik, Ö., & Yavaş Çelik, G. (2025). Utilizing large language models for EFL essay

grading: An examination of reliability and validity in rubric-based assessments. British Journal

of Educational Technology, 56(1), 150–166. [Link]


GUIDELINES FOR AI TEST-SCORING 14

Table 1.

Ten Guidelines for AI Scoring of Psychological Assessments

Do: Do Not:
1. Disclose the Use of AI to Participants 2. Train AI-Scoring Models on AI-
Before They Are Assessed Generated Material
3. Account for Uncertainty in AI Response- 4. Use AI to Score AI-Generated or AI-
Ratings Assisted Responses
5. Predict Human-Scored Criteria to Build a 6. Assume AI-Scoring Models Generalize
Validity Argument Across Cultures
7. Establish A Clear Protocol for Human 8. Directly Compare Participant Scores
Oversight and Auditing Across AI Models or Versions
9. Explain the Use of AI to Participants in 10. Expect Stake-holders to Trust a Black-
the Final Score Report Box Without Justification

View publication stats

Common questions

Powered by AI

AI models might preferentially score AI-generated text highly due to shared characteristics in training data, potentially skewing the validity of genuine assessments. This bias undermines the purpose of evaluating human psychological attributes, necessitating rigorous training on diverse, human-generated examples and discouraging AI assistance in respondent-generated content .

AI faces challenges in scoring open-ended assessments due to the non-numeric nature of these data, which historically relied on human judgment. To address this, AI models must be trained on large, diverse datasets with multiple human judgments to ensure inter-rater reliability. Additionally, incorporating a system to flag responses with low AI confidence can allow human judges to intervene, ensuring scoring accuracy .

Scoring AI-generated responses with AI can lead to misleading validity indicators, as AI models trained on similar data can appear internally valid by simply agreeing with each other. This can diminish the assessment's measure of the actual psychological attribute intended for the human participant. Therefore, it is crucial to prevent AI from scoring AI-generated responses to avoid circular validation problems .

Continuous validation and re-training are essential to mitigate drift, which occurs as cultural shifts and model updates change the context or application of the AI scoring model. Regularly updating AI with current, high-quality human-rated data ensures that it retains reliability and validity over time, preventing decay in assessment quality .

Human oversight should be prioritized in situations where AI models show low confidence in scoring or face edge cases not well represented in the training data. If AI ratings are highly uncertain, human judges should intervene to maintain the integrity and accuracy of the assessment scores .

Cultural diversity is critical because AI models trained on non-representative human judgement data may not yield valid scores across different cultural backgrounds. Ensuring fairness involves using diverse training data, performing measurement invariance analyses, and applying differential item functioning techniques to confirm that AI assessments are culturally fair and unbiased .

The principle 'garbage in, garbage out' implies that AI models will reproduce undesirable traits if trained on low-quality data. For psychological assessments, this means AI must be trained on high-quality human judgment data to ensure accurate scoring. Using AI-generated data can create a circular validity issue, where AI agrees with another AI rather than accurately assessing responses .

Different AI versions might yield non-comparable scores due to varying training data and algorithms, compromising assessment consistency. Mitigation strategies include standardizing training protocols, frequent recalibration with benchmark datasets, and ensuring rigorous version control to maintain comparability and reliability of scores across updates .

Promoting trust involves explaining AI use clearly in final score reports and providing evidence of scoring validity. Transparency through the provision of examples and the rationale behind AI decisions helps demystify the process. Additionally, stakeholders should be informed of AI's ability to meet high inter-rater reliability standards to show that AI decision-making is accountable and reliable .

Using AI for scoring without disclosure may violate ethical standards by undermining transparency and informed consent. Participants have the right to know how their assessments are scored, facilitating trust and understanding of the process. Ethically, it is crucial to disclose AI use to maintain the integrity of psychological evaluations and respect participant autonomy .

You might also like