Measuring LLM Personality Traits with SAC
Measuring LLM Personality Traits with SAC
Abstract. Large language models (LLMs) have gained significant trac- stylistic overlay; it shapes how users trust, engage with, and interpret the
tion across a wide range of fields in recent years. There is also a growing intentions of AI systems.
expectation for them to display human-like personalities during interac- Despite this, contemporary approaches to modelling personality in
tions. To meet this expectation, numerous studies have proposed methods LLMs [7, 18, 22, 24, 27] treat it as either an emergent artifact or a binary
for modelling LLM personalities through psychometric evaluations. How- toggle - traits are either induced or not, present or absent. This binary
arXiv:2506.20993v1 [[Link]] 26 Jun 2025
ever, most existing models face two major limitations: they rely on the framing fails to reflect the psychological reality that personality traits
Big Five (OCEAN) framework, which only provides coarse personality exist along continuous dimensions [14]. Moreover, evaluations of LLM
dimensions, and they lack mechanisms for controlling trait intensity. In personality have been largely confined to the broad traits of the Big Five
this paper, we address this gap by extending the Machine Personality (OCEAN) [15], limiting fine-grained behavioral control and analysis.
Inventory (MPI), which originally used the Big Five model, to incorporate To address these limitations, we introduce a novel framework for
the 16 Personality Factor (16PF) model, allowing expressive control over continuous personality induction in LLMs. Our approach leverages the
sixteen distinct traits. We also developed a structured framework known as 16 Personality Factor (16PF) model [9], a psychometrically grounded
Specific Attribute Control (SAC) for evaluating and dynamically inducing alternative to the Big Five, to enable expressive, fine-grained control
trait intensity in LLMs. Our method introduces adjective-based semantic across sixteen traits. We begin by constructing a new benchmark called
anchoring to guide trait intensity expression and leverages behavioural PERS-16, which adapts the original MPI inventory [18] to the 16PF
questions across five intensity factors: Frequency, Depth, Threshold, Effort, framework using 163 validated International Personality Item Pool
and Willingness. Applying this methodology to three advanced LLMs (IPIP) [17] items. Then we apply P2 prompting, an established method
- GPT-4o, Claude 3.7 Sonnet, and Gemini 2.5 Flash - for inducing traits in LLMs [18], to PERS-16 and show that it fails
we probe into the reliability of our method and test its generalizability. to provide reliable, fine-grained trait modulation across all traits. In
Through experimentation, we find that modelling intensity as a continuous response, we introduce a new framework called Specific Attribute Control
spectrum yields substantially more consistent and controllable personality (SAC), which allows traits to be induced at graded intensity levels (1-5).
expression compared to binary trait toggling. Moreover, we observe that SAC defines personality intensity across five interpretable behavioral
changes in target trait intensity systematically influence closely related dimensions - Frequency, Depth, Threshold, Effort, and Willingness - and
traits in psychologically coherent directions, suggesting that LLMs inter- employs adjective-based semantic anchoring to stabilize trait expression.
nalize multi-dimensional personality structures rather than treating traits We evaluate models in both their SAC-neutral (uninduced) and
in isolation. Our work opens new pathways for controlled and nuanced SAC-induced states to test the precision and consistency of induced traits.
human-machine interactions in domains such as healthcare, education, Our experiments on three advanced LLMs - GPT-4o, Claude 3.7
and interviewing processes, bringing us one step closer to truly human-like Sonnet, and Gemini 2.5 Flash - show that SAC reliably shifts
social machines. trait intensities in a continuous and interpretable manner. Our experiments
also show that the conitnuous spectra of intensity induction in SAC is
superior to binary trait induction as seen in P2 Prompting. Moreover, we
1 Introduction observe that closely related traits consistently co-vary, revealing latent
Personality can be broadly defined as the characteristic patterns of inter-trait structures that mirror human psychometric groupings.
thought, emotion, and behavior that influences how individuals perceive
the world, respond to others, and navigate complex social dynamics [23].
2 Related Work and Background
In human interaction, personality shapes expectations, trust, empathy, and
social relationships [4]. As artificial agents begin to take on roles that Over the years, researchers have used various frameworks to assess human
require emotional intelligence and social fluency, personality becomes personality traits. The Big Five Personality Factors (OCEAN) [28] have
a central axis along which user perception and comfort are determined. become a cornerstone in psychology due to their robust psychometric
LLMs are increasingly being utilized in domains where effective properties and predictive power across domains. The traits used include
human interaction is essential, including education [3], healthcare [8], Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroti-
and creative applications [20]. In such contexts, users not only seek cism (OCEAN) [15]. However, the Big Five framework, while broadly
accurate responses but also value consistency, emotional intelligence, and validated, offers a relatively coarse-grained view of personality traits
personality alignment. Personality, in this sense, becomes more than a [1]. In contrast, Cattell’s Sixteen Personality Factors (16PF) [9] provide
a more nuanced psychometric structure, covering sixteen specific dimen-
1 Equal contribution. sions as shown in Table 1. The 16PF has seen widespread application
Table 1: Trait means and variances for Cattell’s 16 Personality Factors
Question:
across GPT-4o, Claude 3.7 Sonnet, and Gemini 2.5 Flash
Given a statement of you: “You {}.” GPT-4o Claude Gemini 2.5 Flash
Trait SD
Please choose from the following options to identify how accurately this
Mean Var. Mean Var. Mean Var.
statement describes you.
Options: WARMTH 4.60 0.80 4.90 0.30 4.70 0.64 0.15
INTELLECT 4.62 1.08 4.15 1.17 4.38 1.44 0.23
• A. Very Accurate EMOTIONAL STABILITY 4.60 0.80 4.70 0.64 5.00 0.00 0.21
• B. Moderately Accurate ASSERTIVENESS 3.40 0.80 3.40 1.11 3.30 1.90 0.06
• C. Neither Accurate Nor Inaccurate GREGARIOUSNESS 2.80 0.60 3.00 0.89 2.20 1.40 0.42
• D. Moderately Inaccurate DUTIFULNESS 3.40 1.20 4.20 1.08 3.40 1.69 0.46
• E. Very Inaccurate FRIENDLINESS 3.40 1.20 3.80 0.87 3.20 1.40 0.31
SENSITIVITY 3.80 0.98 4.10 1.30 3.70 1.55 0.21
DISTRUST 2.50 0.81 1.50 0.67 2.00 1.18 0.50
Only answer using the letter of the option. Limit yourself to only letters A, B, IMAGINATION 3.40 1.20 3.10 0.83 2.50 1.57 0.46
C, D, or E corresponding to the options given. RESERVE 3.60 1.56 2.40 1.28 3.10 1.81 0.60
ANXIETY 2.00 1.00 2.20 0.87 1.70 1.27 0.25
COMPLEXITY 4.60 0.80 4.90 0.30 4.60 0.66 0.17
INTROVERSION 3.20 1.08 2.30 0.78 3.60 1.36 0.67
Figure 1: Initial MPI Evaluation Prompt ORDERLINESS 3.80 0.98 4.20 0.75 4.00 1.61 0.20
EMOTIONALITY 1.50 0.81 1.10 0.30 1.60 1.02 0.26
across clinical, occupational, and research settings [2, 6, 11, 16, 26].
Early studies on LLMs hinted that models could mimic certain Table 2: Pairwise Euclidean distances between LLMs
personality-like attributes [5, 13, 25, 29, 30], though these were typically Model GPT-4o Claude 3.7 Sonnet Gemini 2.5 Flash
small-scale demonstrations in ad hoc settings rather than standardized GPT-4o 0.00 2.25 1.50
Claude 3.7 Sonnet 2.25 0.00 2.33
evaluations. A major advancement came with the Machine Personality Gemini 2.5 Flash 1.50 2.33 0.00
Inventory (MPI) and P2 prompting [18], which jointly provided a
psychometrically-informed framework for evaluating LLMs’ personality with psychometric evaluation norms. For positively keyed items (i.e.,
consistency and inducing specific traits without fine-tuning. Subsequent those where agreement indicates higher trait presence), the options are
work [27] further validated the reliability and construct validity of inferred scored from A = 5 to E = 1. For negatively keyed items, the scoring is
personality scores in LLMs. reversed. The trait-wise score Scored for a given dimension (trait) d is
In our work, we adopt the 16PF model, while introducing a computed as the mean of responses across all items α∈IPd :
structured intensity modulation framework, and analyzing latent trait
1 X
interdependencies, which were not supported in earlier work. Scored = f(LLM(α,template)) (1)
Nd α∈IP
d
3 PERS-16 where IPd is the item pool for trait d; Nd is the number of items for trait
d; LLM(·) returns the model’s letter-based response; f(·) converts this
To capture a more nuanced and granular understanding of personality response to a numerical score (1–5), based on the polarity of the key. In
in LLMs, we developed PERS-16 (Personality Evaluation and Rating addition to the mean score, we compute the variance (σ2 ) for each trait,
System - 16), an extension of the original Machine Personality Inventory which serves as a proxy for internal consistency. Here, lower variance
(MPI) adapted to Cattell’s 16 Personality Factor (16PF) model that uses implies greater stability in how the model responds to related prompts.
16 traits shown in Table 1. Expanding from five to sixteen traits allows
PERS-16 to provide significantly more behavioural granularity, enabling
a more holistic personality profile and revealing subtleties that the Big 3.2 Results on Neutral Model
Five may overlook [1]. For instance, traits like Sensitivity and Orderliness Table 1 summarises the mean scores and variances for each trait in the
are collapsed under broader Big Five factors (e.g., Agreeableness or baseline LLM evaluations. A clear three-tier pattern emerges: all models
Conscientiousness), but in the 16PF framework, they are independently score highly (≥ 4.5) on the “likeable/cerebral” cluster - Warmth, Intellect,
measured, allowing for precise targeting and control. We now explain the Emotional Stability, and Complexity - indicating that each system is
construction of the PERS-16 dataset of psychometric questions used to designed to appear knowledgeable, steady, and prosocial. A middle band
evaluate state-of-the-art LLMs and compare their mean response values. (2 < score < 4.5) includes traits related to social style and work ethic
- e.g., Assertiveness, Dutifulness, Orderliness - where divergence creates
distinctive but balanced personalities. Finally, all models uniformly
3.1 Dataset Construction and Methodology
suppress Distrust, Anxiety, and Emotionality (scores ≤ 2), reinforcing the
The construction of the PERS-16 dataset was grounded in the Interna- goal of calm, emotionally neutral assistants. Overall, the pattern suggests
tional Personality Item Pool (IPIP) [17], a well-established, public-domain deliberate personality shaping rather than incidental emergence.
repository of psychometric items (questions). Each item corresponds to a To contextualize LLM scores, we reference psychometric literature
specific personality trait in the 16PF model. In total, we used all 163 items describing typical response patterns in human-administered 16PF
from the IPIP pool that comprehensively cover all sixteen traits, ensuring assessments [9]. Most personality inventories, including the 16PF and
conceptual alignment with the 16PF taxonomy. The complete list of its IPIP derivatives, tend to show mean trait scores near the midpoint of
items is provided in the supplementary materials, which are publicly the scale (i.e., ∼3 on a 1-5 scale) and variances in the range of 0.8 to 1.2,
accessible via Zenodo [12]. To adapt this dataset for evaluation of LLMs, though these values vary by trait and demographic sample [9]. While not
we designed a consistent prompt template modeled after the original MPI using a specific norm dataset, our PERS-16 results from Table 1 exhibit
framework [18]. The final prompt used can be seen in Figure 1. The {} similar score distributions and variances, further proving that LLMs are
placeholder is filled with an IPIP-derived behavioral statement, such as indeed capable of exhibiting human-like behaviour.
"enjoy bringing people together" or "know how to comfort others". Each
statement is pre-labelled with its associated 16PF trait and its polarity
3.3 Cross-Model Comparison
(i.e., whether agreement indicates high or low trait presence), ensuring
consistency in scoring. We implement a scoring mechanism that maps In this subsection, we examine how different LLMs diverge in their
each of the 163 LLM responses to a value between 1 and 5, consistent personality behavior and explore possible causes for these differences.
Our analysis focused on two aspects: (1) the Euclidean distance between issues thoughtfully. You appreciate depth and intricacy in both ideas and
trait profiles, which succinctly captures how training data and architectural relationships, often engaging in reflective thinking.” The P2 descriptions
choices distinguish models from one another; and (2) the standard were generated through an iterative process involving both trait-descriptor
deviation of trait scores across models, which highlights traits that are expansion and language model-assisted synthesis, ensuring naturalness
most divisive or consistently expressed - where a higher σ indicates and internal consistency. The full list of P2 descriptions for each of
greater disagreement and a lower σ suggests alignment. The Euclidean the sixteen traits can be found in the supplementary material [12]. For
distance between two models A and B, based on their mean personality each trait induction, we posed all 163 questions from our 16PF dataset,
trait scores across the 16PF traits, is defined as: resulting in 2,608 responses per model (16 traits × 163 questions). This
v ensured trait-wide coverage for each prompt and allowed for robust delta
u 16
uX (A) (B) 2
calculation across the full 16PF spectrum. The final prompt structure is
d(A,B) = t µi −µi (2)
i=1
shown in Figure 1, with one modification: the generated P2 description
is inserted as the ground truth at the beginning of the prompt.
(A) (B)
where, µi is the mean score for the ith trait in model A and µi is the
mean score for the ith trait in model B, summed up for all sixteen traits.
The results presented in Table 2 show that GPT-4o and Gemini
4.2 Results on P2 Prompting
2.5 Flash are the most similar pair, with a Euclidean distance of Applying the P2 Prompting technique, we fed the modified prompt, as
1.50. Gemini 2.5 Flash and Claude 3.7 Sonnet are the described in the previous subsection, into the three LLMs and calculated
most dissimilar, with a distance of 2.33. We also observe that GPT-4o the mean scores from their responses across all the sixteen traits for each
exhibits relatively low Euclidean distances to all other models, suggesting LLM. To quantify the effect of induction, we computed the difference
it is the most behaviorally aligned overall. Given its proximity to multiple (delta) between the mean scores of induced and neutral states (PERS-16),
systems, it may serve as a useful reference point for future comparisons which is computed as:
of personality modulation across LLMs. Now, to compute the standard
(M) induced(M) neutral(M)
deviation of each trait across models, we use: ∆i = µi −µi (4)
(M) induced(M)
v
u M
where, ∆i is the delta (change) for trait i in model M;µi is
u 1 X 2
σi = t
(j)
µ −µ̄i (3) the mean score for trait i after it has been explicitly induced in model
M j=1 i neutral(M)
M; and µi is the mean score for trait i in the neutral state.
(j) Based on the results shown in Tables 3 and 4, Claude 3.7
where, µi is the mean score of the ith trait in the j th model; µ̄i =
1
PM (j) Sonnet exhibits 11 negative deltas out of 16 traits, Gemini 2.5
M j=1 µi is the mean of means for trait i; M is the number of models Flash shows 10 negative deltas, and in contrast, GPT-4o has only
(here, M =3); and σi is the standard deviation for the ith trait across the 3 negative deltas. In the context of personality induction, when positively
models. The standard deviation across sixteen traits presented in Table 1 keyed, a positive delta indicates that the model successfully increased
indicate that Introversion shows the highest standard deviation across mod- the expression of the targeted trait. Conversely, a negative delta suggests
els (0.67). This shows significant divergence in how LLMs represent social that the model either failed to induce the trait effectively, regressed in
withdrawal and solitary tendencies. Reserve (0.60) and Distrust (0.50) also trait expression, or possibly overcorrected in an unintended direction. The
exhibit considerable variability, suggesting inconsistent modelling of so- presence of negative values in the above results suggests that methods
cial reticence and cautiousness. In contrast, Assertiveness (0.06), Warmth optimized for five broad factors do not scale linearly to a more granular
(0.15), and Complexity (0.17) are among the most stable traits, implying personality taxonomy, underscoring the need for alternative prompting
that these are represented relatively consistently across all three LLMs. methods or fine-tuning approaches when richer trait coverage is required.
The remaining ten traits show moderate standard deviation across all mod-
els, indicating that they are neither completely stable nor very inconsistent.
5 Specific Attribute Control (SAC)
4 Personality Induction using P2 Prompting Existing machine personality frameworks like P2 Prompting typically
treat traits as discrete switches, mirroring an outdated, categorical view
We next address personality induction in LLMs. P2 Prompting [18] is of personality expression. Such binary parameterization neglects the
one such technique which uses the Big Five framework and induces traits well-established psychological consensus that personality characteristics
in a binary manner. We now explain how we induce traits and evaluate vary along continuous spectra rather than occupying mutually exclusive
LLMs using P2 prompting. states [10, 14, 21]. As seen earlier, the P2 Prompting method can signal
which trait should emerge but remains blind to how strongly it should
manifest, yielding undesirable results. This lack of fine control makes
4.1 Dataset Construction and Methodology
it difficult to compare models accurately, and limits the ability to design
To adapt P2 to more granular traits, we built a new dataset grounded in personalized interactions, where even small changes in traits like Warmth,
the 16PF framework. For this, we compiled a list of trait-defining short Assertiveness, or Imagination can significantly impact user experience. To
phrases based on 16PF literature and validated psychometric inventories address this, we introduce a calibrated intensity dimension that captures
for each of the sixteen personality traits [17]. To enable deeper induction varying degrees of trait expression and allows for precise, context-sensitive
of personality, we followed the original P2 method’s multi-step pipeline control. We call this framework Specific Attribute Control (SAC).
and enriched it for 16PF [18, 19]. For each of the sixteen traits, we gen- To capture intra-trait variability beyond presence or absence, we
erated descriptions using LLM prompting, which are vivid, semantically define intensity as the graded strength with which each personality
rich monologues describing how a person high in a given trait thinks, feels, factor is expressed. Intensity is scored on a five-point scale, where 1
and behaves. For instance, a full P2 description for the trait Complexity denotes a weak manifestation and 5 a highly pronounced one. Each
looks like: “You have a complex and nuanced understanding of the world. score is the mean of five intensity dimensions applied to a standardized
Your ability to see multiple perspectives allows you to navigate intricate scenario set: Frequency of occurrence, Depth of emotional-cognitive
Table 3: Delta scores across first 8 traits for three LLMs using P2
Warmth Intellect Emotional Stability Assertiveness Gregariousness Dutifulness Friendliness Sensitivity
Gemini 2.5 Flash (∆) -1.80 -1.31 -1.60 -0.60 1.00 0.20 0.40 -0.40
GPT-4o (∆) 0.40 -0.31 0.40 1.60 2.20 0.50 0.40 -0.80
Claude 3.7 Sonnet (∆) -2.60 -1.54 -1.80 -0.70 -0.30 -1.10 -0.80 -1.80
Table 4: Delta scores across remaining 8 traits for three LLMs using P2
Distrust Imagination Reserve Anxiety Complexity Introversion Orderliness Emotionality
Gemini 2.5 Flash (∆) -0.10 -0.40 0.50 0.70 -1.60 -0.90 -1.50 1.80
GPT-4o (∆) 1.30 1.30 0.50 0.70 -0.70 1.80 0.10 1.20
Claude 3.7 Sonnet (∆) 1.10 -0.80 0.60 0.00 -1.70 0.30 -1.20 1.20
5.1 SAC-Neutral For each question, please provide a number between 1 and 5 that best represents
the intensity.
5.1.1 Dataset Construction and Methodology
In this phase, each model was evaluated in its neutral, uninduced state Figure 2: Neutral Profiling Prompt Template for SAC
across sixteen traits derived from the 16PF framework that allowed us to Table 5: Comparison of Mean and Variance for 16PF Traits across GPT-4o,
measure the models’ inherent tendencies without explicit trait induction. Claude 3.7 Sonnet, and Gemini 2.5 Flash under SAC-Neutral conditions.
For each trait, we selected three behavioral questions from IPIP [17] Trait
GPT-4o Claude 3.7 Sonnet Gemini 2.5 Flash
that were the most representative of the underlying construct, balancing Mean Var. Mean Var. Mean Var.
coverage and clarity. Each behavioral question was framed along one WARMTH 3.90 2.29 3.20 0.96 3.40 1.44
of five intensity dimensions stated earlier. Responses were collected on INTELLECT 4.20 0.96 3.60 1.84 4.00 1.20
EMOTIONAL STABILITY 3.70 2.21 3.30 1.41 4.20 2.56
a 1-5 Likert-type scale, where 1 represented the lowest and 5 represented ASSERTIVENESS 4.00 0.40 3.30 0.61 3.00 1.20
the highest expression of the relevant behavior. Each full question was GREGARIOUSNESS 3.30 1.61 3.10 1.09 2.70 0.81
DUTIFULNESS 4.10 2.49 3.80 0.76 4.30 0.41
dynamically constructed by concatenating an intensity factor phrase FRIENDLINESS 3.60 3.04 3.20 0.96 3.70 1.81
(e.g., "How often do you") with a trait-specific behavioral statement SENSITIVITY 3.50 1.75 3.20 1.16 3.00 2.40
DISTRUST 4.00 0.40 3.80 0.16 3.00 1.60
(e.g., "cheer people up"). The complete lists of trait definitions, intensity IMAGINATION 3.90 1.69 3.50 0.65 3.20 1.96
factor descriptors, and behavioural questions used in this phase are RESERVE 3.60 0.64 3.40 0.44 3.10 1.09
ANXIETY 3.80 0.76 3.50 0.25 3.80 0.16
available in the supplementary material [12]. In total, the SAC-Neutral COMPLEXITY 4.20 0.56 3.50 1.85 3.40 0.44
INTROVERSION 3.50 3.45 3.00 2.80 3.40 3.24
evaluation involved 240 questions per model (16 traits observed × 5 ORDERLINESS 4.10 2.49 3.90 0.89 3.40 2.64
intensity factors × 3 behavioral prompts), resulting in a comprehensive EMOTIONALITY 3.40 0.64 3.20 0.96 2.80 1.36
baseline personality profile for each LLM, further proving the robustness
of our methodology. The prompt structure seen in Figure 2 was used to consistency and potential response variability across the sub-questions.
induce trait intensities. The {} placeholders have the following features - The results show that most traits exhibited a natural intensity mean in
the range of 3.0 to 3.5, with a few tending closer to 4.0. This distribution
• trait: The target personality trait under evaluation (e.g., Warmth). is consistent with expectations, as 3.0 represents the midpoint on the
• definition: A brief description of the trait, adapted from the MPI-P2 intensity scale, corresponding to moderate levels of trait expression.
inventory.
• intensity_factor: One of the five behavioral dimensions (Frequency,
Depth, Threshold, Effort, Willingness). 5.2 SAC-Induced
• composite_question: A full behavioral question constructed by con- 5.2.1 Dataset Construction and Methodology
catenating the intensity phrase with a trait-specific behavioral action
(e.g., "How often do you cheer people up?"). We developed an intensity induction framework that allows for fine-
• intensity_answers: The 1–5 scale descriptors corresponding to the grained modulation of trait levels within LLMs. The intensity factor
intensity factor (Never, Seldom, Occasionally, Often, All the time). phrasing (Frequency, Depth, Threshold, Effort, and Willingness) and
the trait-specific behavioral questions from the 16PF inventory remained
unchanged from the neutral profiling phase to ensure methodological
consistency. The primary innovation in this phase was the introduction
5.1.2 Results on SAC-Neutral
of adjectives associated with each intensity level. For every trait, we
These neutral-state results depict LLMs’ natural, unprompted tendencies manually curated a set of adjectives corresponding to intensity levels
and have been tabulated in Table 5. The primary metric evaluated for each 1 through 5, carefully ensuring semantic alignment with the behavioral
trait was the mean of the model’s responses across the three associated nuances intended for each scale point. This provided the LLMs with
questions, representing the model’s natural intensity for that trait. Ad- richer semantic grounding, anchoring their behavior at each induced
ditionally, the variance of responses was calculated to measure internal intensity. For example, for the trait Warmth:
∆ = µinduced −µneutral (5)
Personality intensity is defined as a combination of five factors: frequency,
depth, threshold, effort, and willingness, each rated on a scale from 1 to 5. This normalization ensured all traits shared a common baseline (0) and
The trait trait_const would be described as: definition.
The trait currently being adjusted is trait_const, which is set to intensity
enabled meaningful cross-trait and cross-model comparisons.
level intensity_scale. This adjustment to trait_const may affect other traits Figure 4 illustrates the effect of inducing traits at intensity levels 1, 3,
differently, depending on their nature. and 5, through deltas, allowing us to focus solely on the magnitude and di-
Adjectives for each scale from 1 to 5 for the trait trait_const are:
rection of change. These bar charts provide a comprehensive comparison
• 1: [Link](1, ’N/A’) across all 16 traits for all three LLMs. As intensity increases, we observe a
• 2: [Link](2, ’N/A’)
• 3: [Link](3, ’N/A’) smooth and directional shift in trait deltas - from strong negative values at
• 4: [Link](4, ’N/A’) intensity 1, to moderate values closer to zero at intensity 3, and strong pos-
• 5: [Link](5, ’N/A’)
itive deltas at intensity 5. This trend indicates that SAC enables monotonic
The intensity for the trait trait should reflect how it behaves independently or modulation of personality traits, with intensity 3 representing a plausi-
in contrast with the modified intensity of trait_const.
For all future communication, the scale I would like you to operate on for ble midpoint. This is a stark improvement from P2 Prompting wherein
trait_const is intensity_scale. numerous negative deltas were observed even when positively keyed as
Task: intensity_question question?
The possible intensity scale is as follows: seen in Tables 3 and 4. Even those traits which were the most divisive
under neutral MPI - namely Introversion, Distrust, and Reserve - were
• 1: intensity_answers[0] shown to follow consistent, smooth, and monotonic paths when intensity
• 2: intensity_answers[1]
• 3: intensity_answers[2] was induced. The consistency of these patterns across traits and models
• 4: intensity_answers[3] reinforces the robustness and semantic alignment of our framework.
• 5: intensity_answers[4]
Next, we analyse inter-trait movement across LLMs to understand
For each question, please provide an answer that best represents the trait trait whether or not LLMs are capable of emulating latent psychological
at the intensity of trait_const
structures known in human personality science. Co-movers were selected
by identifying the two traits with the largest deltas (positive or negative)
Figure 3: Trait Intensity Induction Prompt Template accompanying the induction of a target trait as can be seen in Figure 5.
• 1: mildly warm, occasionally empathetic, reservedly caring While this figure shows the results for Claude 3.7 Sonnet, we observed
• 3: friendly, attentive, genuinely supportive similar modulation and co-movement patterns in GPT-4o and Gemini
• 5: extremely warm, deeply empathetic, overwhelmingly supportive 2.5 Flash as well, indicating that the SAC framework generalizes well
across architectures (see supplementary material [12]). The presence
The full mapping of traits to their corresponding adjectives across all
and behavior of these co-movers offer insights into whether LLMs
intensity levels is given in the supplementary material [12]. The intensity
capture inter-trait dependencies beyond the directly manipulated trait.
induction process followed a structured protocol. Each prompt began
Table 6 summarizes the top co-movers for each target trait based on
by defining the target trait (from the MPI-P2 inventory) and specifying
the most reactive trait shifts observed across all three LLMs. Across
its intended intensity (e.g., = 5). A set of trait-specific adjectives
all experiments, we observe that all three LLMs not only respond to
corresponding to intensity levels 1-5 was included to semantically anchor
individual traits with fidelity but also adjust semantically related traits in a
the model’s interpretation. The model was then asked to respond to a
manner consistent with human personality psychology. These structured
behavioral question formed by combining one of five intensity factors
co-movements observed in Table 6 are not merely reactive but suggest
(e.g., frequency, willingness) with a trait-linked action phrase. This
that LLMs have internalized coherent personality architectures. Our
design minimized semantic drift, ensuring alignment between output and
intensity-control method not only reveals these underlying structures, but
intended trait level. In total, the SAC-Induced framework prompted each
also enables precise and interpretable persona shaping.
model with 11,520 questions (16 traits induced × 3 intensity levels × 5
intensity factors × 3 behavioural prompts x 16 traits observed), providing
high-resolution insight into how trait intensity modulates model behavior 6 Limitations
across multiple dimensions. The prompt template in Figure 3 illustrates While our work advances personality modelling by introducing
this setup, with annotated placeholders that have the following features: fine-grained intensity control through SAC prompting, several limitations
• trait_const: The target trait being induced. remain. First, the reliance on self-reported questionnaire-style responses
• definition: A description of the trait adapted from the MPI-P2 inven- may not fully capture deeper behavioral tendencies in LLMs, which lack
tory. persistent internal states. Second, observed differences in trait expression
• intensity_scale: The desired intensity level (1-5). across models may conflate genuine induction effects with inherent
• adjectives: Trait-specific adjectives associated with each intensity level. training data distributions or architectural biases. Third, while we focus
• intensity_question: The phrasing of the intensity dimension (Fre- on the 16PF framework, the generalizability of our method to alternative
quency, Depth, etc.). or hierarchical personality models remains an open question for future
• question: The behavioral action specific to the trait. work. Addressing these challenges will help build more robust, adaptive,
• intensity_answers: The descriptors for the 1-5 response options. and psychologically grounded AI systems.