Clinical Psychology: Interviewing Techniques
Clinical Psychology: Interviewing Techniques
Contents
Interview...............................................................................................................................................3
Observation.........................................................................................................................................30
Psychological Testing...........................................................................................................................45
Conclusion...........................................................................................................................................99
Reference...........................................................................................................................................100
Interview
An interview is a face-to-face social interaction between an interviewer and a
respondent, aimed at obtaining specific information with minimal bias and maximum
efficiency. Its success depends on the quality of interaction between both parties, as well as
factors such as accessibility, cognition, and motivation. The respondent’s verbal and non-
verbal responses to the interviewer’s questions provide important insights and can influence
the interviewer’s perceptions. Similarly, during the interview, the respondent forms
impressions about the interviewer, and these impressions may affect their answers.
There are two main types of interviews: formal and informal. A formal interview is
one in which pre-prepared questions are asked in a fixed sequence, with responses recorded
systematic procedure for collecting information and can be conducted smoothly even by less
experienced interviewers. The structured nature of formal interviews often results in higher
reported validity compared to informal interviews; however, they can be expensive and time-
consuming, and their validity may still be lower than that of certain other assessment
An informal interview, on the other hand, does not have a predetermined set of
questions or a fixed order. The interviewer has the flexibility to decide what to ask and how
to ask it, allowing for a deeper exploration of the respondent’s personality and behavior. This
type, also known as an unstructured interview, is valued for its adaptability but comes with
certain limitations. It is more prone to interviewer bias, requires greater skill, tact, and subject
knowledge, and the data obtained is often difficult to compare or analyze due to variations in
questions, language, and responses. While formal interviews offer structure and ease of
execution, informal interviews provide richer insights but demand more expertise and present
functions that give it an advantage over other data collection methods: description and
exploration.
Description involves using interviews to gain valuable insights into the interactive
aspects of social life. Since interviews typically involve extended verbal interaction between
individuals, they allow the interviewer to better understand the respondent’s perspective on
the subject being studied. This deeper understanding helps the interviewer interpret the
respondent’s social life, which might otherwise appear abstract or be reduced to mere
statistics.
identify new variables for study, clarify concepts, and stimulate the development of
hypotheses for future research. This function allows researchers to probe respondents’
description.
interviews usually occur early in a client’s contact with a clinic, with the primary though not
exclusive goal of clarifying the clinician’s understanding of the client’s problems to plan
behavior. In both forms, the client’s problems and needs remain central, and even in the
earliest assessment interview, the clinician plays a therapeutic role by attending to the client’s
distress and supporting problem-solving. Thus, learning about the client’s concerns and
facilitating solutions are interconnected parts of a continuous process that requires genuine
testing. The results of these are combined to determine the therapeutic plan. However, for
most clients, it is both less stressful and more effective to consolidate these procedures into a
single initial interview that initiates the clinical process. This initial interview serves multiple
purposes:
Unlike older diagnostic interviews, the initial interview is less rigid, following a free-
flowing exchange in which the client’s concerns are expressed in their own words while still
As an assessment tool, the clinical interview is essential for exploring the client’s
conscious concerns, emotions, and problems as they experience them. Clients share details
about their life circumstances, relationships, successes, failures, and aspects of their
personality, all framed in personal meaning. This information emerges either spontaneously
or in response to the interviewer’s prompts. Throughout, the client is also responding to the
unique interpersonal context of the interview and the presence of the clinician, whose
qualities and behavior influence what the client shares, emphasizes, or feels. Unlike
psychological testing, which relies on passive, nonhuman stimuli, the interview is an active,
interpersonal encounter in which the clinician acts as both participant and observer.
Information from the interview includes the client’s self-reported feelings and life
history, observable behaviors, both deliberate and unintended, and reactions to the clinician’s
real or perceived actions. The clinician must skillfully attend to the client’s words, observe
their behavior, and remain aware of their own influence on the interaction tasks requiring
purpose” (Bingham & Moore, 1924), their structure and content differ depending on the goal.
Highly structured formats, such as census-taking or surveys, contrast with the open-ended
nature of clinical interviews. Although social research and clinical interviews share certain
aim to collect consistent, comparable data from many individuals to meet the researcher’s
objectives, with less emphasis on the participant’s needs. Clinical interviews, however, are
Kinds of Interviews
In modern clinical practice, there is a tendency to reduce the number and variety of
pre-treatment assessment interviews, which in the past were often carried out by different
professionals in different roles and might not involve the person who would later assume
therapeutic responsibilities. The envisioned initial interview addresses many areas typically
covered in specialized interviews, though in less depth, and is conducted by the person who is
most likely to continue as the client’s therapist. This approach streamlines clinical practice
and avoids the discomfort of being passed between multiple professionals and having to
repeat the same information. From the client’s perspective, this means entering the helping
process immediately rather than feeling they must first go through a series of preliminary
assessments.
Despite this shift, specialized interviews remain common in clinical settings. There
are situations where dividing responsibilities among professionals with specific expertise is
advantageous, as well as cases where the in-depth information obtained through targeted
interviews is essential for quality clinical care or research. Therefore, clinicians should
maintain the ability to conduct focused interviews tailored to the purposes of traditional
evaluation, modelled after the diagnostic procedures commonly used in medicine. Emerging
largely from hospital practice within the Kraepelinian tradition, it is still most frequently
applied to patients with psychotic disorders. The primary focus of this interview is on the
patient’s symptoms, aiming to document the type, severity, duration, history, and likely
progression of the psychiatric condition as precisely as possible. The format often follows a
1) Intellect and thought processes evaluating the ability to think accurately, quickly,
5) Insight and self-concept evaluating awareness of the illness, possible causes, and
self-perception.
The primary purpose of the intake interview is to introduce the patient to the clinic
and determine whether its services align with their needs. This process focuses on
understanding the patient’s goals, motivation for seeking treatment, expectations of the clinic,
and potential alternative options. During the interview, the patient is informed about clinic
procedures, fees, schedules, and other relevant details to help them make informed decisions
social workers, who possess specialized knowledge about clinic operations and alternative
community resources. Based on this initial meeting, arrangements are made for follow-up
visits or, if more suitable, referral to another agency that can better address the patient’s
needs. While centered on immediate concerns and the applicability of clinic services, the
In many cases, patients first make contact by telephone. While this is often a practical
means of scheduling the initial appointment, it can also serve as a first step for individuals
tentatively seeking help. Clinicians may receive inquiries such as, “Could you tell me what
you do in your clinic?” or “What would you think of a person who…?” The telephone allows
patients to express concerns anonymously, shielding their identity while still revealing fears
or preoccupations. Because these medium lacks many of the nonverbal cues available in face-
to-face interactions, telephone interviews require significant skill and patience to identify the
caller’s primary concerns and, when appropriate, encourage them to attend an in-person visit.
gather a detailed history of the patient. Unlike the diagnostic interview, which focuses
primarily on symptoms and specific problems, the social history interview seeks to develop a
comprehensive understanding of the patient’s life course and current personal and social
discuss topics such as early childhood experiences, family relationships, education, hobbies,
employment history, dating and sexual relationships, current social activities, marriage, work,
events rather than on exploring emotionally significant experiences tied to the patient’s
their current personality structure and functioning. Equally important is gaining insight into
the patient’s present life situation, including the stresses and environmental realities that
contacts, such as a spouse, parents, or others with whom the patient has a significant
relationship. This approach is typically used when the patient is unable or unwilling to
young children. In the dynamic tradition, aside from child guidance work, many clinicians
hold strong clinical and ethical reservations about interviewing third-party informants. Their
primary focus is on the patient’s subjective world, including how they perceive others’
Such collateral information should never be gathered without the patient’s explicit
consent. There are, however, situations in which both clinician and patient recognize the
value of involving a relevant other such as a spouse or parent both to share and to receive
information. These sessions are most appropriate after a therapeutic relationship is well
established and are rarely warranted during the initial assessment phase, except in cases
consent and sometimes at their own request, is seen by a more experienced consultant. The
structure of this interview varies depending on the issues under discussion. Historically, it
often took the form of a diagnostic evaluation conducted by a senior authority. While
consultation has been deemphasized in the therapeutic model, it has regained relevance in
community-based work, where clinicians may operate through individuals such as teachers
These interviews are brief and aim to identify individuals who require further, more detailed
evaluation.
Transfer, Furlough, or Discharge Interview. These interviews take place when a
prepared for discharge. They help ensure continuity of care, clarify expectations, and address
professionals. The primary purpose of this interview is to establish rapport and prepare the
patient for the testing process. The clinician should explain the nature and purpose of the
tests, the types of activities involved, the intended use of the results, and assure the patient of
confidentiality. Additionally, the clinician should identify personal or situational factors that
may affect test interpretation. Failure to give this process sufficient attention can result in
clinical insensitivity, undermining trust and potentially affecting the quality of the
assessment.
generally more structured and narrowly focused. Their format and content are determined
primarily by the objectives of the study rather than the individual needs of the patient. While
the interviewer should remain professional and considerate, their role is to gather the same
type of information from each participant, following a consistent sequence of topics and
asking identical questions in a similar manner. This uniformity, often achieved by having the
same interviewer conduct all sessions, reduces the open-ended flow characteristic of clinical
wording, sequence, probing techniques, and recording methods—that influence the reliability
and validity of the data. Ethical standards require that participants are never subjected to
harm or humiliation, and that all research procedures are carried out only with their informed
and voluntary consent. This is especially important in clinical settings, where patients may
mistakenly believe that the research is part of their treatment or that participation is necessary
explain the purpose and nature of the study and to assure patients that their decision to
physicians, psychiatrists, courts, schools, employers, social service agencies, and other
organizations. In these cases, clients are referred with the purpose of answering a specific
intellectual disability, or whether a proposed custody arrangement would serve a child’s best
interests.
The primary objective of this type of interview is to address the referral question,
making it essential that the question is stated with clarity. Vague or overly broad requests
such as “Provide a profile of Mr. Q” or “Will Ms. Y be a good parent?” are not sufficiently
precise. Likewise, requests like “Please test my child’s IQ so I can prove to the schools that
he should be in the gifted class” should prompt caution regarding the appropriateness of
conducting the assessment without first clarifying the parent’s motives and needs. Once the
referral question is clearly defined, it guides the nature and scope of the assessment
conducted.
treatment often have limited knowledge about what to expect or what is expected of them,
particularly if they have never interacted with mental health professionals before. To reduce
uncertainty and promote comfort, many clinicians conduct orientation interviews or dedicate
portions of other interviews to familiarize clients with the assessment, treatment, or research
Orientation interviews serve at least two key purposes. First, by inviting clients to ask
questions and share comments, clinicians can identify and address misconceptions that might
otherwise hinder treatment progress. Second, these interviews help clients understand the
procedures ahead and their role within them (Couch, 1995). For example, clinicians may
explain that those who gain the most from treatment are typically open, cooperative,
committed, and actively engaged in addressing their problems. In this way, orientation
interviews can help align clinician efforts with clients most willing to participate fully in the
These interviews are equally important in research contexts. While researchers may
withhold certain details of the study design to prevent bias, they remain ethically obligated to
ensure that participants understand the nature of the tasks involved and any associated risks.
In research, orientation interviews not only fulfil informed consent requirements but also
enhance participant motivation an especially critical factor in long-term clinical trials and
longitudinal studies.
provide clients with information and assess their understanding after a significant event.
These interviews may take place in various contexts. For instance, clients who have
often want to know the results, how the information will be used, and who will have access to
it. Such concerns are heightened when the assessor has served as a consultant to institutions
such as schools or courts. In these cases, debriefing interviews can help reduce client anxiety
by clarifying the procedures for safeguarding confidential information and offering a
Debriefing interviews also occur in situations such as with soldiers following combat
experiences (Adler, Bliese, McGurk, Hoge, & Castro, 2011), with clients who have attempted
realistic simulation exercises (Morse, 2012). Regardless of the setting, the primary aim of
these interviews is to help individuals process and understand the event, maximizing potential
relationship. In cases where psychological treatment has ended successfully, these interviews
serve to address unresolved matters, express and receive gratitude, offer reminders for
managing future challenges, make plans for follow-up contact, and reassure clients of their
ability to maintain progress independently. These steps help make the transition from
When treatment ends less successfully such as when clients discontinue early
termination interviews can provide valuable insight into the factors contributing to dropout.
This information can inform adjustments to treatment structure in the future (Hummelen,
Termination and debriefing interviews can also take place after research participation.
information about the study. At times, researchers may also conduct interviews to better
Crisis Interviews. When individuals in crisis arrive at a clinical facility or reach out
interviews, aiming to provide support, gather essential assessment information, and offer
2008).
In these situations, the interviewer approaches the client with calmness and
acceptance, asks targeted and relevant questions (e.g., “Have you ever attempted suicide?” or
“What medications do you currently have at home?”), and addresses the urgent issue directly
or connects the client to appropriate services. For some individuals, one or two well-executed
crisis interviews may resolve the immediate problem and conclude contact, particularly when
the need for help is temporary and tied to specific circumstances. For others, the crisis
interview may serve as the entry point to further assessment and ongoing treatment.
people from different ethnic backgrounds, often due to limited cultural knowledge. For
as paranoia rather than caution about the mental health system. Similarly, an Asian client’s
hesitance to disclose may reflect cultural norms rather than resistance or lack of insight.
Some mental health concerns are culture-specific and don’t fit neatly into standard
includes anxiety and physical symptoms but doesn’t exactly match any DSM diagnosis.
Cultural differences in how distress is expressed such as Asian clients reporting more
physical symptoms like nausea or dizziness can lead to misinterpretation. These patterns may
reflect cultural communication styles rather than actual differences in disorder rates.
Cultural values also influence interviews. In Hispanic and Asian cultures,
interdependence and family obligations may be more valued than personal independence,
While it’s unrealistic for clinicians to know every cultural variation, they should
competence involves learning about common cultural patterns, being aware of the
backgrounds of clients they serve, and openly discussing cultural concerns during sessions.
Clinicians should acknowledge their own limitations and seek guidance from colleagues with
relevant expertise.
education and practice can help build cultural awareness and skills. Research in this area is
expanding, with studies examining topics such as the validity of diagnostic tools across
languages, evaluating cross-cultural training, and improving services for refugees. Clinicians
are not expected to know everything, but they should recognize when their cultural
A good interview space should be quiet, comfortable, and free from distractions. The
furniture doesn’t need to be fancy, but both clinician and patient should have similar chairs—
equal in height and comfort—positioned to maintain personal space without feeling too
distant. The environment should encourage open conversation, so it’s best to avoid settings
where one person towers over the other or where a large desk creates unnecessary barriers.
Some clinicians use a desk or table for note-taking, while others prefer more casual
seating with a small coffee table. Personal touches like books, art, or mementos can make the
room more inviting, as long as they are tasteful and professional. Sensitive items such as
personal letters or patient files should never be left out. The room must also ensure privacy—
patients shouldn’t be able to hear others (or be overheard), and interruptions should be
confidentiality.
While the clinician needs to recall important details later, note-taking during the
interview is debated. Writing too much can distract from listening and maintaining eye
contact. However, jotting down brief key words is fine if it doesn’t disrupt communication.
Even small gestures like leaning forward or saying “uh-huh” can show interest without
Recording interviews common in training and research should only be done with the
patient’s full knowledge and consent. They should have the option to erase the recording
afterward if they wish. The microphone should be visible, both for transparency and for better
sound quality. The biggest challenge is often for the clinician, who may feel self-conscious
about being recorded. Preparations should be made before the session starts to avoid
nonverbal communication as well. While more revealing, it follows the same rules as audio
Finally, not all valuable clinical interactions happen in the consultation room—
hallways, waiting areas, and casual encounters can also provide meaningful insights. Clinical
communication can take place in any setting where patient and clinician meet.
Stages in the Initial (Assessment) Interview
The initial interview, has four major purposes which generally emerge sequentially during
the session:
In a similarly general way, the session moves from emphasis on present feelings,
toward concern with the patient’s past experiences, and ends oriented toward future plans and
actions.
The clinician takes the role of host and the patient that of guest, with the first few
practical gestures like arranging chairs, offering an ashtray, and exchanging introductions.
From the outset, the clinician’s manner communicates genuine concern for the patient’s
comfort and an understanding of the apprehension that can come with a first visit.
A simple opening question, such as “Tell me what brings you here?”, is usually
sufficient. Some patients will immediately share details about their problems, while others
might focus on why they came to this particular clinic at this time. In either case, early
questions can reveal the patient’s perspective on their difficulties, the degree to which they
accept personal responsibility, and whether they view their issues as psychological or
externally caused. Often, a recent crisis has triggered the visit, even if the problems are long-
standing. Understanding the circumstances leading to the decision to seek help is an
Clinicians should not assume that a patient’s presence means they fully acknowledge
their need for help sometimes the visit is prompted by someone else’s insistence. Early in the
interview, questions should be brief and minimal, allowing the patient to bring up topics
important to them. While intriguing details may arise, deeper probing can wait until after the
patient has freely covered the issues most pressing to them. Questions that impose structure—
such as asking about exact times and places—should be postponed to better understand the
Although the clinician may guide the conversation when the patient gets stuck or
repetitive, the goal is to give as much freedom as the patient can handle. Some may feel lost
without direction, in which case gentle prompts like “That’s fine, but why don’t you tell me
more about…” can help. The aim is neither aggressive questioning nor passive listening, but
Patients often feel anxious due to the unfamiliar setting, the discomfort of sharing
vulnerabilities with a stranger, the emotional strain of revisiting painful events, and
uncertainty about the clinician’s assessment. The clinician must respond to this anxiety with
respect, attentiveness, and warm yet objective understanding. Supporting the patient through
the first meeting is best done by example rather than lengthy explanation, and the clinician
should acknowledge painful emotions without overdoing sympathy. Statements like “It’s hard
to talk about…” convey empathy, while reassurances like “Don’t worry, lots of people feel
Patients may also have questions about the clinician or the clinic. When possible, such
discussions should be postponed until later in the session so the focus remains initially on the
patient’s concerns. If direct questions are asked, they should be answered honestly but briefly.
Ultimately, the purpose of the opening phase is to set a tone of trust and openness,
encouraging the patient to share their situation in a way that feels personally meaningful.
The central portion of the interview focuses on gathering enough information to form at
least a preliminary understanding of the patient’s problems and personality. The clinician
aims to learn:
1. The patient’s current symptoms and concerns (“presenting problems”), the reasons for
seeking help now, the nature of any immediate crisis, and current life circumstances.
2. Whether recent stressful events may have disrupted coping mechanisms and triggered
major traits, emotional patterns, defenses, and conflicts — particularly those related to
the current difficulties, as well as any notable changes in behavior. Early life
treatment.
Since not all of these areas can be covered in depth, the clinician focuses on obtaining
enough detail to create working hypotheses, with the option of follow-up assessment sessions
Discussion typically begins with the patient’s immediate concerns, then moves
outward to related past events. After learning about the present distress, the clinician might
ask questions like, “How long has this been happening?” or “What was life like before then?”
Identifying events that preceded the onset often provides important insight. These can be
either negative (e.g., loss of a loved one) or positive but still disruptive (e.g., birth of a child),
as both can significantly challenge established coping patterns. The clinician remains alert to
There is no rigid checklist of questions; the flow is guided by the patient’s narrative,
emotional emphasis, and avoidance patterns. Attention is paid to both what is said and what is
left unsaid. If the patient mentions a topic briefly and moves on, the clinician may later return
to it with neutral prompts (e.g., “You mentioned your brother earlier — could you tell me
more?”). Sometimes, a more probing question (e.g., “I wonder why?”) may be used, though
Even with cooperative, articulate patients, the information gathered may feel
incomplete or disorganized. Distress, painful memories, and defensive behavior can lead to
fragmented stories, inconsistencies, or seemingly trivial focus points. The clinician’s role is to
work through this confusion without adding to it, piecing together a realistic picture of a
By the conclusion of this stage, the clinician should have a provisional formulation that
includes:
emotional functioning
This working understanding informs decisions about whether treatment is appropriate, what
type might be most effective, and what the initial goals should be. The clinician also
evaluates the patient’s openness to therapy, motivation for change, insight, and psychological
readiness, while considering practical and social factors. At this stage, possible outcomes
range from ongoing therapy, referral to other services, or immediate intervention in urgent
Before concluding the interview, it is important to help the patient regain composure,
provide relevant information, and collaborate on the next steps. Although the closing stage is
not primarily for assessment, additional insights about the patient may still emerge, just as in
earlier phases.
The interview process can be emotionally taxing for the patient, as it often involves
recalling distressing memories and sharing vulnerable emotions with someone relatively
unfamiliar. Such disclosures may provoke anxiety or other uncomfortable feelings. Therefore,
before ending the session, the clinician should aim to restore the patient’s sense of stability
and, ideally, leave them with some hope or a feeling of progress. This should not be achieved
through superficial reassurances or clichéd statements that minimize the seriousness of the
concerns. Instead, the clinician should communicate empathy for the challenges of discussing
personal issues, acknowledge the significance of the patient’s difficulties, and convey
optimism for future improvement. Expressions of genuine concern and practical suggestions
However, the clinician should avoid granting unconditional approval or indulging unrealistic
demands, as this may reinforce maladaptive patterns rather than prepare the patient for the
challenges of therapy.
The clinician should share their tentative understanding of the patient’s concerns and
prognosis, or treatment in ways that resemble expectations from medical consultations. While
avoiding an overly authoritative stance, the clinician can explain the differences between
psychological and medical approaches and summarize the situation, available options, and
their professional recommendation. The ultimate decision rests with the patient, who should
be fully informed about the nature, accessibility, cost, and procedures of potential treatments.
temporary and likely to resolve with situational change, or because it stems from chronic
After the patient departs, the clinician should take time to reflect on the session,
reviewing notes and recalling the sequence of events without the pressure of immediate
response. This reflective process often reveals patterns or relationships that were less
apparent during the interaction. Writing a detailed account and integrating observations into a
theoretical framework can deepen understanding, while care must be taken to preserve the
interpersonal communication in which both intended and unintended messages are exchanged
between participants. Accurate clinical assessment relies heavily on the clarity and
Language is an imperfect system for representing the full range of human perceptions
and concepts. Accurately expressing nuances of emotion, for example, can be difficult.
Effective communication requires that both the sender and receiver not only have the
necessary vocabulary but also share the same meaning for the words used. A lack of shared
vocabulary can hinder interaction, as word usage varies across social classes, ethnic groups,
regions, professions, and age groups. People with similar social characteristics are generally
clinician and the patient. Teaching the patient the clinician’s terminology can be challenging
and potentially alienating, while adopting the patient’s vocabulary can sometimes appear
spontaneous word choices while avoiding excessive use of technical jargon or artificial
imitation of the patient’s idiom. Communication should draw from shared vocabulary and
reference. For example, statements such as “I was offered a good job” or “She is an
interesting woman” are open to multiple interpretations depending on the speaker’s personal
standards or social background. The clinician must determine the context of such terms by
asking clarifying questions (e.g., “Do you mean in terms of duties, hours, or salary?”).
Clinicians generally begin with neutral, direct, and open-ended questions that allow
patients to present their concerns in their own words. For example, asking “Tell me
something about your marriage” is preferable to “Is your marriage happy?” Later, more
targeted or provocative questions may be appropriate, but leading questions, abrupt shifts in
patient is unwilling or unable to speak about sensitive topics directly. For example, asking
“What might make a person consider suicide?” can elicit valuable information about suicidal
Nonverbal Communication
Information in an interview is conveyed not only through words but also through
nonverbal behaviors. Some nonverbal acts, such as placing a finger to the lips to signal
“quiet,” are intentional. However, gestures, body movements, posture, gait, facial
expressions, and vocal patterns can also reveal information unintentionally. These elements
offer valuable cues for informal assessment. While clinicians have traditionally focused more
on patients’ verbal expressions, skilled practitioners also remain attentive to their bodily
behavior. Freud (1905) observed that individuals cannot fully conceal their inner states—if
silent, they may still reveal themselves through subtle physical actions. Nonverbal behaviors
can reflect stable aspects of personality or clinical condition, or they can communicate
temporary emotional states. For example, severely restricted individuals may display tense
major depression often move and speak slowly and lethargically. Short-term bodily cues,
such as fidgeting, stammering, trembling, clenched fists, or facial flushing, can indicate
emotions such as embarrassment, anxiety, or anger. Some expressions are universal due to
shared physiological responses, while others are shaped by culture or are entirely unique to
the individual.
Research has shown that personally meaningful gestures often recur throughout an interview.
Mahl (1968), for example, documented two recurring gestures from a patient, “Mrs. B.”—
turning her palms upward and outward, and playing with her rings—and noted how their
frequency varied across interview minutes. When compared with the patient’s verbal content
and interaction with the interviewer, four relationships emerged: (1) some gestures conveyed
the same meaning as the accompanying verbal statements, such as palms-up gestures when
expressing helplessness; (2) some gestures appeared unrelated to the immediate verbal
content but foreshadowed later themes, such as playing with a wedding ring before later
discussing marital dissatisfaction; (3) some gestures contradicted verbal claims, as when an
apprentice machinist dropped a pencil precisely when insisting his work was flawless; and (4)
some gestures were directly influenced by the interviewer’s responses, as when Mrs. B.’s
In interview contexts, reliability refers to the consistency with which clients provide
the same information across different occasions or to different interviewers. Validity refers to
the degree to which interview data or conclusions accurately reflect the truth. These two
interview-based information.
Reliability
responses over repeated interviews (test–retest reliability) and by measuring the extent to
which different evaluators agree on their conclusions from interviews with the same client
(interrater reliability). A common method for evaluating interrater reliability involves having
multiple clinicians view recorded interviews and independently provide ratings or other
inferences. This approach has been widely used in studies on the consistency of DSM
diagnoses (Widiger et al., 1991), evaluations of client progress (Goins et al., 1995),
assessments of therapeutic alliance quality in intake sessions with immigrants (Shechtman &
Because interview formats and purposes vary greatly, no single reliability value
applies universally (Craig, 2009). Nonetheless, reliability is generally higher when interviews
are conducted close together in time and when adult clients are asked about non-sensitive
demographic information (Ross et al., 1995). Reliability tends to decrease when the interval
between interviews is long, when the interviewee is a young child, or when topics involve
sensitive matters such as illegal drug use, sexual behavior, or traumatic events (Schwab-Stone
Although sensitive topics are often the focus of clinical interest, structured interviews
can help elicit more reliable responses in such cases. Commonly used structured formats,
such as the Diagnostic Interview Schedule for Children (DISC), have demonstrated
satisfactory test–retest reliability (Flisher et al., 2012). Similarly, the interrater reliability of
the Structured Clinical Interview for DSM-IV (SCID) scales is generally moderate to
Reliability can also be affected by the population for which the interview method was
developed. A format that performs well with one cultural group may not produce the same
allows clinicians to use certain interviews with more confidence. For example, the DISC has
been shown to work reliably with South African children (Flisher et al., 2012), and the
seek evidence on both cross-cultural reliability and validity before applying interview
Validity
alter information. This risk increases for individuals with intellectual disabilities (Heal &
Sigelman, 1995), brain disorders (West et al., 1991), or a desire to conceal facts about
behaviors such as drug use, sexual activity, criminal conduct, or prior hospitalizations
(Morrison et al., 1995). Conversely, some clients may exaggerate or fabricate symptoms to
create the impression of a mental disorder a behavior known as malingering which has
prompted the development of specialized detection methods (Rogers et al., 1991). More
broadly, the tendency to present oneself in a particular way to a mental health professional,
Validity can be evaluated through multiple approaches, such as ensuring all relevant
aspects of a topic are covered (content validity), comparing interview results with established
measures of the same construct (concurrent validity), or testing the ability of interview data to
predict future outcomes (predictive validity). The latter two methods require an external
benchmark, often referred to as the gold standard (Komiti et al., 2001). Structured diagnostic
interviews are frequently validated against this standard of clinical judgment (Zetin & Glenn,
1999), although in some cases clinical judgment is validated against structured interviews
instance, when an interview can accurately distinguish clients likely to drop out of therapy
personality disorder.
Given the wide variation in interview purposes and administration styles, no single
conclusion applies to all interviews. However, formats that are more structured and validated
validity tends to improve when there is strong rapport between interviewer and client, and
Interviews, as a research method, offer both benefits and limitations. Gordon (1969) outlines
2. Efficient information gathering – They enable the quick and targeted collection of
relevant data.
3. Interpretive accuracy – Direct interaction helps ensure that responses are understood
4. Contextual control – The interviewer can manage the setting and timing of questions
to maintain relevance.
to assess accuracy.
Despite these strengths, interviews have notable disadvantages:
1. Interviewer’s variability – Differences in how interviewers ask questions, interpret
unstructured formats.
3. Validity and dependability concerns – Respondents may not behave or answer as they
truly would outside the interview context, reducing the accuracy of verbal reports.
interview data, and approaches vary in completeness and precision, which can affect
dependability.
While structured interviews can reduce some of these disadvantages, certain limitations—
Observation
In behavioral research, questionnaires and interviews are widely used methods for
data collection. However, there are circumstances in which these tools are not appropriate or
meaningful. For instance, when the aim is to study behavior in a natural context and focus on
involve watching participants in experimental settings, noting the expressions and reactions
of respondents during interviews, or even observing how people behave while answering
survey questions. While all investigators inevitably have some degree of direct contact with
and listening to the behavior of others over time without manipulating or controlling the
situation. It focuses on recording and interpreting findings in a way that allows for analysis
and discussion. As such, observation entails the careful selection, recording, and coding of
1968):
1. It occurs within a natural social context in which the individual’s behavior is studied.
Although it most often takes place in natural settings, observation can also be applied
2. It captures significant events or interactions that influence the relationships among the
3. It identifies recurring patterns in social life by comparing and contrasting data from a
made during the research process, giving it a more scientific and methodical basis.
Goals of Observational Assessment
evident that the interpretation of observations was often subjective, as different observers
might notice different aspects or assign varying meanings to the same behavior. To address
this issue, observational assessments have been made increasingly structured, similar to the
progression seen with interviews. Modern observational techniques typically specify the
exact behaviors to be observed and outline clear procedures for recording, combining, and
interpreting the data. While many of these structured methods were initially designed for
research purposes, improvements in their reliability, validity, and user-friendliness have led to
Supplementing Self-Reports
Self-reports collected through interviews or certain tests can often be unreliable. Most
individuals find it challenging to give objective and unbiased accounts of their own behavior,
especially when it concerns emotionally charged situations. For instance, a distressed couple
may struggle to accurately describe their interactions, particularly during conflicts. Likewise,
individuals with conditions such as dementia may be unable to provide precise self-reports,
even when trying their best. In such cases, observational data tends to offer more valid and
dependable insights.
This is common among participants in programs for smoking cessation, drug rehabilitation,
or alcohol treatment, which is why such self-reports are often supplemented with input from
family members or biological tests that can confirm the presence of certain substances.
Traditional assessment often assumes that interviews and test responses are sufficient
for understanding a client’s personality and problems, as these responses are believed to
reflect general traits that guide behavior. Clinicians with this perspective view observations as
clinicians see observational data as direct samples of behavior that reveal meaningful person–
situation interactions. They place less emphasis on inferring stable personality traits and
instead focus on the specific conditions under which behaviors occur. By using observation-
based functional analyses, clinicians can avoid the high level of inference involved in sign-
oriented testing and interviews. These observational methods focus on gathering factual,
concrete data, reducing the risk of misinterpretation. They help identify the situations most
likely to elicit problematic behaviors, the triggers that provoke them, and the consequences
that reinforce and maintain them. In comparison, traditional tests and interviews are not
Observations conducted within the client’s natural physical and social settings can
offer the most accurate representation of their experiences and difficulties. Such assessments
tend to have strong ecological validity and often reveal situational factors that assist
workplace contexts. Tailoring interventions in this way can enhance the likelihood of
virtual reality. Some methods blend both approaches to meet specific assessment needs.
but acknowledged. Regardless of the approach, selecting target behaviors and using
systematic recording methods are essential to minimize bias and improve accuracy (Haro et
al., 2006; Antony & Barlow, 2010; Haynes & O’Brien, 2000; Hersen & Beidel, 2012;
Madden et al., 2013; Miller & Leffard, 2007; Repp & Horner, 2000).
researchable data, Reiss (1971) distinguishes between systematic observation, which follows
explicit, predetermined principles and scientific logic, and unsystematic observation, which is
casual and lacks specific guidelines. Observation can also be categorized by the investigator’s
group activities, either overtly or covertly, to gather realistic and meaningful data, though it
may lack precision, be time-consuming, and risk bias from personal involvement. In contrast,
structured manner, making it more reliable and representative. Both approaches have distinct
strengths and limitations, with choice depending on the research goals and context (Reiss,
1971).
Naturalistic Observation
Natural settings like the home, school, or workplace offer a realistic and contextually
relevant backdrop for understanding a client’s behavior and the factors that influence it.
When conducted subtly, naturalistic observation can capture behavior without distortion from
identifying and monitoring the nature and progression of issues presented for clinical
intervention, which may include eating disorders, intrusive thoughts, maladaptive social
behaviors, parenting challenges, psychotic symptoms, and other concerns such as classroom
which a researcher becomes part of a tribe, subculture, or social group to document its
characteristics and members’ behaviors, often producing ethnographic accounts (Mead, 1928;
Williams, 1967). To reduce the influence of outsider presence, observations may also be
conducted by individuals already within the client’s everyday environment, using systems
that allow for minimally intrusive recording of behavior frequency, intensity, duration, or
form. Early home-based clinical observations were often unsystematic (Ackerman, 1958), but
more reliable systems, such as Patterson et al.’s (1969) approach for conduct-disordered
children and the Home Observation for Measurement of the Environment (HOME) for
assessing developmental influences (Glad et al., 2012), have since been developed.
Behavioral observation tools are also widely used in schools, playgrounds, and similar
settings for clinical and educational purposes (Bunte et al., 2013; Nock & Kurtz, 2005;
Ollendick & Greene, 1990). In classrooms, observers may focus on a single child, multiple
children in sequence, or the entire class (Milich & Fitzgerald, 1985). Observation is equally
important in hospital settings, where tools exist to assess nonverbal indicators of pain in
dementia patients (Lints-Martindale et al., 2012). Additionally, hospitals and clinics serve as
valuable settings for supervising and evaluating trainees learning assessment skills, with
structured observation measures helping ensure accurate and constructive feedback
(Holmboe, 2004).
Self-Observation
intensity of various events, such as exercise, headaches, positive thoughts, hair pulling,
smoking, eating patterns, stress, sleep disturbances, anxiety, and health-promoting activities.
Typically, the client and therapist decide together on specific target behaviors, and the client
keeps a record or diary noting when these behaviors happen, the circumstances surrounding
them, and often the thoughts linked to these moments. However, self-monitoring in areas like
drug use can sometimes lead to underreporting or, less commonly, overreporting (Clark &
Winters, 2002). Despite these challenges, many brief self-report tools for symptoms and
Using insiders as observers of adult behavior for clinical purposes is less common but
still practiced. For instance, when helping clients quit smoking, clinicians might request
Lichtenstein, & McIntyre, 1983). Such third-party reports are also gathered during
assessments of alcoholism or drug use (e.g., Frank et al., 2005), sexual behavior (e.g., Rosen
& Kopel, 1977), marital interactions (Johnson, 2002; Jouriles & O’Leary, 1985), and other
adult behaviors.
behavior. For example, school grades, arrest records, and court files have been used to
evaluate treatments for delinquent youth and adult offenders (Davidson, Redner, Blakely,
Mitchell, & Emshoff, 1987; Rice, 1997). Changes in academic grade point averages have
served as indicators of improvement in test anxiety (Allen, 1971). These life records, also
allows clinical psychologists and other behavioral scientists to study people’s behavior
In clinical research, unobtrusive measures can be used to test theories about the
signs of schizophrenia by analyzing home videos of children as they grew up—a method that
took advantage of many families' tendency to record videos (Walker, Grimes, Davis, &
Smith, 1993). Trained observers reviewed videos of individuals who later developed
schizophrenia, along with videos of their same-sex siblings who did not. The study found that
long before diagnosis, children who eventually developed schizophrenia showed significantly
more negative facial expressions compared to their siblings, with some differences evident
Controlled Observation
events can disrupt the assessment. For example, a client might move out of the observer’s
view or receive help from someone else when facing a stressor, raising the question of how
the client would have responded without assistance. Another drawback is the sometimes-
lengthy waiting periods before infrequent behaviors, such as a child’s tantrum or a family
specific, controlled situations where clients’ reactions to planned, standardized events can be
observed. This method, known as controlled observation, allows for greater control over the
assessment environment, similar to how psychological tests are administered. Controlled
observations are also referred to as analog behavior observation (ABO), situation tests, or
contrived observations.
methods to evaluate personality traits and behavioral skills. For example, in the Operational
Stress Test, pilot candidates were asked to operate an aircraft flight simulator while being
“You’re making too many errors”; Melton, 1947). The assessor rated candidates’ responses to
criticism and stress, complementing the objective performance data from the simulator.
The Office of Strategic Services (OSS), which later became the CIA, used
leadership in potential espionage agents and special personnel. Candidates were tasked with
building a five-foot cube-shaped frame from large wooden poles and blocks resembling a
giant Tinker Toy set. Two “assistants” (actually psychologists), named Kippy and Buster,
were assigned to interact with the candidate. Kippy acted passively and only followed orders,
often getting in the way, while Buster was aggressive, offering impractical suggestions,
criticizing, and highlighting the candidate’s weaknesses. Their role was to create as many
obstacles and annoyances as possible within 10 minutes. The candidates became so frustrated
that none completed the construction in the allotted time (OSS, 1948). Since then, less intense
versions of these situational tests have been used for personnel selection.
Psychologists sometimes create pretend situations where clients are asked to act out
how they usually behave. This technique, called role-playing, has been used by therapists for
many years and is important in group, psychodynamic, and humanistic therapies. But it only
became a regular part of clinical assessment in the late 1960s. Since then, role-playing tests
have been commonly used to observe children’s social and safety skills, parent-child
During these role-plays, clients’ or trainees’ actions are recorded and then rated by
observers based on many factors like how appropriate their responses are, how assertive or
anxious they seem, how quickly they reply, how long they talk, their body language (posture,
Sometimes, clinicians use a staged event that looks natural to the client but is actually
controlled. For example, social skills of psychiatric patients have been tested by having them
talk to a stranger who was instructed to challenge them in specific ways, like forgetting their
name or asking about themselves during the conversation. Similar setups have been used to
test parenting skills by having parents role-play with a doll or respond to videos of child
misbehavior.
Because these observations can involve some deception or privacy concerns, they
must be carefully designed to protect the client’s well-being and dignity. Those who support
these methods note that they are best for measuring specific behaviors, like refusal, rather
Physiological Measures
Some performance tests measure physical responses like heart rate, breathing, blood
pressure, sweating, muscle tension, and brain activity in response to different situations. A
well-known example is Gordon Paul’s (1966) study, where he measured heart rate and
sweating just before people gave a speech to identify those with speech anxiety. These
measures were also used after anxiety treatments to see if the treatments worked (see also
Recently, clinical psychologists have been using these physiological tests more often
because they study many health issues with psychological aspects, such as insomnia,
headaches, chronic pain, sexual problems, digestive disorders, HIV/AIDS, and diabetes (see
Chapter 12, “Health Psychology”). For example, in assessing sexual arousal and dysfunction,
male participants listen to or watch erotic audio or videos, while a device measures changes
in penis size (called phallometric measurement). Larger responses mean higher sexual
arousal. Some research found that child pornography offenders with stronger responses were
more likely to be pedophiles (Seto, Cantor, & Blanchard, 2006). However, clear patterns of
arousal for different sexual offenses haven’t been identified yet, but there is hope that new
technology, like pupil dilation measures (Rieger, 2012), might help in the future.
As interest grows in how psychological factors affect health and illness, the use of
especially since companies now offer affordable, portable devices and virtual reality systems
by a computer. Sometimes this is shown on a screen, but often clients use headsets, helmets,
and gloves that provide visual, sound, and sometimes touch sensations. This technology
allows precise control of three-dimensional stimuli, making the experience feel very real for
clients. During these sessions, clinicians can gather self-reports, observe behavior, and collect
physiological or other types of data. These measures have shown that VR is valuable both for
assessment and treatment purposes (Côtè & Bouchard, 2005; Gorman, 2006; Llobera et al.,
2013).
seem mostly unfounded. Flight simulators, for example, have long been effective tools for
training and assessing both civilian and military pilots. Except for some clients with autism
settings (Standen & Brown, 2005). For instance, Lew and colleagues (2005) used a driving
simulator to evaluate the long-term driving abilities of clients with traumatic brain injuries.
The VR test predicted future driving performance better than an actual on-road test. Another
study (Jouriles, Rowe, McDonald, Platt, & Gomez, 2010) found that combining role-playing
with VR stimuli produced stronger and more distinct responses than role-playing alone.
situations. During a BAT, clients face a feared stimulus while observers note the type and
level of avoidance behavior shown. Although informal BATs with children date back to the
1920s (e.g., Jones, 1924a,b), it wasn’t until the early 1960s that systematic avoidance testing
For example, in a study on systematic desensitization for snake phobia, Peter Lang
and David Lazovik (1963) asked clients to enter a room with a harmless caged snake and to
gradually approach, touch, and eventually hold it. Observers rated avoidance based on
whether the clients could look at, touch, or hold the snake. Other feared animals like rats,
spiders, cockroaches, and dogs have also been used in various BAT versions. Over time, the
simple “look–touch–hold” scoring system has been replaced by more advanced methods.
Nowadays, virtual reality technology is often used to create simulated environments where
A key step in ensuring that observational assessments are reliable and valid is clearly
defining what behavior will be measured. This means deciding exactly which aspects of
behavior to watch for and how to categorize them, based on the assessor’s understanding of
what the behavior means. For example, when assessing assertiveness, one clinician might
look at how well a client can say no to unreasonable requests, while another might focus on
how often the client shows positive emotions directly. Although this issue of defining
behaviors may never have one perfect answer, the process of checking an observational
system’s reliability and validity always starts by asking what specific behaviors should be
recorded.
When using observational assessment, clinicians must be careful that clients don’t
being watched. The situation itself can influence how clients act through social cues or
expectations (called demand characteristics; Orne, 1962) that suggest what behaviors are
appropriate. For example, if a couple is told, “We want to see how much you actually fight,”
they might behave in a way that shows more or less conflict than usual.
situations and responded. Their assertiveness was rated on a scale from 1 to 5. They heard the
recordings twice, but the instructions varied: in the “low-demand” condition, they were told
possible. Results showed that participants acted more assertively when given high-demand
instructions. If the instructions stayed the same for the second test, their assertiveness stayed
consistent, but if the instructions changed (more or less demanding), their assertiveness
changed accordingly. This shows that what participants think they are expected to do affects
Other research on anxiety assessment has found that instructions, the presence of the
experimenter, the physical environment, and other factors also influence how much fear
Although various methods have been suggested to reduce this situational bias
(Bernstein & Nietzel, 1977; Borkovec & O’Brien, 1976), it can never be completely
eliminated. Whenever the conditions during observation differ from those when the client is
not being observed, we cannot be sure the observed behavior will reflect their usual behavior
elsewhere. The best clinicians can do is minimize cues that might influence clients’ actions.
behaviors.
change over time. For example, a couple might show a lot of anger during one discussion but
much less a week later, causing hostility scores to vary widely (Heyman, 2001). Establishing
costly. Because of this, interrater reliability—how much different observers agree—is often
more important.
Two main factors affect interrater reliability: task complexity and observer training.
Simplifying the task usually improves reliability—for instance, using a 15-category coding
system instead of 100 categories helps observers agree more easily. Training is also key; if
observers don’t have clear definitions for behaviors (like what counts as laughter), their
ratings can differ a lot. Observers tend to be more careful during training and practice, but
may become less attentive when collecting real data if they feel they aren’t being monitored
Modern, well-designed observational systems with trained observers can achieve high
reliability, often with coefficients above 0.80 or 0.90 (Antony & Barlow, 2010; Harwood,
Beutler, & Groth-Marnat, 2011). Less structured observations usually have lower reliability.
At first, observing behavior seems like the most valid way to assess clinical issues.
For example, if we see aggression in a couple, aren’t we directly measuring aggression? The
answer depends on three things: (a) whether the behaviors being coded (like yelling) truly
represent aggression, (b) whether the data accurately capture how much aggression is
happening during the observation, and (c) whether the client’s behavior in the observation
conclusions from other methods. For example, if people rated as assertive by their peers also
show more refusal of unreasonable requests during observation, then the observation
correlates with peer ratings and shows convergent validity. The stronger the correlation
between observational data and other sources—like interviews, physiological tests, self-
reports, or others’ opinions—the higher the convergent validity (e.g., Andersson, Miniscalco,
with brain injury can drive safely, it shows predictive validity (Lew et al., 2005).
Observations that focus on clearly defined behaviors and sample repeated actions in realistic
settings tend to have better predictive validity. For example, one study found that
observations of preschool children’s behavior predicted their future behavior better than
Clinicians now have many observational assessment tools to choose from. Many are brief,
focus on specific symptoms or behaviors, and have decent reliability and validity. However,
many tools still lack thorough validation. Validating these instruments can take a long time
because there is no absolute “gold standard” to compare against (Lang & Kleijnen, 2010).
Because of this, clinicians and researchers can sometimes feel unsure about which tools are
Psychological Testing
Historically, psychological testing was the main, and sometimes the only,
responsibility of clinical psychologists. Although their roles have expanded to include other
(Lubin & Lubin, 1972; Wellner, 1968). Psychologists originally entered the clinical field due
to their expertise in developing and applying tests, which remains a major contribution to
In clinical practice, test results are combined with information from interviews,
observations, and other assessments to guide diagnosis, treatment planning, and evaluation.
Tests provide standardized, objective measures that help clinicians understand patients, plan
interventions, track changes over time, and compare individuals. While debate persists about
the appropriate role of tests in clinical decision-making, effectively used test data can
improve the quality of care. Beyond immediate clinical benefits, examining test behavior
such as evaluating the effectiveness of a treatment. This research can use existing
standardized tests or require the creation of new measures. In either case, a solid
understanding of testing principles is crucial. Psychological tests are therefore essential tools
in clinical practice, research, and training, and clinicians must be skilled in their use while
psychometric theory, general testing methods, and resources focused on personality and
and statistics, students should appreciate the challenges addressed by researchers like Ghiselli
(1964) and Nunnally (1967). For broad overviews of psychological testing, Cronbach (1970)
and Anastasi (1968) provide detailed discussions on test development, application, and
evaluation across domains such as intelligence, abilities, and personality. Reference works
II (1974), and Personality Tests and Reviews, II (1975)—offer critical reviews of hundreds of
tests and are regularly updated. Before administering any test, students and clinicians should
consult these reviews along with the test manual prepared by the test’s author. The APA’s
Standards for Educational and Psychological Tests and Manuals (1974) also provides a
concise overview of test development, distribution, use, and evaluation. Further surveys of
clinical assessment methods can be found in Holt (1968), Lanyon and Goodstein (1971), and
Kleinmuntz (1967), with Holt’s work reflecting the ego-psychological tradition of Murray
and Rapaport. More detailed discussions in this tradition are available in Rapaport, Gill, and
Mayman (revised by Holt, 1968), Schafer (1948, 1954), and Allison, Blatt, and Zimet (1968).
To apply tests effectively, students must engage with this literature and gain practical
experience through observation and supervised practice. Being tested themselves helps
students understand the client’s perspective and fosters empathy. Experienced clinicians
obtain richer information from testing than novices or literature alone would suggest. They
develop personal standards, sensitivity to subtle cues, and awareness of their influence on the
testing process. They learn when to modify procedures, generate and test hypotheses during
assessment, and adjust interpretations according to the case. Expert practitioners, such as
psychotherapists “listen with a third ear.” However, expertise can also lead to overreliance on
a particular test, potentially overlooking more suitable alternatives. To avoid this, test experts
must continuously study their methods, engage in research, compare interpretations with
Psychological Test
a person’s behavior in order to gain insights into their abilities, characteristics, or functioning.
While the concept of a “test” is familiar from everyday experiences such as school exams,
driving tests, or job skill assessments psychological tests specifically focus on measuring
or attitudes.
These tests often produce numerical scores, such as an IQ of 110 or a percentile rank,
which indicate how an individual compares to others on a given trait. In some cases, tests
classify individuals into categories, such as color-blind versus not, or psychotic versus non-
psychotic. Results may appear as a single score, a profile of multiple scores, or qualitative
observations requiring expert interpretation. For example, the Wechsler Adult Intelligence
Scale (WAIS) provides an overall IQ, detailed subscale scores, and specific item responses
hospitals. More recently, there has been interest in assessing the psychological qualities of
allow many people to be assessed at the same time and are usually self-administered. Some
tests are specifically designed for group use, while others originally created for individual
administration have been adapted for group settings. For example, the Rorschach and
Thematic Apperception Test (TAT) are traditionally given one-on-one to allow for discussion
and exploration, which helps produce meaningful responses. However, they can also be
administered to groups by projecting the stimuli and having participants record their answers,
though this approach sacrifices some detail in favor of saving time and reducing costs. Tests
like the Minnesota Multiphasic Personality Inventory (MMPI) work equally well in both
formats because they rely less on interpersonal interaction. In clinical contexts, group testing
is typically used for screening or research purposes, while in educational, vocational, and
Tests vary in the materials used and the tasks required. Performance tests generally
puzzle, or tying shoelaces. In contrast, verbal tests require language-based responses. While
all tests involve some form of performance, the term “performance test” is typically reserved
Tests also differ in how specific and limited the tasks are. Structured tests, like
multiple-choice exams, have clearly defined questions and responses, making them easier to
score and standardize. Unstructured tests, such as essay exams or open-ended prompts, give
the participant more freedom to respond in their own way. While unstructured tests can reveal
more about a person’s unique concerns, they are harder to evaluate objectively. The term
stimuli. These are considered “performance measures” in the sense that the individual’s
behavior during the task is analyzed, but they rely on indirect indicators rather than direct
questioning.
Direct vs. Indirect Measures
Some tests use direct indicators tasks that have an obvious link to the skill or trait
being measured. For instance, vocabulary tests directly assess word knowledge, and timed
typing tests directly measure typing ability. Other tests use indirect indicators, where the
relationship between the task and the trait is less obvious. For example, Rorschach responses
are interpreted for what they reveal about personality, not for the literal content of the
answers. Indirect measures can limit a person’s ability to deliberately control how they are
perceived. Interestingly, even self-report tests can be indirect if items are empirically linked
to traits in ways not apparent from their face value, as is the case with some MMPI items.
According to Cronbach (1970), tests can also be classified based on whether they
determine how well a person can function under optimal conditions common in intelligence,
ability, and achievement testing. These tests provide clear instructions, encourage doing one’s
best, and use objective scoring criteria. Typical-performance tests, on the other hand, measure
what a person usually does, such as their personality traits, interests, or attitudes. These do
not have right or wrong answers, and the goal is not to “score high” but to reveal consistent
Personality assessments can be created using three main approaches: (1) Rational-theoretical,
(2) Empirical, and (3) Internal consistency (Lanyon & Goodstein, 1971). These methods are
not mutually exclusive—some tests may incorporate elements from all three but each reflects
In the rational-theoretical approach, test items are selected based on logical reasoning
or theoretical assumptions about the trait being measured. In ability testing, the relationship
between the task and the construct is usually direct, such as measuring typing proficiency by
having a person type a standard passage or assessing driving skill through a simulator. In
personality testing, items may explicitly describe relevant traits (e.g., “I am often frightened”
the Thematic Apperception Test (TAT) is based on the assumption that individuals project
their needs, feelings, and expectations onto ambiguous images when constructing stories,
while the Blacky Pictures Test uses images reflecting Freud’s psychosexual stages to identify
conflict types. Overall, rational-theoretical tests rely primarily on the developer’s judgment
whether based on common sense or formal theory rather than on statistical analysis for item
selection.
Empirical Approach
The empirical method selects test items based on their demonstrated ability to
distinguish between specific groups, regardless of whether the content appears logically
related to the construct being measured. For example, to predict leadership potential,
researchers might compare responses from individuals known to be leaders, such as club
presidents or team captains, with those from non-leaders, retaining only the items that
consistently differentiate the two groups. A classic example is the Minnesota Multiphasic
Personality Inventory (MMPI), in which items were included solely for their ability to predict
group membership. While this approach can produce highly predictive measures, it has the
limitation that some items may appear unrelated to the intended trait and could correlate with
the criterion group merely by chance; therefore, cross-validation with new samples is
essential.
The internal consistency method uses statistical techniques, most commonly factor
analysis, to identify patterns within a large set of items. In this approach, a broad pool of
items often selected through rational methods is administered to a large group, and responses
are analyzed to determine which items cluster together, indicating that they measure the same
underlying construct. These item clusters, or factors, are then interpreted as psychological
variables. For example, Cattell’s Sixteen Personality Factor Questionnaire (16PF) was
Similarly, factor analysis has been applied to existing instruments, such as the Minnesota
including extraversion and neuroticism, which align with Eysenck’s core trait model.
Measurement of Intelligence
Tuckman (1975) noted that the history of modern testing is essentially the history of
testing intelligence or mental ability. The term mental test was first introduced by J.M. Cattell
in 1890. Cattell, an American psychologist who studied with Wilhelm Wundt in Germany and
Francis Galton in Britain, was a strong advocate of what became known as the “brass
sensory thresholds and reaction times. This approach assumed that keen sensory abilities
were central to intelligence. However, Cattell’s own student demonstrated that sensory test
results, such as reaction times or color naming, had little to no relationship with academic
performance. Consequently, psychologists abandoned these measures as indices of
intelligence.
Over the years, many attempts have been made to define and measure intelligence,
but there remains no single universally accepted definition. This has made it important to
implying that administering such a test means we accept its definition. While this practical
approach avoided theoretical disputes, it faced issues—such as the vast variety of tests
available, cultural biases, and the lack of a common meaning across societies—which would
described it as the ability to reason with abstract ideas, understand well, maintain clear goals,
purposeful thinking, and self-corrective judgment. Binet saw intelligence as a single but
complex mental process measurable through diverse mental tasks. His tests, later refined by
others, grouped tasks by age levels and introduced the concept of the intelligence quotient
states that all mental activities share a general factor (g) along with specific factors (s) unique
to each activity. The g factor represents the core of intellectual ability and is more important
than the s factors, as it predicts performance across various situations. Tests like the Raven
Progressive Matrices and Cattell’s Culture Fair Intelligence Test are designed to measure g. A
major challenge to Spearman’s theory arose from findings that dissimilar tests could show
higher-than-expected correlations, suggesting overlapping abilities rather than purely separate
factors.
distinction is between individual tests and group tests. An individual intelligence test is
administered to one person at a time, such as the Binet-Simon Scale. A group intelligence test
is designed for simultaneous administration to multiple people. Group tests became popular
during World War I with the development of the Army Alpha and Beta tests the Alpha being
Individual tests are more suitable for clinical settings, schools, military services, and
situations requiring careful, detailed examination. They allow the examiner to observe the
examinee closely, adapt instructions if needed, and evaluate specific responses. They are
especially useful for young children, individuals with disabilities, or those who cannot be
Group tests are typically used in educational, vocational, and large-scale assessment
contexts where efficiency is a priority. These tests are standardized, economical, and faster to
administer, making them suitable for screening large populations. However, they limit
In terms of construction, individual tests are more challenging and costly to develop
due to the need for specialized materials and trained administrators. Group tests, on the other
hand, can be prepared for large-scale use with relatively fewer resources. Norms for group
tests tend to be more reliable because they are based on large samples, but individual tests
written responses and language comprehension can be unsuitable for individuals unfamiliar
with the test language or those with limited literacy skills. In such cases, non-verbal group
tests like the Army Beta, Raven Progressive Matrices, or Cattell’s Culture Fair Test—are
preferred.
Culture-free and culture-fair tests aim to minimize the effects of cultural background
on performance. While culture-free tests attempt to remove cultural influence entirely (often
using abstract, non-verbal tasks), this is difficult to achieve completely. Culture-fair tests,
such as Cattell’s, focus on minimizing bias through the use of abstract problem-solving and
Binet’s work in France, beginning in 1889. Assigned by the French Government to create a
method for identifying children with learning difficulties, Binet collaborated with physician
Simon at the asylum of Saint-Yon to produce what became the Binet–Simon Scale in 1905.
This scale contained 30 items graded by difficulty to assess the mental development of school
children. In 1908, the scale was revised to extend the age range and to create ordered item
sets, sparking interest in countries such as Germany, England, Belgium, Switzerland, Italy,
and the USA. Following feedback, Binet and Simon further revised the scale in 1911,
In the USA, several adaptations emerged, the most notable being at Stanford University.
Earlier translations included Goddard’s 1908 English version of the 1905 scale and revisions
by Kuhlmann (1912), Yerkes (1915, 1923), and Terman (1916). Terman’s revision, called the
Stanford–Binet, introduced the concept of the Intelligence Quotient (IQ) and became widely
used. The 1937 revision offered two equivalent forms (L and M), extending the age range and
improving standardization. The 1960 revision combined these into a single form (L–M),
retained core features of the 1937 version, eliminated weaker items, and covered ages 2 to 22
years and 11 months. It measured seven ability areas: language, reasoning, memory, social
Administration time was about one hour, with scoring expressed as a deviation IQ.
A further revision in 1972 provided updated normative data based on a sample of 2,100
children, making it the first version to include nonwhite individuals. The fourth edition, SB4
(1986), represented the most comprehensive revision, incorporating the gf–gc theory of
intelligence (Horn & Noll, 1997). This theory distinguishes between fluid intelligence (gf)—
the capacity to solve novel problems and acquire new knowledge—and crystallized
model:
1. Top level – General intelligence (g factor) representing shared variance across all
tasks.
memory.
format, starting with items matched to the subject’s estimated ability (based on age or
vocabulary performance) and branching into progressively harder or easier items depending
on responses. This ensures that each individual is assessed appropriately without unnecessary
In Binet’s modern scale, the g-factor (general intelligence) sits at the top of the model,
representing the shared variance across all tasks. Below it are group factors, including
crystallized abilities (knowledge and skills acquired through learning), fluid–analytic abilities
(capacity to solve novel problems), and short-term memory (ability to retain and use
information briefly). Crystallized abilities are divided into verbal reasoning and quantitative
measured by four tests, quantitative reasoning by three tests, abstract/visual reasoning by four
The modern scale provides both an overall g score and specific scores for content areas. To
avoid uneven item distribution, the age-scale format was replaced with separate content-
based tests. For example, all vocabulary items are grouped into one test, while all matrix
items form another. Each subtest belongs to one of four main areas:
The basal age is the lowest point where two consecutive items of similar difficulty are
passed, while the ceiling is reached when at least three of four items are failed. This approach
maintains age differentiation, meaning that item difficulty increases in alignment with
developmental ability. Scores from the 15 tests can be converted to standard age scores (mean
= 50, SD = 8) and grouped into the four content areas (mean = 100, SD = 16).
with strong standardization, clear scoring procedures, and applicability across normal, gifted,
children.
Challenging administration – since items are scored as the test progresses, making it
The fifth edition (SB5), introduced in 2003, reviewed items for gender, ethnic,
cultural, and socio-economic fairness. It produces Full Scale IQ, Verbal IQ, and
Nonverbal IQ, and uses a 15-subtest structure. A key update was adjusting the standard
deviation to 15, aligning it with other major IQ tests, enhancing comparability across
measures.
Wechsler Scale
intelligence, David Wechsler, working at Bellevue Hospital in New York, created a new
intelligence test in 1939 for individuals aged 10 and older. This was called the Wechsler–
Bellevue scale, initially consisting of two forms (Form I and II), each with 10 subtests—5
In 1955, this scale was revised and renamed as the Wechsler Adult Intelligence Scale
(WAIS), containing 11 subtests (6 verbal and 5 performance). It was updated again in 1981
(WAIS-R) and then in 1997 (WAIS-III).The WAIS-III includes 14 subtests, 7 verbal and 7
Comprehension Judgment
Each set of seven subtests (verbal and performance) forms the verbal scale and
performance scale, respectively. Each subtest yields a raw score, which can be converted into
standard/scaled scores (mean of 10, SD of 3). Age-adjusted and reference-group norms are
respective subtests are summed. Both are deviation IQs with a mean of 100 and SD of 15. A
In addition, WAIS-III provides four index scores for a more detailed assessment:
These indices help identify specific cognitive strengths and weaknesses. For example,
verbal comprehension relates to language skills, perceptual organization relates to spatial and
reasoning abilities, working memory concerns temporary mental storage, and processing
Lastly, WAIS-III allows for pattern analysis, which compares performance across
different subtests. This can help identify conditions like schizophrenia, where there's typically
Kaufman Scales
The Kaufman Scales, developed during the 1980s and 1990s, were designed to
incorporate modern developments in test construction and are administered individually for a
range of purposes. The three main assessments within this series are the Kaufman Assessment
Battery for Children (K-ABC), the Kaufman Adolescent and Adult Intelligence Scale
(KAIT), and the Kaufman Brief Intelligence Test (K-BIT). The K-ABC is intended for
children aged 2½ to 12½ years and consists of 16 subtests grouped into five major scales:
measures intelligence through mental processing abilities, with a focus on sequential and
simultaneous thinking.
of brain functioning, Sperry’s split-brain research, and Neisser’s cognitive processing theory.
mentally organizing information in a specific order and solving problems step-by-step, while
mental processing abilities, and its nonverbal scale is particularly useful for assessing
children who are linguistically diverse or have language impairments. Scores from the
subtests are standardized with a mean of 10 and a standard deviation of 3, and can be
converted to a global score with a mean of 100 and standard deviation of 15, along with
percentile and age norms. Reliability and validity are generally strong, with split-half
Despite its strengths, the K-ABC has been criticized. Jensen (1984) argued that it is
less predictive of school achievement than the Binet and Wechsler scales and places
excessive emphasis on rote learning. Other critiques point to theoretical limitations, over-
In 2004, the second edition, KABC-II, was released to assess cognitive abilities in
children and adolescents aged 3 to 18 years. Unlike its predecessor, KABC-II offers two
Luria’s neuropsychological model of processing. It includes 18 subtests divided into core and
supplemental types. Under Luria’s model, the scales include simultaneous processing,
sequential processing, learning ability, and planning ability. In the CHC model, simultaneous
memory (Gsm), learning ability is redefined as long-term storage and retrieval (Glr), and
planning ability is identified as fluid reasoning (Gf). The CHC model also adds a fifth scale,
Psychologists have developed many nonverbal intelligence tests that can be given to
individuals or groups and used across different cultures. Two well-known examples are
people aged 5 to older adults. It consists of 60 multiple-choice items arranged from easy to
difficult. The test measures reasoning ability, especially Spearman’s “g” factor of intelligence,
by asking test-takers to identify the missing piece in a pattern of shapes or designs. The
designs may require recognizing simple shapes or solving more complex problems involving
analogies, patterns, and logical relationships. RPM has no time limit. It comes in three
versions:
CPM (Coloured Progressive Matrices) for young children or special needs groups
The RPM has high reliability and validity. Test–retest reliability scores generally
range from .70 to .90, and internal consistency is usually in the .80s and .90s. It tends to be
more accurate when compared with performance-based tests than with verbal tests.
This is another nonverbal test of intelligence where the person is asked to draw a man
as accurately as possible. First developed in 1926 and revised in 1963, it now includes three
scales: the Man Scale, the Woman Scale, and the Self Scale. The Woman Scale requires
drawing a woman, while the Self Scale involves drawing oneself (often used like a
personality test). Each drawing is scored based on details included, with a total possible score
of 70 points. The scores are converted into standard scores (mean 100, SD 15). It can be used
for children aged 3 to 15 years 11 months, but is best suited for ages 3–10. The test has
adequate reliability and validity, with correlations to other intelligence tests usually
above .50.
Although originally made for cross-cultural use, these tests are also widely used in
clinical and counselling contexts. Some were adapted from military group intelligence tests
such as the Army Alpha and Beta tests during World War I. The Alpha test measured verbal
skills, while the Beta test measured nonverbal ability for people with low literacy or non-
intelligence across a wide range of grade levels, from kindergarten through the twelfth grade.
It contains several verbal and nonverbal items suitable for both young children and older
students, including those who may be handicapped in verbal expression. The results of the
KAT can be expressed as verbal, quantitative, and total scores, and may also be reported in
percentile bands or as deviation IQs. The test is known for its strong psychometric properties,
with split-half and test-retest reliability coefficients in the low .90s and validity coefficients
ranging from the low .80s to the high .90s. These values indicate that the KAT is an
extremely sound and reliable measure of group nonverbal intelligence, correlating highly
The Henson–Nelson Test is another widely used group mental ability test designed for
all grade levels. It provides a single score that is believed to reflect general intelligence. The
test consists of 90 items and takes approximately 30 minutes to complete. It has provisions
for raw scores that can be converted into deviation IQs as well as percentiles, based on
distributions by age and grade. The test has demonstrated strong reliability, with split-half and
test-retest coefficients typically in the .90s, and validity coefficients ranging from .50 to .84.
The primary advantage of the Henson–Nelson Test is that it can predict academic success
effectively, although it primarily measures a single factor of intelligence and may not assess
The Cognitive Abilities Test (CAT) is another important group intelligence test that
particularly suited for assessing the intelligence of well-educated individuals as well as those
for whom English is a second language. The test includes extensive practice exercises and
two or three days. Reliability coefficients for the CAT are generally high, with the
quantitative and verbal subtests showing reliability in the low .90s and the nonverbal test in
the high .90s. Despite its strengths, some researchers have noted that the CAT’s effectiveness
depends on the degree to which its norms are representative and that more supportive data are
Many types of intelligence tests have been developed in India over the years,
reflecting the nation’s growing contribution to the field of psychometric assessment. One of
the pioneers in this area was Professor S. M. Mohsin, who constructed an intelligence test in
Hindustani in the late 1930s. In the first Mental Measurement Handbook for India, Mohsin’s
work was documented with a list of 103 tests designed to measure intelligence in different
Indian languages. His tests have been recognized for their reliability, validity, and satisfactory
functioning under the National Council of Educational Research and Training (NCERT), has
played a significant role in documenting and publishing Indian intelligence tests through its
Several popular Indian intelligence tests have emerged over the years. These include
the General Mental Ability Test for Children developed by R. P. Shrivastava and K. Saxena,
the Group Test of Intelligence by P. Ahuja, and the Group Test of General Intelligence by G.
C. Ahuja. Other significant contributions include the Verbal Intelligence Scale by R. K. Ojha
and Ray Chaudhary, the Test of General Intelligence for College Students by S. K. Pal and K.
S. Mishra, and the Group Test of Intelligence by R. K. Tandon. Additionally, the Emotional
Intelligence Scale by A. Hude, S. Pethe, and Upindhar Dhar, the Social Intelligence Scale by
N. K. Chadha and Usha Ganesan, and the Indian Child Intelligence Test by Usha Khire have
added to the diversity of intelligence assessments available in India. Most of these tests are
intelligence tests in India. He reported that around 70% of the tests were constructed and
standardized during the 1960s and 1970s, marking a period of active growth in indigenous
test development. However, a noticeable decline was observed during the 1980s. Among
these tests, 67% were verbal, 23% were nonverbal, and 10% were performance-based.
Furthermore, about 51% of the tests were developed in Hindi, emphasizing their applicability
to Indian linguistic and cultural contexts. These tests, primarily group-administered, were
often designed for the assessment of children and adolescents aged between 13 and 17 years,
making them suitable for educational and psychological evaluation in the Indian context.
The terms intelligence, aptitude, and achievement are often used in psychology and
education, but it’s important to distinguish between them. Intelligence is a general ability,
while aptitude is a potential for learning or developing skills in a specific area, and
achievement reflects what a person has already learned or accomplished. Aptitude refers to
the ability to acquire or develop a skill (e.g., learning a language, playing an instrument),
Aptitude tests are designed to predict future performance in a particular skill or field,
while achievement tests measure what has already been learned. Intelligence tests and
aptitude tests both look at potential, but achievement tests are more focused on past learning.
Psychometricians also use the term “ability” as a broad category that covers both aptitude and
achievement.
In practice, the boundary between aptitude and achievement tests is not always clear,
but research suggests that aptitude tests generally predict future performance better, while
Aptitude Tests
Aptitude tests can be divided into multiple aptitude tests and special aptitude tests.
Multiple aptitude tests assess a variety of abilities through different subtests, while special
aptitude tests focus on one specific skill. Early aptitude tests were based on general
The Differential Aptitude Test, first published in 1947, is one of the most widely used
multiple aptitude batteries. It has been revised several times, with the fifth edition released in
1992. The DAT includes eight subtests measuring verbal reasoning, numerical ability, abstract
reasoning, mechanical reasoning, clerical speed and accuracy, spatial relations, spelling, and
language usage. It is mainly used for vocational guidance for students from grade 8 upward.
Test results are given as percentile ranks, and a composite score (verbal + numerical
reasoning) is often used as an index of general scholastic ability. The DAT has also been
The General Aptitude Test Battery was developed by the U.S. Employment Service in
1962 for use in the armed forces and other fields. Based on factor analysis, it originally
measured 12 aptitudes, later reduced to 9: general learning ability, verbal aptitude, numerical
aptitude, spatial aptitude, form perception, clerical perception, motor coordination, finger
dexterity, and manual dexterity. The 12 tests include both verbal and nonverbal items, some
The Flanagan Aptitude Classification Test was developed for vocational counseling
and employee selection. It originated from research into job elements that distinguish
between successful and unsuccessful workers. Out of 21 identified job elements, 19 tests
were developed, with 2 more still in progress. The complete battery requires three testing
sessions and takes more than two and a half hours to administer.
has ten subtests such as general science, arithmetic reasoning, word knowledge, and
composites measure verbal, mathematical, and general ability, while occupational composites
measure skills for mechanical, clerical, electronics, and health/social fields. The ASVAB is
known for high reliability and validity and helps match individuals to suitable training
programs.
Although less common today, special aptitude tests still exist because they offer more
flexibility in measuring specific skills not fully covered by large aptitude batteries. Examples
include tests for artistic, musical, creative, vision, hearing, and motor skills.
Sensory Tests
Early sensory tests, such as those by Galton, aimed to measure sensory abilities like
vision and hearing but initially failed to connect with intellectual performance. Modern
sensory tests now measure dimensions like visual acuity (near and far), depth perception, and
muscle control of the eyes. Common tools include the Snellen chart and instruments like the
Ortho-Rater and American Optical Sight Screener, used for school screenings and job
selection. Hearing tests measure auditory acuity using pure-tone audiometers to detect sound
at different frequencies and directions. The Massachusetts Hearing Test allows testing of
These measure coordination of hand, arm, and/or leg movement for task performance.
Common examples include the Crawford Small Parts Dexterity Test, Stromberg Dexterity
Test, Purdue Pegboard, Bennett Hand Tool Dexterity Test, and the Complex Coordination
Test.
Artistic Aptitude Tests
Artistic aptitude tests are designed to measure an individual’s artistic ability. These
assessments often focus on two aspects artistic appreciation and productive ability. Tests of
artistic appreciation evaluate the ability to recognize and discriminate between artistic
qualities, while tests of productive ability measure the skill to create art. For example, a
person might be able to judge art effectively but lack the skill to paint.
One historically significant example is the McAdory Art Test (1929), though it is now
obsolete. The Meier Art Judgement Test (developed in 1929) is a well-known measure of
artistic appreciation, assessing traits such as manual skill, aesthetic intelligence, imagination,
and perceptual ability. In this test, participants are shown two similar paintings—one by a
professional artist and another altered in aspects like symmetry, shading, or color use—and
Another example, the Graves Design Judgement Test (1948, revised 1951), asks
participants to choose the most appealing design from sets containing an original and
variations. The Horn Art Aptitude Inventory measure’s productive artistic ability, such as
arranging common objects or geometric figures in aesthetically pleasing ways. The test
includes two timed parts: arranging geometric shapes into appealing designs and selecting the
Some other well-known measures of artistic ability include the Knapp Art Ability Test
and the Leverenz Tests for assessing core visual art skills.
For nearly four decades, the University of Iowa, under psychologist Carl Seashore,
conducted extensive research on measuring musical aptitude. This led to the development of
the Seashore Measures of Musical Talents by Seashore, Lewis, and Saetveit (1939). The test
evaluates aspects such as pitch, loudness, rhythm, time, timbre, and tonal memory. Designed
for students from grade 4 onward, it involves comparing pairs of tones and making judgments
about differences in their characteristics. For example, in the pitch test, the examinee decides
whether the second tone is higher or lower than the first, with difficulty increasing as pitch
differences narrow. In the rhythm test, examinees decide whether rhythmic patterns are
identical or different. The time test measures the ability to compare tone durations. Timbre
involves distinguishing between tonal qualities of sounds, and tonal memory tests memory
While results are reported for each subset and converted into percentile scores,
psychologists note that the test may yield limited meaning for children under age 10 and
primarily measures sensory discrimination rather than complete musical aptitude (Nunnally,
1970).
Another approach, the Wing Standardized Test of Musical Intelligence (Wing, 1941,
1962), is widely used by music teachers because it emphasizes skills needed for musical
training. This battery includes chord analysis, pitch discrimination, tuning, harmony, rhythm,
intensity, and phrasing. The first three subtests require complex sensory discrimination, while
the later tests measure aesthetic judgment of musical quality. This test takes about an hour but
Although mechanical and clerical aptitudes are often included in larger multiple
aptitude batteries, they can also be assessed through independent tests. This is important
because these areas are essential for specific occupations and cannot always be given full
comprehension. Examples include the Minnesota Paper Form Board Test, the Space Relations
Test, the Differential Aptitude Test (DAT) Mechanical Reasoning subtest, the SRA
Clerical aptitude tests generally assess perceptual speed and accuracy. Examples
include the Clerical Speed and Accuracy Test of the DAT and the Minnesota Clerical Test,
both of which require comparing letters or numbers and identifying matches. These tests
usually have two parts: a practice trial and a timed trial, with the goal of completing as many
correct matches as possible within a time limit. The Minnesota Clerical Test also includes
tasks such as the Number Comparison and Name Comparison tests, which involve quickly
Achievement Tests
Achievement tests, also referred to as proficiency tests, are widely used in psychological
and educational research. Before examining their various types, it is helpful to understand
adjust their instruction to improve outcomes, while learners may be motivated to work
educational content and teaching methods. They help determine whether the material
is effective and relevant, identify outdated content, and suggest new topics. They also
provide critical feedback that can help make instruction more engaging and aligned
strengths and weaknesses. Instruction can then be modified to fit these needs, such as
Like aptitude tests, achievement tests can be divided into general achievement batteries
and special achievement tests. General batteries assess overall educational achievement from
primary to secondary school and cover various skills, such as arithmetic, spelling, reading,
and map reading. Examples include the Iowa Tests of Basic Skills, the California
Achievement Tests, and the SRA Achievement Series. Many also form part of national
programs like the Adult Basic Learning Examination (ABLE), which is aimed at
Special achievement tests focus on specific subjects or skills, often for diagnostic
purposes. For example, the Stanford Diagnostic Reading Test identifies reading difficulties
and is available in different versions for various grade levels. The Durrell Tests measure
specific reading skills like comprehension, word recognition, and reading rate, while the
These standardized tests are often paired with end-of-course examinations, which
provide a coordinated measure of student performance across different subjects. This allows
Essay-type tests require relatively free or extended responses to questions that assess
experience. They are often used to evaluate student achievement in the classroom and can
way. They are generally easy for teachers to prepare and administer.
However, the use of essay tests as achievement measures has declined because of
drawbacks such as the difficulty of ensuring objectivity, time-consuming scoring, and the
limited number of questions that can be included. As a result, many essay tests are being
replaced by standardized achievement tests with objective items that can be scored quickly
and consistently. The main distinction between essay and achievement tests is that
achievement tests may also be used for purposes beyond classroom evaluation, such as
Although achievement tests are valuable tools for assessing performance, they have
certain limitations:
1. Scores from achievement tests cannot always be used to judge a student’s potential
because they may not reflect all factors influencing classroom performance.
2. Achievement tests, unlike intelligence and aptitude tests, are more difficult to
construct.
achievement.
Tests of Creativity
Creativity tests measure divergent thinking the ability to generate multiple solutions
to a problem from limited information contrasting with convergent thinking, which focuses
It has two sections verbal and figural. The verbal section includes three parts:
2. Product improvement – the examinee is shown an object and asked for ideas to
improve it.
3. Unusual uses – the examinee lists as many uses as possible for a given object.
The figural section involves drawing tasks where the examinee is given a simple shape
and must incorporate it into a creative drawing. Each response is scored for fluency (number
Another example is the Remote Associates Test (RAT) by Mednick (1971), intended
mainly for high school students. In this test, the examinee is given three words and must find
a fourth word that links them. The test emphasizes the ability to connect seemingly unrelated
ideas.
Personality Test
traits. Personality measurement aims to describe and study traits such as social traits (e.g.,
prestige), personal conceptions (e.g., values, attitudes, and self-perception), and adjustment
(e.g., emotional stability and the absence of maladaptive behaviors like hysteria or
psychoses). These traits, which are often interrelated, help explain a person’s temperament,
character, and ability to adapt, and can be organized into a few broad categories to better
their feelings, environment, and reactions to others. These can be categorized into five types:
those measuring specific social traits like confidence or extroversion; those assessing
adjustment to different life aspects; those identifying pathological traits such as depression or
paranoia; those screening individuals into categories based on certain conditions; and those
measuring attitudes, interests, and values. Despite the differences, all share the principle that
behavior reflects underlying traits, and the presence or absence of certain behaviors indicates
trait presence.
observation of individuals in set situations. Observers record behavior, and these observations
Projective techniques, the most widely used, involve presenting individuals with
ambiguous situations or stimuli and asking them to respond, revealing unconscious thoughts,
among the most common tools for measuring personality. They assess traits, types, states, and
related aspects such as self-concept. Personality traits indicate a person’s consistent patterns
of behavior, while personality types describe general categories of people, and self-concept
reflects a stable set of beliefs about oneself. In these tests, individuals respond to written
statements, often marking them as “True” or “False” to show if they apply. This structured or
objective method differs from projective methods because each item is clear and
unambiguous, with specific guidelines for responses. Research shows that constructing such
link test items with specific traits. For example, if a person marks “True” for “I like to
participate in social activities,” it is assumed they enjoy socializing. This approach, used to
assess face validity, was prominent before and after World War I. The first major personality
inventory, the Woodworth Personal Data Sheet (1920), aimed to identify recruits with
emotional issues through 116 “Yes/No” items. Its success inspired other inventories like the
Bell Adjustment Inventory and Bernreuter Personality Inventory, which measured areas such
as social life, health, and specific traits. While initially popular, the logical-content method
faced criticism for relying on unverified assumptions about honesty and interpretation.
Research revealed that such assumptions often lacked empirical support, leading to the rise of
statistical analyses and empirical data to select test items. It involves creating two groups: a
criterion group, made up of individuals with specific traits or conditions (e.g., hysteria,
schizophrenia), and a control group from the general population. Test items that best
distinguish between these groups are included in the test, though their face validity may be
low. These items are later cross-validated with a different criterion group to ensure reliability.
Scores from the control group are used to set standard values, enabling the comparison of
new test-takers’ results. After construction and validation, further research is done to interpret
the meaning of endorsed items. Examples of tests developed with this method include the
containing 550 items, it included ten clinical scales (e.g., Hypochondriasis, Depression,
Schizophrenia, Mania, and Social Introversion) and four validity scales (Lie, Infrequency,
Correction, and Cannot Say). The clinical scales identify various psychological disorders,
while the validity scales check for honesty, defensiveness, or test-taking issues. Over time,
MMPI was revised into MMPI-2 and MMPI-A (for adolescents). MMPI-2 contains 567
statements, retains the original clinical and validity scales, and adds new content and
supplementary scales, along with three new validity scales: the Back F (Fb) scale, Variable
Response Inconsistency (VRIN), and True Response Inconsistency (TRIN), which assess
consistency and care in responses. MMPI-2 is widely used, with over 10,000 studies
supporting its research and clinical utility. MMPI-A is tailored to adolescents, addressing
school, family, and developmental issues. Both versions have strong reliability, although
some problems include item overlap, high intercorrelation among scales, and limited
generalizability. Despite limitations, MMPI and MMPI-2 remain among the most popular
1957, is a self-report personality test originally containing 480 items, later revised to 434 in
its third edition. Unlike the MMPI, which focuses on identifying psychopathology, the CPI
targets normal personality traits in non-clinical populations and is based on the criterion-
group strategy. Many CPI items are adapted from the MMPI, but it uses them to assess
positive personality attributes. The CPI has 20 scales, including three validity scales—Well-
being (Wb), Good Impression (Gi), and Communality (Cm)—which help detect inconsistent
or socially desirable responding. The remaining 17 scales measure traits such as Dominance,
Sociability, Self-control, Tolerance, Achievement, and Flexibility. Like the MMPI, scores are
Research shows the CPI is effective for predicting academic performance, leadership,
and success in various professional roles. Reliability is comparable to the MMPI, with short-
term test–retest coefficients ranging from .49 to .90. However, it shares MMPI’s limitations,
dimensions. Pioneers like Guilford, Cattell, and Eysenck analyzed large sets of personality
data to extract core traits, leading to inventories such as Guilford’s Inventory of Factors
STDCR, which measured social introversion, thinking introversion, depression, and other
which measures ten dimensions including General Activity, Friendliness, Emotional Stability,
and Masculinity. Although innovative, this factor-analytic approach is less popular today,
Cattell’s Personality Questionnaire (16 PF). Raymond B. Cattell, along with his
collaborators, developed the Sixteen Personality Factor Questionnaire (16 PF) based on a
1949, has undergone several revisions and is currently available in its fifth edition. It is
designed to assess normal personality traits for individuals aged 16 and above. The
best representing their typical behavior. The 16 PF identifies 16 primary personality factors,
each representing a bipolar trait dimension (e.g., reserved vs. outgoing, trusting vs.
suspicious). These traits were derived through extensive statistical analysis to capture the
The latest edition contains 185 carefully selected items, including a set of 15 problem-
solving questions that form the Reasoning Scale, serving as a brief indicator of mental ability.
The 16 PF also includes indices to detect random responses and social desirability bias. Each
Scale B (Less Intelligent–More Intelligent) assesses reasoning ability; and Scale C (Stable–
Emotionally Unstable) evaluates emotional control and ego strength. Other dimensions such
Cattell also developed the Cattell Fair Intelligence Test to assess fluid intelligence
independent of cultural and educational influences. The 16 PF has shown high reliability and
validity, with test–retest coefficients ranging from .65 to .93 and internal consistency often
exceeding .80. For younger age groups, the Junior 16 PF and Children’s Personality
Overall, the 16 PF remains one of the most comprehensive and empirically validated
tools for assessing normal-range personality, offering insights valuable for clinical,
lifelong effort to measure both normal and abnormal aspects of personality through a factor-
personality — Psychoticism (P), Extraversion (E), and Neuroticism (N) — which Eysenck
“Yes” or “No” responses and includes a Lie (L) scale to check the honesty and validity of
responses. A Junior EPQ version is also available for children aged 7 to 15, containing 81
statements.
The P Scale evaluates traits associated with psychoticism, which is not equivalent to
egocentricity, and impulsivity. Individuals scoring high on this scale tend to show poor
concentration, lack of empathy, insensitivity, and disregard for social norms and conventions.
They are often seen as unconventional or even peculiar by others. Conversely, a low score
The E Scale measures extraversion and its opposite, introversion. High scorers are
typically outgoing, sociable, fun-loving, and seek excitement and social interaction. They
enjoy being around people and prefer active engagement. Low scorers, on the other hand, are
instability. High scorers on this scale tend to be anxious, easily upset, moody, and
overemotional, whereas low scorers are calm, stable, and emotionally balanced.
Research using the EPQ has extensively examined the biological and behavioral
correlates of these dimensions, particularly the link between extraversion and introversion.
Findings suggest that extraverts seek external stimulation, are more conditioned to sexual
arousal, and are more suggestible than introverts. In contrast, introverts show greater
Psychometric studies have established the EPQ as a reliable and valid instrument,
with one-month test–retest reliabilities of .78 for Psychoticism, .89 for Extraversion, .86 for
Neuroticism, and .84 for the Lie scale. The construct validity has also been supported through
Theoretical Stratergy
The theoretical strategy in personality testing involves selecting test items based on
predefined dimensions of personality and ensuring that each item aligns with the underlying
theory. This contrasts with factor-analytic approaches, which identify traits from statistical
patterns in large datasets. The theoretical strategy aims for a homogeneous scale by selecting
items that measure only one dimension, producing results consistent with the theoretical
framework.
Edwards Personal Preference Schedule. One early example of this method is the
Edwards Personal Preference Schedule (EPPS), created by Edwards (1954) using Murray’s
theory of personality needs (1938). The EPPS contains 15 primary needs, each represented by
14 paired statements. Test-takers choose the statement that best describes them, a method
called ipsative scoring, which compares the strength of each trait relative to the person’s other
Endurance, Heterosexuality, and Aggression, plus a Consistency scale that checks for
response stability. Raw scores are converted into percentiles for interpretation.
While the EPPS is valued for its theoretical grounding, forced-choice format, and
non-threatening nature (especially for college and general adult populations), it has
limitations. These include difficulty in comparing individuals due to its ipsative nature,
potential social desirability bias, and questions about the advisability of converting ipsative
The Personality Research Form (PRF) and the Jackson Personality Inventory
(JPI) were developed by Jackson (1967; 1976a, 1976b) based on Murray’s theory of needs.
These tests represent an effort to create structured personality measures using a theoretical
approach. While developing the PRF, Jackson had two major goals: first, to design a
scientifically sound tool for personality research, and second, to create a test that could assess
The PRF was originally developed in two parallel forms that differed in the number of
items and scales. The shorter forms (A and B) included 15 scales and 300 items, while the
longer forms (AA and BB) contained 22 scales and 440 items, including two validity scales—
Desirability and Infrequency—which detect response biases such as carelessness and social
endurance, exhibition, harm avoidance, impulsivity, nurturance, order, play, sentience, social
recognition, succorance, understanding, and desirability and infrequency. These scales were
designed to represent bipolar traits, meaning they measure both the presence and absence of a
colorfulness, whereas a low score suggests shyness and avoidance of social situations.
A revised version, known as the Form E of PRF, was later developed to improve the
test and extend its applicability to diverse populations beyond college students. Through
modern item-selection techniques, Jackson reduced the number of items from 440 to 352,
while retaining the same 22 personality scales. The PRF thus became a comprehensive and
psychometrically sound tool widely used in both research and applied psychology.
The Jackson Personality Inventory (JPI) was developed after the PRF with the aim of
providing a more practical and appealing personality assessment for both research and
applied settings. The JPI consists of 320 true-false items divided into 16 scales, each
containing 20 items. It was designed primarily for high school and college students, as well
as adults. The 16 scales measure traits such as breadth of interest, complexity, conformity,
Both the PRF and the JPI are known for their balanced nature in terms of true-false
keying and content coverage, minimizing response bias. They have been standardized on
large populations, and the construction procedures ensured a high level of reliability and
validity. Overall, these inventories represent significant advances in the structured assessment
developed by Isabel Briggs Myers and Katharine Cook Briggs (Myers, 1962; Myers &
McCaulley, 1985) and is one of the most widely used personality assessment tools in the
types, which proposed that individuals differ in how they perceive the world and make
decisions. The MBTI aims to assess these psychological preferences among normal
versus Introversion (I), which reflects how individuals gain and expend energy—either
through interaction with others or through solitude. The second dimension, Sensation (S)
versus Intuition (N), represents opposing ways of perceiving information, with “S” types
focusing on concrete facts and “N” types preferring abstract patterns and possibilities. The
third dimension, Thinking (T) versus Feeling (F), contrasts logical, objective decision-
making with value-based, empathetic reasoning. The final dimension, Judging (J) versus
Perceiving (P), indicates one’s orientation to the outer world—whether a person prefers
types—each represented by a four-letter code such as INTP, ENFJ, or ISFP. For instance, an
INTP type represents a person who is introverted, intuitive, thinking, and perceptive—
someone who is typically quiet, analytical, enjoys solving logical problems, and shows deep
interest in ideas and theoretical concepts. Each preference direction is also scored
The primary objective of the MBTI is to determine an individual’s position along the
functioning. It is based on the assumption that every person has inherent preferences in the
way they interpret and respond to the world. These preferences influence their values,
Over the years, numerous studies have supported the practical applications of the
MBTI in various fields. It has been extensively used to explore relationships between
personality and financial success (Mabon, 1998), leadership development (Fitzgerald, 1997),
communication styles (Loffredo & Opt, 1998), career counseling (McCaulley & Martin,
1995), and interpersonal effectiveness (Doerries & Ridley, 1998). The MBTI continues to be
understand their own personality patterns and how these influence their interactions, choices,
where individuals choose from a list of adjectives those that describe them.
One widely used checklist is Gough’s Adjective Checklist (ACL), containing 350
based on Murray’s needs theory and draw from several established personality frameworks,
such as Berne’s Transactional Analysis theory (1961, 1966) and Wells’s theory of creativity
and intelligence.
Another significant tool is the Student Self-Concept Scale (SSCS), which is grounded
in Bandura’s theory of self-efficacy. The SSCS assesses three primary domains of self-
concept academic, social, and self-image. Respondents rate not only how true a statement is
for them but also how important it is, reflecting their confidence in achieving certain traits or
goals.
and Whalen (1990), is another major tool for high school and college students. It explores the
global self-esteem scale along with six specific facet scales related to the social aspects of
self-concept.
Similarly, the Pier–Harris Children’s Self-Concept Scale consists of 80 self-
children. Another well-known measure is the Tennessee Self-Concept Scale, a simple paper-
A more experiential approach comes from Carl Rogers’ theory of self, using the Q-
Sort technique. In this method, individuals sort a set of self-descriptive statements into
categories ranging from “least descriptive” to “most descriptive.” This is done twice — once
to represent their real self and again for their ideal self. The degree of similarity between the
two sorts indicates self-esteem and self-acceptance, with larger discrepancies suggesting
Combination strategy
In modern times, test developers have adopted a combination strategy that merges
various methods to create structured personality assessments. Two widely recognized tools
that follow this strategy are the Millon Inventories and the NEO Personality Inventories.
personality disorders through the Millon Clinical Multiaxial Inventories (MCMI). First
published in 1977 as MCMI, it was revised over time to MCMI-II and MCMI-III. These
inventories are second only to the MMPI in assessing pathological traits and are known for
being more clinically practical. The MCMI-III, which contains 175 true/false statements, is
shorter and easier to administer than MMPI-2. Its primary purpose is to diagnose clinical
The first version, MCMI (1977), consisted of 24 scales focused on personality traits
and aimed to refine MMPI’s clinical application. Later revisions—MCMI-II and MCMI-III—
were made to align with changes in the DSM, including new categories and scales for
disorders such as depressive and post-traumatic stress disorders. The MCMI-III includes 24
clinical scales grouped into four main categories: Clinical Personality Patterns, Severe
Personality Pathology, Clinical Syndromes, and Severe Syndromes. Each of these includes
scales representing various personality and clinical disorders at different levels of severity.
Additional modifying indices, such as Disclosure, Desirability, and Debasement, help assess
Millon later expanded his work with two new instruments the Millon Adolescent
Clinical Inventory (MACI) and the Millon Index of Personality Styles (MIPS). The MACI is
designed for adolescents aged 13 to 19 and was developed from the earlier Millon Adolescent
academic guidance. The MIPS, on the other hand, evaluates personality in normal adults,
Overall, the Millon Inventories demonstrate strong reliability and validity, making
them widely accepted and psychometrically sound tools in both clinical and counseling
contexts.
(1992), is one of the most recent and significant personality assessment tools. The inventory
was created using factor analysis and grounded in theoretical concepts of personality
structure. It measures the Big Five personality dimensions — Neuroticism (N), Extraversion
(E), Openness to Experience (O), Agreeableness (A), and Conscientiousness (C). The term
“Big Five” highlights that each of these broad dimensions encompasses several specific traits,
Initially, Costa and McCrae (1992, 1995) focused on just three traits—Neuroticism,
Extraversion, and Openness—hence the title “NEO.” Later, they expanded the model by
including Agreeableness and Conscientiousness to align with the Five-Factor Model (FFM).
Each of these five domains is divided into six narrower facets, resulting in a total of 30 facets
(5 domains × 6 facets). Each facet is assessed using 8 items, producing a 240-item inventory.
Respondents rate each statement on a five-point Likert scale, indicating their level of
agreement or disagreement.
vulnerability.
and assertiveness.
mindedness.
The NEO-PI-R has demonstrated strong reliability and validity across various populations
and correlates well with other personality instruments such as Goldberg’s (1992) adjective
empirical, logical, and statistical foundations in test construction. Because of its wide
applicability and psychometric strength, it is highly regarded for evaluating personality traits
globally.
In India, a variety of self-report personality inventories have been developed, with some
Sohoni (1953) created a Temperament and Character Test for high school children,
which demonstrated a reliability coefficient ranging from 0.44 to 0.54 and a validity
Singh (1967) designed an Adjustment Inventory for college students, assessing five
domains — home, health, society, emotion, and education — using 102 Yes-No items.
Its internal consistency ranged between 0.92 and 0.94, and the validity coefficient
Bengalee (1964) developed the Multiphasic Personality Inventory, later known as the
assessed five areas of adjustment: personal and social adjustment, unhealthy parent
1. Parent Attitude Scale – to assess unhealthy parental attitudes, including five subscales:
Prasad (1974) created another adjustment inventory focusing on parental, home, social, and
by Singh & Bhanot, 2002). The inventory has 105 items, and its reliability coefficients range
from 0.70 to 0.89, while internal consistency values range from 0.55 to 0.84.
Mohsin & Hussain (1981) adapted the Bell Adjustment Inventory (Students’ Form) in
Hindi, with 135 items (revised to 124 in 1987). Its test-retest reliability ranged
between 0.70 and 0.92, and split-half reliability between 0.73 and 0.93.
for ages 11–20 years, with 70 items and a test-retest reliability of 0.73.
reliability coefficient of 0.87 for the Extraversion scale (E) and 0.82 for the
adaptation of personality inventories that are both culturally appropriate and psychometrically
Projective Test
the projective hypothesis, these techniques assume that people unconsciously project their
inner thoughts, emotions, and conflicts onto unclear stimuli. This allows psychologists to
1. Projective Hypothesis:
This central idea suggests that when individuals interpret vague or ambiguous stimuli,
their responses reveal their inner emotions, motives, and experiences. For instance, a
2. Unstructured Nature:
Because there are no fixed answers, the variety of responses helps reveal unique
3. Disguised Purpose:
The intent of these tests is often hidden from participants to ensure genuine,
feelings, making them useful for exploring inner conflicts and hidden emotional
7. Idiographic Approach:
comparing them with group norms. They are used to develop a personalized
psychological profile.
Although projective methods have been criticized for their low reliability and questionable
validity, they remain popular in psychoanalytic and humanistic approaches. Critics argue that
results can vary across examiners and lack empirical support. However, clinicians value these
methods for the rich, qualitative insights they offer into an individual’s unconscious world,
the most famous projective techniques. It involves showing individuals ten inkblot cards and
asking them to describe what they see. The responses are analyzed to uncover personality
traits, emotional patterns, and potential psychological issues such as anxiety, aggression, or
creativity.
Although early criticism centered on its poor reliability, lack of standardization, and
subjective interpretation, it remains a widely used clinical tool. To improve its scientific
credibility, John Exner developed the Comprehensive System (1968), which standardized
administration and scoring. Exner’s system introduced over 50 scoring categories, focusing
on response structure and content to form a structural summary. This method achieved higher
inter-rater reliability and became the most accepted system for the Rorschach. Despite
ongoing debates over validity and the complexity of scoring, the test continues to be valued
Morgan and Henry Murray (1935). It uses a series of ambiguous black-and-white pictures,
and participants are asked to create stories about what is happening in each scene. The stories
these factors, known as thema, explains how behavior is shaped by both inner drives and
external situations. The TAT was created to explore these dynamics and remains an important
sets for men, women, and children. Typically, 20 cards are selected for use. Participants
create stories describing what led to the scene, what is happening, and how it will end. The
examiner records the responses verbatim and may prompt for clarification. Interpretation
focuses on identifying the main character (hero/heroine), their motives, and the press
(situational forces). Repeated themes help reveal unconscious drives and personal conflicts.
Evaluation
Practitioners may use different subsets of cards, resulting in inconsistent results across
studies.
2. Differences in Administration and Scoring:
Reliability is low due to open-ended responses and varied scoring. Some research
supports its use for specific traits like achievement or affiliation, but overall validity
remains limited.
Interpretations can reflect the examiner’s own perceptions, making results subjective.
While the TAT’s use has declined, adaptations such as TEMAS (for minority groups),
Senior Apperception Test (for the elderly), and Children’s Apperception Tests remain
Projective Drawings
drawings to infer personality and emotional functioning. These tests gained popularity for
relationships, and personality traits. They are widely used in clinical and educational
from low reliability and poor validity. While general impressions of drawings may be
somewhat consistent, interpretations based on details like head size, pencil pressure, or eyes
Illusory Correlation
real evidence. Even trained psychologists have been found to fall into this trap, interpreting
exploratory and therapeutic contexts. They are often valued for the emotional and relational
dynamics and emotional expression rather than providing formal diagnoses. Tests like Kinetic
Family Drawing and Draw-A-Family remain common in family therapy, where they help
subjectivity, they offer invaluable insights into the unconscious aspects of personality that
structured and objective tests often fail to capture. Techniques such as the Rorschach Inkblot
Test, Thematic Apperception Test (TAT), and Projective Drawing Tests allow individuals to
express hidden emotions, internal conflicts, and underlying motives in a way that transcends
verbal communication. These tools are especially useful in clinical and counseling settings,
where understanding the client’s inner world is essential for diagnosis and therapeutic
intervention. Thus, while projective techniques should be used cautiously and supplemented
with other methods, they remain a vital component of comprehensive personality assessment,
Reference
Korchin, J. S. (2004). Modern clinical psychology: Principles of intervention in the clinic and
Kramer, P.G., Bernstein, A. D & Phares, V (2014). Introduction to Clinical Psychology (8th
Singh, A.K. (2019). Tests, measurements and research methods in behavioral sciences (6th