Psychological Assessment
Psychological Testing and History
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Testing and Assessment Assessor is the key to the process of selecting
o 1905: Alfred Binet published a test designed to tests and/or other tools of evaluation
help place Paris school children in appropriate Requires an educated selection of tools of
classes evaluation, skill in evaluation, and thoughtful
o Testing – refer to everythingꟷfrom organization and integration of data
administration of test to interpretation of the test Entails logical problem-solving that brings to bear
scores many sources of data assigned to answer the
▪ Once used to describe the group of screening referral question
individuals of thousands of military recruits o Test – measuring device or procedure
o Psychological Assessment – gathering and o Psychological Test – device or procedure
integration of psychology-related data for the designed to measure variables related to
purpose of making psychological evaluation psychology
▪ Educational – evaluate abilities and skills ▪ Content – subject matter
relevant in school context ▪ Format – form, plan, structure,
▪ Retrospective – draw conclusions about arrangement, layout
psychological aspects of a person as they ▪ Item – a specific stimulus to which a person
existed at some point in time prior to the responds overtly and this response is being
assessment scored or evaluated
▪ Remote – subject is not in physical proximity ▪ Administration Procedures – one-to-one
to the person conducting the evaluation basis or group administration
▪ Ecological Momentary – “in the moment” ▪ Score – code or summary of statement,
evaluation of specific problems and related usually but not necessarily numerical in
cognitive and behavioral variables at the very nature, but reflects an evaluation of
time and place that they occur performance on a test
▪ Collaborative – the assessor and assesee ▪ Scoring – the process of assigning scores to
may work as “partners” from initial contact performances
through final feedback ▪ Cut-Score – reference point derived by
▪ Therapeutic – therapeutic self-discovery and judgement and used to divide a set of data
new understanding are encouraged into two or more classification
▪ Dynamic – describe interactive approach to ▪ Psychometric Soundness – technical quality
psychological assessment that usually ▪ Psychometrics – science of psychological
follows the model: evaluation > intervention measurement
of some sort > evaluation ▪ Psychometrist or Psychometrician – refer to
o Psychological Testing – process of measuring professional who uses, analyzes, and
psychology-related variables by means of interprets psychological data
devices or procedures designed to obtain a o Achievement Test – measurement of the
sample of behavior previous learning
Testing o Aptitude – refers to the potential for learning or
Usually numerical in nature acquiring a specific skill
Could be individual or by group in administration o Intelligence – refers to a person’s general
Test administrators can be interchangeable potential to solve problems, adapt to changing
without affecting the evaluation environments, abstract thinking, and profit from
Requires technician-like skills in terms of experience
administration and scoring o Human Ability – considerable overlap of
Yield a test score or a series of test score achievement, aptitude, and intelligence test
Assessment o Structured Personality tests – provide
Answers the referral question through the use of statement, usually self-report, and require the
different tools of evaluation subject to choose between two or more
Administered individually alternative responses
Psychological Assessment
Psychological Testing and History
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Projective Personality Tests – unstructured, and b. extent to which they understand and agree to
the stimulus or response are ambiguous the rationale for the assessment
o Interview – method of gathering information c. capacity and willingness to cooperate
through direct communication involving d. amount of physical or emotional distress
reciprocal exchange e. amount of physical discomfort
▪ Panel Interview (Board Interview) – more f. alertness level
than one interviewer participates in the g. predisposed to agree or disagree when
assessment presented with stimulus statements
▪ Motivational Interview – used by counselors h. received prior coaching
and clinicians to gather information about i. portraying themselves in good or bad light
some problematic behavior, while j. “luckiness” or have “bad luck” on multiple-
simultaneously attempting to address it choice achievement test
therapeutically o Psychological Autopsy – on the basis of archival
o Portfolio – samples of one’s ability and records, artifacts, and interviews previously
accomplishment conducted with the deceased assess or people
o Case History Data – refers to records, who knew him or her
transcripts, and other accounts in written, o Other parties: organizations, companies,
pictorial, or other form that preserve archival government that could sponsor the
information, official and informal accounts, and development of the test
other data and items relevant to an assessee What?
▪ Case study – a report or illustrative account Educational Setting
concerning a person or an event that was o Achievement Test – evaluates accomplishment
compiled on the basis of case history data or the degree of learning that has taken place
▪ Groupthink – result of the varied forces that o Diagnostic Test – refers to a tool of assessment
drive decision-makers to reach a consensus used to help narrow down and identify areas of
o Behavioral Observation – monitoring of actions deficit to be targeted for intervention
of others or oneself by visual or electronic ▪ Diagnosis – description or conclusion
means while recording quantitative and/or reached on the basis of evidence and opinion
qualitative information regarding those actions o Informal Evaluation – nonsystematic
▪ Naturalistic Observation – observe humans assessment that leads to the formation of an
in natural setting opinion or attitude
o Role Play – defined as acting an improvised or Clinical Settings
partially improvised part in a stimulated o Used to help screen for or diagnose behavior
situation problems
▪ Role Play Test – assesses are directed to act o Tests could be intelligence tests, personality
as if they are in a particular situation tests, neuropsychological tests, or other
o Other tools include: computer, physiological specialized instruments
devices (biofeedback devices) o Usually individualized
Who, What, Why, How, and Where? Counseling Settings
Who? o May occur in environments as diverse as school,
o Test Developers – create tests or other methods prisons, and governmental or privately owned
of assessment institutions
o Test User – clinicians, counselors, o Goal: improve the client in terms of adjustment,
psychologists, HR personnel, consumer productivity, or some related variables
psychologists, experimental psychologists, and Geriatric Settings
social psychologists o Quality of Life – variables related to perceived
o Test taker – taking the test stress, loneliness, sources of satisfaction,
o Test takers vary in terms of: personal values, quality of living conditions, and
a. amount of test anxiety quality of friendships and other social support
Psychological Assessment
Psychological Testing and History
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Dementia – loss of cognitive functioning that assessee or by means of alternative methods
occurs as the result of damage to or loss of brain designed to measure same variables
cells o Accommodation – adaptation of a test,
o Pseudodementia – a severe depression that procedure, or situation of one test for another, to
mimics dementia make the assessment more suitable for an
Business and Military Settings assessee with an exceptional needs
o A wide range of achievement, aptitude, interest, Where?
motivational, and other tests may be employed in o Test Catalogues – contain only a brief
the decision to hire as well as in related description of the test and seldom contain the
decisions regarding promotions, transfers, job kind of detailed technical information that a
satisfaction, etc. prospective user might require
o Psychological Tests involves the engineering o Test Manuals – detailed information concerning
and design of products and environment the development of a particular test and
o Can work in marketing to help “diagnose” what technical information relating to it should be
can be improved with the brand found in the test manual
Governmental and Organizational Credentialing o Others: professional books, journals, online
o Governmental licensing, certification, or general databases
credentialing of professionals History
Academic Research Settings o Testing programs was first held in China as early
o Conducting any sort of research typically entails as 2200 B.C.E. for Civil Service
measurement of some kind, and any o 1733: Abraham De Moivre introduced the basic
academician who ever hopes to publish notion of sampling error
research should ideally have a sound knowledge o 1859: Charles Darwin argued that chance
of measurement principles and tools of variation in species would be selected or
assessment rejected by nature according to adaptivity and
Other Settings survival value
o Judiciary, program evaluation ▪ Humans had descended from the Ape as a
How? result of such genetic variations
o Test users should only use tests that are o 1869: Francis Galton explored and quantify
necessary and appropriate for the individual individual differences of people
being tested ▪ Classified people according to the “natural
o Test user must be prepared and suitably trained gifts” and to ascertain their “deviation from an
to administer the test properly average”
o Protocol – refers to the form or sheet or booklet ▪ Pioneered the use of a statistical concept
on which a testtaker’s responses are entered central to psychological experimentation and
o Rapport – working relationship between the testing: the coefficient of correlation
examiner and the examinee o Karl Pearson developed Product-Moment
o Test users who have responsibility for Correlation Technique
interpreting scores or other test results have an o First experimental psychology laboratory was
obligation to do so in accordance with founded by Wilhelm Wundt in Germany
established procedures and ethical guidelines ▪ Wundt focused on how similar people are and
o Alternate Assessment – for children who, as a viewed individual differences as frustrating
result of a disability, could not otherwise source of error
participate in state- and district-wide o James McKeen Cattell – coined the term Mental
assessments Test
▪ An evaluative or diagnostic procedure or o Charles Spearman – originated the concept of
process that varies from the usual, test reliability as well as building mathematical
customary, or standardized way a framework for the statistical technique of factor
measurement is derived, either by virtue of analysis
some special accommodation made to the
Psychological Assessment
Psychological Testing and History
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Victor Henri – collaborated with Alfred Binet on o Hermann Rorschach – developed Rorschach
papers suggesting how mental tests could be Inkblot test
used to measure higher mental processes o Henry Murray & Christiana Morgan – developed
o Emil Kraepelin – early experimentation with the Thematic Apperception Test
word association technique as a formal test o 1943: Minnesota Multiphasic Personality
▪ One of the founding founders of modern Inventory was published
psychiatry ▪ to use empirical methods to determine the
▪ Classification and diagnosis of mental meaning of a test response
disorders o Factor Analysis – method of finding the minimum
▪ Dementia Praecox number of dimensions (factors) to account for a
o Lightner Witmer – “Little know founder of large number of variables
Clinical Psychology” o J.R. Guilford – made the first serious attempt to
▪ Founded the first psychological clinic in US use factor analytic technique in the development
o 1895: Alfred Binet and Victor Henri published of a structured personality test
several articles in which they argued for the o Raymond Cattell – introduced 16PF
measurement of abilities o Beginning of 1980s, several major branches of
o 1905: Alfred Binet and Theodore Simon published applied psychology emerged such as
the first intelligence test designed to help neuropsych, health psych, forensic psych, and
identify Paris schoolchildren with ID child psych
▪ Considered standardization sample end
▪ Representative Sample – one that comprises
individuals similar to those for whom the test
is to be used
▪ 1908: Mental Age was determined
▪ L.M. Terman revised Binet’s test for US use
o 1939: David Wechsler introduced Adult
Intelligence Test
▪ Intelligence was the aggregate or the global
capacity of the individual to act purposefully,
to think rationally, and to deal effectively with
his environment (Weschler, 1939)
o Binet devised his intelligence test into group
intelligence test in response to the US Military’s
need for screening of recruits for WWI
o Lewis M. Terman, Robert M. Yerkes, and others
developed Army Tests for recruits
▪ Army Alpha – for literate folks
▪ Army Beta – for illiterate folks
o Robert Woodworth was assigned the task of
developing a measure of adjustment and
emotional stability that could be administered
quickly and efficiently to groups of recruits
▪ Disguised as Personal Data sheet
o Woodworth Psychoneurotic Inventory – first
self-report measure of personality to identify
soldiers at risk for shell shock
o Projective Test – one in which an individual is
assumed to project onto some ambiguous
stimulus his or her own unique feelings
Psychological Assessment
Statistics
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
Scales of Measurement 4. Ratio – has true zero point (if the score is zero, it
o Measurement – the act of assigning numbers or means none/null)
symbols to characteristics of things according to ▪ Easiest to manipulate
rules Describing Data
o Descriptive Statistics – methods used to provide o Distribution – defined as a set of test scores
concise description of a collection of quantitative arrayed for recording or study
information o Raw Scores – straightforward, unmodified
o Inferential Statistics – method used to make accounting of performance that is usually
inferences from observations of a small group of numerical
people known as sample to a larger group of o Frequency Distribution – all scores are listed
individuals known as population alongside the number of times each score
o Magnitude – the property of “moreness” occurred
o Equal Intervals – the difference between two o Independent Variable – being manipulated in the
points at any place on the scale has the same study
meaning as the difference between two other o Quasi-Independent Variable – nonmanipulated
points that differ by the same number of scale variable to designate groups
units ▪ Factor – for ANOVA
o Absolute 0 – when nothing of the property being o Post-Hoc Tests – used in ANOVA to determine
measured exists which mean differences are significantly
o Scale – a set of numbers who properties model different
empirical properties of the objects to which the o Tukey’s HSD test – allows the compute a single
numbers are assigned value that determines the minimum difference
▪ Continuous Scale – takes on any value within between treatment means that is necessary for
the range and the possible value within that significance
range is infinite Measures of Central Tendency
▪ Discrete Scale – can be counted; has distinct, o Measures of Central Tendency – statistics that
countable values indicates the average or midmost score
o Error – refers to the collective influence of all the between the extreme scores in a distribution
factors on a test score or measurement beyond ▪ Goal: Identify the most typical or
those specifically measured by the test or representative of entire group
measurement o Mean – the average of all the raw scores
▪ Measurement with continuous scale always ▪ Equal to the sum of the observations divided
involve with error by the number of observations
o Four Levels of the scales of Measurement: ▪ Interval and ratio data (when normal
1. Nominal – involve classification or distribution)
categorization based on one or more ▪ Point of least squares
distinguishing characteristics ▪ Balance point for the distribution
▪ Label and categorize observations but do not o Median – the middle score of the distribution
make any quantitative distinctions between ▪ Ordinal, Interval, Ratio
observations ▪ Useful in cases where relatively few scores
▪ mode fall at the high end of the distribution or
2. Ordinal – rank ordering on some characteristics relatively few scores fall at the low end of the
is also permissible distribution
▪ median ▪ In other words, for extreme scores, use
3. Interval – contains equal intervals, has no median
absolute zero point (even negative values have ▪ Identical for sample and population
interpretation to it) ▪ Also used when there has an unknown or
▪ Zero value does not mean it represents none undetermined score
Psychological Assessment
Statistics
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
▪ Used in “open-ended” categories (e.g., 5 or
more, more than 8, at least 10)
▪ For ordinal data
o Mode – most frequently occurring score in the
distribution
▪ Bimodal Distribution – if there are two scores
that occur with highest frequency
▪ Not commonly used
▪ Useful in analyses of qualitative or verbal
nature
▪ For nominal scales, discrete variables
▪ Value of the mode gives an indication of the
shape of the distribution as well as a measure
of central tendency
Measures of Variability o Symmetrical Distribution – right side of the
o Variability – an indication how scores in a graph is mirror image of the left side
distribution are scattered or dispersed ▪ Has only one mode and it is in the center of the
o Measures of Variability – statistics that describe distribution
the amount of variation in a distribution ▪ Mean = median = mode
o Range – equal to the difference between highest o Skewness – nature and extent to which
and the lowest score symmetry is absent
▪ Provides a quick but gross description of the o Positive Skewed – few scores fall the high end of
spread of scores the distribution
▪ When its value is based on extreme scores of ▪ The exam is difficult
the distribution, the resulting description of ▪ More items that was easier would have been
variation may be understated or overstated desirable in order to better discriminate at
o Quartile – dividing points between the four the lower end of the distribution of test scores
quarters in the distribution
▪ Specific point
▪ Quarter – refers to an interval
▪ Interquartile Range – measure of variability
equal to the difference between Q3 and Q1
▪ Semi-interquartile Range – equal to the
interquartile range divided by 2
o Standard Deviation – equal to the square root of
the average squared deviations about the mean
▪ Equal to the square root of the variance
▪ Variance – equal to the arithmetic mean of the ▪ Mean > Median > Mode
squares of the differences between the o Negative Skewed – when relatively few of the
scores in a distribution and their mean scores fall at the low end of the distribution
▪ Distance from the mean ▪ The exam is easy
Normal Curve ▪ More items of a higher level of difficulty would
o Also known as Gaussian Curve make it possible to better discriminate
o Bell-shaped, smooth, mathematically defined between scores at the upper end of the
curve that is highest at its center distribution
o Asymptotically = approaches but never touches
the axis
o Tail – 2 – 3 standard deviations above and below
the mean
Psychological Assessment
Statistics
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
▪ Raw score that fell in the mean has T of 50
▪ Raw score 5 standard deviations about the
mean would be equal to a T of 100
▪ No negative values
▪ Used when the population or variance is
unknown
o Stanine – a method of scaling test scores on a
nine-point standard scale with a mean of five (5)
and a standard deviation of two (2)
o Linear Transformation – one that retains a direct
▪ Mean < Median < Mode numerical relationship to the original raw score
o Skewed is associated with abnormal, perhaps o Nonlinear Transformation – required when the
because the skewed distribution deviates from data under consideration are not normally
the symmetrical or so-called normal distributed
distribution o Normalizing the distribution involves stretching
o Kurtosis – steepness if a distribution in its center the skewed curve into the shape of a normal
▪ Platykurtic – relatively flat curve and creating a corresponding scale of
▪ Leptokurtic – relatively peaked standard scores, a scale that is technically
▪ Mesokurtic – somewhere in the middle referred to as Normalized Standard Score Scale
o Generally preferrable to fine-tune the test
according to difficulty or other relevant
variables so that the resulting distribution will
approximate the normal curve
o STEN – standard to ten; divides a scale into 10
units
▪ High Kurtosis = high peak and fatter tails
▪ Lower Kurtosis = rounded peak and thinner
tails
Standard Scores
o Standard Score – raw score that has been
converted from one scale to another scale
Hypothesis Testing
o Z-Scores – results from the conversion of a raw
o Statistical method that uses a sample data to
score into a number indicating how many SD
evaluate a hypothesis about a population
units the raw score is below or above the mean
o Alternative Hypothesis – states there is a
of the distribution
change, difference, or relationships
▪ Identify and describe the exact location of
o Null Hypothesis – no change, no difference, or no
each score in a distribution
relationship
▪ Standardize an entire distribution
o Alpha Level or Level of Significance – used to
▪ Zero plus or minus one scale
define concept of “very unlikely” in a hypothesis
▪ Have negative values
test
▪ Requires that we know the value of the
o Critical Region – composed of extreme values
variance to compute the standard error
that are very unlikely to be obtained if the null
o T-Scores – a scale with a mean set at 50 and a
hypothesis is true
standard deviation set at 10
o If sample data fall in the critical region, the null
▪ Fifty plus or minus 10 scale
hypothesis is rejected
▪ 5 standard deviation below the mean would
be equal to a t-score of 0
Psychological Assessment
Statistics
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
o The alpha level for a hypothesis test is the ▪ Coefficient of Determination – an indication of
probability that the test will lead to a Type I error how much variance is shared by the X- and Y-
o Directional Hypothesis Test or One-Tailed Test – variables
statistical hypotheses specify either an increase o Spearman Rho/Rank-Order Correlation
or a decrease in the population mean Coefficient/Rank-Difference Correlation
o T-Test – used to test hypotheses about an Coefficient – frequently used if the sample size is
unknown population mean and variance small and when both sets of measurement are in
▪ Can be used in “before and after” type of ordinal
research ▪ Developed by Charles Spearman
▪ Sample must consist of independent Pearson R
observationsꟷthat is, if there is not Interval/ratio + interval/ratio
consistent, predictable relationship between Two continuous variables
the first observation and the second Spearman Rho
▪ The population that is sampled must be Ordinal + Ordinal
normal Point-Biserial Coefficient
▪ If not normal distribution, use a large sample Dichotomous
Two nominal + continuous variable (interval/ratio)
Kendall’s Coefficient
3 or more rank/ordinals (ratings)
Ordinal + ordinal + ordinal
Correlation and Inference
Phi or Fourfold Coefficient
o Correlation Coefficient – number that provides
Nominal + nominal
us with an index of the strength of the
All Dichotomous variables
relationship between two things
Rank Biserial Correlation
o Correlation – an expression of the degree and
direction of correspondence between two things Nominal + Ordinal (Rating)
▪ + & - = direction Tetrachoric Correlation R
▪ Number anywhere to -1 to 1 = magnitude Continuous + Continuous but both are measured
▪ Positive – same direction, either both going as Nominal (e.g., Passed or Not Passed, rather
up or both going down than grades itself)
▪ Negative – Inverse Direction, either DV is up o Outlier – extremely atypical point located at a
and IV goes down or IV goes up and DV goes relatively long distance from the rest of the
down coordinate points in a scatterplot
▪ 0 = no correlation o Regression Analysis – used for prediction
▪ Predict the values of a dependent or
response variable based on values of at least
one independent or explanatory variable
▪ Residual – the difference between an
observed value of the response variable and
the value of the response variable predicted
from the regression line
▪ The Principle of Least Squares
▪ Standard Error of Estimate – standard
deviation of the residuals in regression
o Pearson r/Pearson Correlation analysis
Coefficient/Pearson Product-Moment ▪ Slope – determines how much the Y variable
Coefficient of Correlation – used when two changes when X is increased by 1 point
variables being correlated are continuous and o T-Test (Independent) – comparison or
linear determining differences
▪ Devised by Karl Pearson
Psychological Assessment
Statistics
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
▪ 2 different groups/independent samples +
interval/ratio scales (continuous varriables)
▪ Equal Variance – 2 groups are equal
▪ Unequal Variance – groups are unequal
o T-test (Dependent)/Paired Test – two groups
nominal (either matched or repeated measures)
+ continuous scales
o One-Way ANOVA – 3 or more IV, 1 DV comparison
of differences
o Two-Way ANOVA – 2 IV, 1 DV
o Critical Value – reject the null and accept the
alternative if [ obtained value > critical value ]
o P-Value (Probability Value) – reject null and
accept alternative if [ p-value < alpha level ]
Norms
o Norms – refer to the performances by defined
groups on a particular tests
o Age-Related Norms – certain tests have
different normative groups for particular age
groups
o Tracking – tendency to stay at about the same
level relative to one’s peers
o Norm-Referenced Tests – compares each
person with the norm
o Criterion-Referenced Tests – describes specific
types of skills, tasks, or knowledge that the test
taker can demonstrate
end
Psychological Assessment
Assumptions about Psychological Testing and Assessment
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Assumption 1: Psychological Traits and States Exist provide insight to it, to gauge the strength of that
o Trait – any distinguishable, relatively enduring trait
way in which one individual varies from another o Measuring traits and states means of a test
▪ Permit people predict the present from the entails developing not only appropriate tests
past items but also appropriate ways to score the test
▪ Characteristic patterns of thinking, feeling, and interpret the results
and behaving that generalize across similar o Cumulative Scoring – assumption that the more
situations, differ systematically between the testtaker responds in a particular direction
individuals, and remain rather stable across keyed by the test manual as correct or
time consistent with a particular trait, the higher that
▪ Psychological Trait – intelligence, specific testtaker is presumed to be on the targeted
intellectual abilities, cognitive style, ability or trait
adjustment, interests, attitudes, sexual Assumption 3: Test-Related Behavior Predicts Non-
orientation and preferences, Test-Related Behavior
psychopathology, etc. o The tasks in some tests mimics the actual
o States – distinguish one person from another behaviors that the test user is attempting to
but are relatively less enduring understand
▪ Characteristic pattern of thinking, feeling, o Such tests only yield a sample of the behavior
and behaving in a concrete situation at a that can be expected to be emitted under nontest
specific moment in time conditions
▪ Identify those behaviors that can be Assumption 4: Test and Other Measurement
controlled by manipulating the situation Techniques have strengths and weaknesses
o Psychological Traits exists as construct o Competent test users understand and
▪ Construct – an informed, scientific concept appreciate the limitations of the test they use as
developed or constructed to explain a well as how those limitations might be
behavior, inferred from overt behavior compensated for by data from other sources
▪ Overt Behavior – an observable action or the Assumption 5: Various Sources of Error are part of
product of an observable action the Assessment Process
o Trait is not expected to be manifested in behavior o Error – refers to something that is more than
100% of the time expected; it is component of the measurement
o Whether a trait manifests itself in observable process
behavior, and to what degree it manifests, is ▪ Refers to a long-standing assumption that
presumed to depend not only on the strength of factors other than what a test attempts to
the trait in the individual but also on the nature of measure will influence performance on the
the action (situation-dependent) test
o Context within which behavior occurs also plays ▪ Error Variance – the component of a test
a role in helping us select appropriate trait terms score attributable to sources other than the
for observed behaviors trait or ability measured
o Definition of trait and state also refer to a way in o Potential Sources of error variance:
which one individual varies from another 1. Assessors
o Assessors may make comparisons among 2. Measuring Instruments
people who, because of their membership in 3. Random errors such as luck
some group or for any number of other reasons, o Classical Test Theory – each testtaker has true
are decidedly not average score on a test that would be obtained but for the
Assumption 2: Psychological Traits and States can be action of measurement error
Quantified and Measured Assumption 6: Testing and Assessment can be
o Once the trait, state or other construct has been conducted in a Fair and Unbiased Manner
defined to be measured, a test developer o Despite best efforts of many professionals,
consider the types of item content that would fairness-related questions and problems do
occasionally rise
Psychological Assessment
Assumptions about Psychological Testing and Assessment
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o In al questions about tests with regards to testtakers in a given period of time rather than
fairness, it is important to keep in mind that tests norms obtained by formal sampling methods
are tools ꟷthey can be used properly or o Standardization – the process of administering a
improperly test to a representative sample of testtakers for
Assumption 7: Testing and Assessment Benefit the purpose of establishing norms
Society o Sample – a portion of the universe of people that
o Considering the many critical decisions that are represents the whole population
based on testing and assessment procedures, o Sampling – process of selecting sample
we can readily appreciate the need for tests Probability Sampling
What is a “Good Test”? Random sampling, randomization is used to
o Includes clear instructions for administration, select samples
scoring, and interpretation; offered economy in Simple Random Sampling
the time and money it took to administer, score, Every element in the population has an equal
and interpret; and, measures what it purports to chance of being selected as part of the sample
measure Easy, cheap, remove all risk of bias
o Reliability – consistency of the measuring tool Systematic Sampling
▪ The precision with which test measurers and Every nth item or person after is picked
the extent to which error is present in Researcher can choose the interval at which
measurements items are picked
o Validity – measure what it is supposed to Stratified Sampling
measure Random selection within predefined groups
o A good test is one that trained examiners can More risk of bias due to stratifying
administer, score, and interpret with a minimum
Cluster Sampling
difficulty
Groups rather than individual units of the target
▪ Yields actionable results that will ultimately
population are selected randomly
benefit individual testtakers or society at
Non-Probability Sampling
large
Researchers pick items or individual based on
Norms
their research goals or knowledge
o Norm-Referenced Testing and Assessment –
method of evaluation and a way of deriving Convenience Sampling
meaning from test scores by evaluating an Selected based on their availability
individual testtaker’s score and comparing it to Quota Sampling
scores of a group testtakers Achieve a spread across the target population by
▪ Yield information on a testtaker’s standing or specifying who should be recruited for a survey
ranking relative to some comparison group of according to certain groups or criteria
testtakers Purposive Sampling
o Norms – usual, average, normal, standard, Chosen consciously based on their knowledge
expected, or typical and understanding of the research question at
▪ Test performance data of a particular group hand or their goals
of testtakers that are designed for use as a Snowball or Referral Sampling
reference when evaluating or interpreting People recruited to be part of a sample are asked
individual test scores to invite those they know to take part, who are then
o Normative Sample – group of people whose asked to invite their friends and family and so on
performance on a particular test is analyzed for Helpful when the researcher doesn’t know very
reference in evaluating the performance of much about the target population and has no easy
individual testtakers way to contact or access them
o Norming – process of deriving norms o After obtaining sample for standardization, the
o User Norms or Program Norms – consists or test developer will administer the test according
descriptive statistics based on a group of to the standard set of instructions that will be
Psychological Assessment
Assumptions about Psychological Testing and Assessment
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
used with the test and also describe the
recommended setting for giving the test
o Percentile – an expression of the percentage of
people whose score on a test or measure falls
before a particular raw score
▪ Percentage Correct – refers to the
distribution of raw scores, specifically, to the
number of items that were answered
correctly multiplied by 100 and divided by the
total number of items
o Age Norms – average performance of different
samples of testtakers who were at various ages
at the time the test was administered
o Grade Norms – developed by administering the
test to representative samples of children over a
range of consecutive grade levels
▪ Developmental Norms – norms developed on
the basis of any trait, ability, skill, or other
characteristics that is presumed to develop,
deteriorate, or otherwise affected by
chronological age, school grade, or stage of
life
o National Norms – derived from a normative
sample that was nationally representative of the
population at the time norming study was
conducted
o Local Norms – provide normative information
with respect to the local population’s
performance on some test
o Fixed Reference Group Scoring System – the
distribution of scores obtained on the test from
one group of testtakers (fixed reference group)
is used as the basis for the calculation of test
scores for future administrations of the test
o Criterion-Referenced Tests and Assessment –
method of evaluation and a way of deriving
meaning from test scores by evaluating an
individual’s score with reference to a set
standard (criterion)
o Domain- or Content-Referenced Testing and
Assessment – how scores relate to a particular
content area or domain
end
Psychological Assessment
Reliability, Validity, Utility
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
Reliability test scoreꟷand thus the reliability can be
o Dependability or consistency affected
o Consistency of the instrument o Measurement Error – all of the factors
o A test may be reliable in one context and associated with the process of measuring some
unreliable in another variable, other than the variable being measured
o Reliability Coefficient – index of reliability, a o Random Error – source of error in measuring a
proportion that indicates the ratio between the targeted variable caused by unpredictable
true score variance on a test and the total fluctuations and inconsistencies of other
variance variables in the measurement process
o Classical Test Theory – a score on an ability tests ▪ “Noise”
is presumed to reflect not only the testtaker’s ▪ E.g., physical events that happened while
true score on the ability being measured but also test is happening
error o Systematic Error – source of error in a
▪ Errors of measurement are random measuring a variable that is typically constant or
o Error – refers to the component of the observed proportionate to what is presumed to be the true
test score that does not have to do with the value of the variable being measured
testtaker’s ability o Sources of Error Variance:
a. Item Sampling/Content Sampling – refer to
variation among items within a test as well as to
variation among items between tests
▪ The extent to which testtaker’s score is
affected by the content sampled on a test and
by the way the content is sampled is a source
of error variance
b. Test Administration
▪ Testtaker’s motivation or attention,
environment, etc.
▪ Testtaker variables and Examiner-related
Variables
c. Test Scoring and Interpretation
▪ Type I – “false-positive”; an investigator ▪ May employ objective-type items amenable
rejects a null hypothesis that is true to computer scoring of well-documented
▪ Type II – “false-negative”; fails to reject null reliability
hypothesis that is false in the population ▪ If subjectivity is involved in scoring,
▪ Can reduce the likelihood of type 1 and 2 Reliability Estimates
errors by increasing the sample size Test-Retest Reliability
o Variance – useful in describing sources of test o Time Sampling
score variability o An estimate of reliability obtained by correlating
▪ True Variance – variance from true pairs of scores from the same people on two
differences different administrations of the test
▪ Error Variance – variance from irrelevant, o Appropriate when evaluating the reliability of a
random sources test that purports to measure something that is
o Reliability refers to the proportion of total relatively stable such as personality trait
variance attributed to true variance o The longer the time that passes, the greater the
o The greater the proportion of the total variance likelihood that the reliability coefficient will be
attributed to true variance, the more reliable the lower
test o Coefficient of Stability - when the interval
o Error variance may increase or decrease a test between testing is greater than 6 months
score by varying amounts, consistency of the Parallel Forms and Alternate Forms Reliability
Psychological Assessment
Reliability, Validity, Utility
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
o Item Sampling o Randomly assign items to one or the other half of
o Coefficient of Equivalence – the degree of the test or assign odd-numbered items to one
relationship between various forms of test can half and even-numbered to the other half (odd-
be evaluated by means of an alternate forms or even reliability)
parallel forms coefficient of reliability o Divide the test by content so that each half
o Parallel Forms – each form of the test, the contains items equivalent with respect to
means and the variances are equal content and difficulty
▪ Same items, different o Spearman-Brown Formula – allows a test
positionings/numberings developer or user to estimate internal
▪ Parallel Forms Reliability – estimate of the consistency reliability from a correlation of two
extent to which item sampling and other halves of a test
errors have affect test scores on version of o Reliability of the test is affected by the length.
the same test when, for each form of the test, Usually, reliability increases as length increases
the means and variances of observed test o Spearman-Brown may be used to estimate the
scores are equal effect of the shortening on the test’s reliability
o Alternate Forms – simply different version of a o SBF also be used to determine the number of
test that has been constructed so as to be items needed to attain a desired level of
parallel reliability
▪ Alternate Forms Reliability – estimate of the o If the reliability of the original test is relatively
extent to which these different forms of the low, then it may be impractical to increase the no.
same test have been affected by sampling of items, so they should develop a suitable
error, or other error alternative
o Two administrations with the same group are o Or increase reliability by creating new items,
required clarifying the test instructions, or simplifying the
o Test scores may be affected by factors such as scoring rules
motivation, fatigue, or intervening events such Inter-item Consistency
as practice, learning, or therapy o Refers to the degree of correlation among all the
o Some testtaker might do better on a specific items on a scale
form of a test but not a function of their true o Calculated from a single administration of a
ability but simply bec of the particular items that single form of a test
were selected for inclusion in the test o Useful in assessing Homogeneity
o Minimizes the effect of memory for the content of ▪ Homogeneity – if a test contains items that
a previously administered form of the test measure a single trait (unifactorial)
o Certain traits are presumed to relatively stable ▪ Heterogeneity – degree to which a test
in people measures different factors (more than one
o The means and the variances of the observed trait); source of error variance
scores are equal for two forms o More homogenous = higher inter-item
Internal Consistency consistency
Split-Half Reliability o KR-20 – used for inter-item consistency of
o Obtained by correlating two pairs of scores dichotomous items
obtained from equivalent halves of a single test o KR-21 – if all the items have the same degree of
administered once difficulty (speed tests)
o Useful when it is impractical or undesirable to o Coefficient Alpha – appropriate for use on tests
assess reliability with two tests or to administer containing non-dichotomous items
a test twice ▪ Help answer questions about how similar
o Simply diving the test in the middle is not sets of data are
recommended because it is likely that this ▪ Check consistency across terms of an
procedure would spuriously raise or lower the instrument with responses with varying
reliability coefficient credit
Psychological Assessment
Reliability, Validity, Utility
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
o Average Proportional Distance – measure used ▪ As individual differences decrease, a
to evaluate internal consistency of a test that traditional measure of reliability would also
focuses on the degree of differences that exists decrease, regardless of the stability of
between item scores individual performance
▪ Not connected to the number of items on a o Classical Test Theory – everyone has a “true
measure score” on test
Inter-scorer Reliability ▪ True Score – genuinely reflects an
o The degree of agreement or consistency individual’s ability level as measured by a
between two or more scorers with regard to a particular test
particular measure o Domain Sampling Theory – estimate the extent
o Used for coding nonverbal behavior to which specific sources of variation under
o Coefficient of Inter-scorer Reliability defined conditions are contributing to the test
o Observer Differences scores
o Kappa Statistics is used ▪ Considers problem created by using a limited
▪ Fleiss Kappa – determine the level of number of items to represent a larger and
agreement between two or more raters when more complicated construct
the method of assessment is measured on a ▪ Test reliability is conceived of as an objective
categorical scale; best way; more than 2 measure of how precisely the test score
raters assesses the domain from which the test
▪ Cohen’s Kappa – each classify N items into C draws a sample
mutually exclusive categories; rates the ▪ Generalizability Theory – based on the idea
same thing, corrected for how often that the that a person’s test scores vary from testing
raters may agree by chance; only 2 raters to testing because of the variables in the
Using and Interpreting Coefficient of Reliability testing situations
o Tests designed to measure one factor ▪ Universe – test situation
(Homogenous) are expected to have high degree ▪ Facets – number of items in the test, amount
of internal consistency and vice versa of review, and the purpose of test
o Dynamic – trait, state, or ability presumed to be administration
ever-changing as a function of situational and ▪ According to Generalizability Theory, given
cognitive experience the exact same conditions of all the facets in
o Static – barely changing or relatively the universe, the exact same test score
unchanging should be obtained (Universe score)
o Restriction of range or Restriction of variance – ▪ Decision Study – developers examine the
if the variance of either variable in a usefulness of test scores in helping the test
correlational analysis is restricted by the user make decisions
sampling procedure used, then the resulting o Item Response Theory – the probability that a
correlation coefficient tends to be lower person with X ability will be able to perform at a
o Power Tests – when time limit is long enough to level of Y in a test
allow test takers to attempt all times ▪ Latent-Trait Theory
o Speed Tests – generally contains items of ▪ The computer is used to focus on the range of
uniform level of difficulty with time limit item difficulty that helps assess an
▪ Reliability should be based on performance individual’s ability level
from two independent testing periods using ▪ Difficulty – attribute of not being easily
test-retest and alternate-forms or split- accomplished, solved, or comprehended
half-reliability ▪ Discrimination – degree to which an item
o Criterion-Referenced Tests – designed to differentiates among people with higher or
provide an indication of where a testtaker stands lower levels of the trait, ability or etc.
with respect to some variable or criterion ▪ Dichotomous – can be answered with only
one of two alternative responses
Psychological Assessment
Reliability, Validity, Utility
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
▪ Polytomous – 3 or more alternative o Content Validity – describes a judgement of how
responses adequately a test samples behavior
Reliability and Individual Scores representative of the universe of behavior that
o Standard Error of Measurement – provide a the test was designed to sample
measure of the precision of an observed test o When the proportion of the material covered by
score the test approximates the proportion of material
▪ Standard deviation of errors as the basic covered in the course
measure of error o Test Blueprint – a plan regarding the types of
▪ Provides an estimate of the amount of error information to be covered by the items, the no. of
inherent in an observed score or items tapping each area of coverage, the
measurement organization of the items, and so forth
▪ Higher reliability, lower SEM Criterion-Related Validity
▪ Used to estimate or infer the extent to which o Criterion-Related Validity – a judgement of how
an observed score deviates from a true score adequately a test score can be used to infer an
▪ Standard Error of a Score individual’s most probable standing on some
▪ Confidence Interval – a range or band of test measure of interestꟷthe measure of interest
scores that is likely to contain true scores being criterion
o Standard Error of the Difference – can aid a test o Criterion – standard on which a judgement or
user in determining how large a difference decision may be made
should be before it is considered statistically ▪ Characteristics: relevant, valid,
significant uncontaminated
o Standard Error of Estimate – refers to the ▪ Criterion Contamination – occurs when the
standard error of the difference between the criterion measure includes aspects of
predicted and observed values performance that are not part of the job or
when the measure is affected by “construct-
irrelevant ” (Messick, 1989) factors that are
not part of the criterion construct
Concurrent Validity
o If the test scores obtained at about the same time
as the criterion measures are obtained
o Extent to which test scores may be used to
estimate an individual’s present standing on a
Validity criterion
o Validity – a judgment or estimate of how well a o Economically efficient
test measures what it supposed to measure Predictive Validity
▪ Evidence about the appropriateness of o Measures of the relationship between test
inferences drawn from test scores scores and a criterion measure obtained at a
▪ Inferences – logical result or deduction future time
▪ May diminish as the culture or times change o Researchers must take into consideration the
o Validation – the process of gathering and base rate of the occurrence of the variable, both
evaluating evidence about validity as that variable exists in the general population
o Validation Studies – yield insights regarding a and as it exists in the sample being studies
particular population of testtakers as compared o Base Rate – the extent to which a particular trait,
to the norming sample described in a test behavior, characteristic, or attribute exist in the
manual population
o Face Validity – a test appears to measure to the o Hit Rate – defined as the proportion of people a
person being tested than to what the test test accurately identifies possessing a
actually measures particular trait, behavior, etc.
Content Validity
Psychological Assessment
Reliability, Validity, Utility
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
o Miss Rate – fails to identify having that particular would have different test scores than those
characteristic who really possesses that construct
o False Positive – miss; the test predicted that they o Convergent Evidence – if scores on the test
do possess a particular trait but actually not undergoing construct validation tend to highly
o False Negative – miss; the test predicted they do correlated with another established, validated
not possess a particular trait but actually do test that measures the same construct
o Validity Coefficient – correlation coefficient that o Discriminant Evidence – a validity coefficient
provides a measure of the relationship between showing little relationship between test scores
test scores and scores on the criterion measure and/or other variables with which scores on the
▪ Usually, Pearson R is used, however other test being construct-validated should not be
correlation coefficients could be used correlated
depends on the type of data o Factor Analysis – designed to identify factors or
▪ Affected by restriction or inflation of range specific variables that are typically attributes,
▪ Validity Coefficient need to be large enough to characteristics, or dimensions on which people
enable the test user to make accurate may differ
decisions within the unique context in which a ▪ Employed as data reduction method
test is being used ▪ Identify the factor or factors in common
o Incremental Validity – the degree to which an between test scores on subscales within a
additional predictor explains something about particular test
the criterion measure that is not explained by ▪ Explanatory FA – estimating or extracting
predictors already in use factors; deciding how many factors must be
Construct Validity retained
o Construct Validity – judgement about the ▪ Confirmatory FA – researchers test the
appropriateness of inferences drawn from test degree to which a hypothetical model fits the
scores regarding individual standing on variable actual data
called construct ▪ Factor Loading – conveys info about the
o Construct – an informed, scientific idea extent to which the factor determine the test
developed or hypothesized to describe or score or scores
explain behavior Validity, Bias, and Fairness
▪ Unobservable, presupposed traits that may o Bias – factor inherent in a test that
invoke to describe test behavior or criterion systematically prevents accurate, impartial
performance measurement
o One way a test developer can improve the ▪ Prejudice, preferential treatment
homogeneity of a test containing dichotomous ▪ Prevention during test dev through a
items is by eliminating items that do not show procedure called Estimated True Score
significant correlation coefficients with total test Transformation
scores o Rating – numerical or verbal judgement that
o If it is an academic test and high scorers on the places a person or an attribute along a
entire test for some reason tended to get that continuum identified by a scale of numerical or
particular item wrong while low scorers got it word descriptors known as Rating Scale
right, then the item is obviously not a good one ▪ Rating Error – intentional or unintentional
o Some constructs lend themselves more readily misuse of the scale
than others to predictions of change over time ▪ Leniency Error – rater is lenient in scoring
o Method of Contrasted Groups – demonstrate (Generosity Error)
that scores on the test vary in a predictable way ▪ Severity Error – rater is strict in scoring
as a function of membership in a group ▪ Central Tendency Error – rater’s rating would
▪ If a test is a valid measure of a particular tend to cluster in the middle of the rating
construct, then the scores from the group of scale
people who does not have that construct
Psychological Assessment
Reliability, Validity, Utility
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
▪ One way to overcome rating errors is to use scores on a criterion measure – passing,
rankings acceptable, failing
▪ Halo Effect – tendency to give high score due o Might indicate future behaviors, then if
to failure to discriminate among conceptually successful, the test is working as it should
distinct and potentially independent aspects o Taylor-Russel Tables – provide an estimate of
of a ratee’s behavior the extent to which inclusion of a particular test
o Fairness – the extent to which a test is used in an in the selection system will improve selection
impartial, just, and equitable way o Selection Ratio – numerical value that reflects
o Attempting to define the validity of the test will be the relationship between the number of people
futile if the test is NOT reliable to be hired and the number of people available to
be hired
o Base Rate – percentage of people hired under
the existing system for a particular position
o One limitation of Taylor-Russel Tables is that the
relationship between the predictor (test) and
criterion must be linear
o Naylor-Shine Tables – entails obtaining the
difference between the means of the selected
and unselected groups to derive an index of what
the test is adding to already established
Utility procedures
o Utility – usefulness or practical value of testing Brogden-Cronbach-Gleser Formula
to improve efficiency o Used to calculate the dollar amount of a utility
o Can tell us something about the practical value gain resulting from the use of a particular
of the information derived from scores on the selection instrument
test o Utility Gain – estimate of the benefit of using a
o Helps us make better decisions particular test
o Higher criterion-related validity = higher utility o Productivity Gains – an estimated increase in
o One of the most basic elements in utility analysis work output
is financial cost of the selection device Some Practical Considerations
o Cost – disadvantages, losses, or expenses both o High performing applicants may have been
economic and noneconomic terms offered in other companies as well
o Benefit – profits, gains or advantages o The more complex the job, the more people
o The cost of test administration can be well worth differ on how well or poorly they do that job
it if the results is certain noneconomic benefits o Cut Score – reference point derived as a result
Utility Analysis of a judgement and used to divide a set of data
o Utility Analysis – family of techniques that entail into two or more classifications
a cost-benefit analysis designed to yield ▪ Relative Cut Score – reference point based
information relevant to a decision about the on norm-related considerations (norm-
usefulness and/or practical value of a tool of referenced); e.g, NMAT
assessment ▪ Fixed Cut Scores – set with reference to a
How is Utility Analysis Conducted? judgement concerning minimum level of
Expectancy Data proficiency required; e.g., Board Exams
o Expectancy table – provide an indication that a ▪ Multiple Cut Scores – refers to the use of
testtaker will score within some interval of two or more cut scores with reference to
Psychological Assessment
Reliability, Validity, Utility
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
one predictor for the purpose of
categorization
▪ Multiple Hurdle – multi-stage selection
process, a cut score is in place for each
predictor
▪ Compensatory Model of Selection –
assumption that high scores on one
attribute can compensate for lower scores
Methods for Setting Cut Scores
o Angoff Method – setting fixed cut scores
▪ low interrater reliability
o Known Groups Method – collection of data on the
predictor of interest from group known to
possess and not possess a trait of interest
▪ The determination of where to set cutoff
score is inherently affected by the
composition of contrasting groups
o IRT-Based Methods – cut scores are typically set
based on testtaker’s performance across all the
items on the test
▪ Item-Mapping Method – arrangement of
items in histogram, with each column
containing items with deemed to be
equivalent value
▪ Bookmark Method – expert places
“bookmark” between the two pages that are
deemed to separate testtakers who have
acquired the minimal knowledge, skills,
and/or abilities from those who are not
o Method of Predictive Yield – took into account the
number of positions to be filled, projections
regarding the likelihood of offer acceptance, and
the distribution of applicant scores
o Discriminant Analysis – shed light on the
relationship between identified variables and
two naturally occurring groups
end
Psychological Assessment
Test Development
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Test Conceptualization ▪ Multidimensional – more than one
o Test Development – an umbrella term for all that dimension
goes into the process of creating a test ▪ Comparative and Categorical
o Test Conceptualization – brain storming of ideas ▪ Rating Scale – grouping of words,
about what kind of test a developer wants to statements, or symbols on which judgments
publish of the strength of a particular trait are
o Questions to ponder on when conceptualizing indicated by the testtaker
for new tests: ▪ Summative Scale – final score is obtained by
1. What is the test designed to measure? summing the ratings across all the items
2. What is the objective? ▪ Likert Scale – scale attitudes, usually
3. Is there a need for this kind of test? reliable;
4. Who will use the test? ▪ Thurstone Scale - involves the collection of a
5. Who will take the test? variety of different statements about a
6. What content will the test cover? phenomenon which are ranked by an expert
7. How will the test be administered? panel in order to develop the questionnaire
8. What is the ideal format of the test? ▪ Method of Paired Comparisons – produces
9. Should more than one form of test be developed? ordinal data by presenting with pairs of two
10. What special training will be required of test stimuli which they are asked to compare
users for administering or interpreting the test? ▪ Comparative Scaling – entails judgments of
11. What types of responses will be required of a stimulus in comparison with every other
testtaker’s? stimulus on the scale
12. Who benefits from an administration of this test? ▪ Categorical Scaling – stimuli are placed into
13. Is there potential harm? one of two or more alternative categories
14. How will meaning be attributed to scores on this that differ quantitatively with respect to
test? some continuum
o Pilot Work/Pilot Study/Pilot Research – ▪ Guttman Scale – yields ordinal-level
preliminary research surrounding the creation measures
of a prototype of the test o Item Pool – reservoir or well from which the
▪ Attempts to determine how best to measure a items will or will not be drawn for the final
targeted construct version of the test
▪ Entail lit reviews and experimentation, ▪ A comprehensive sampling provides a basis
creation, revision, and deletion of preliminary for content validity of the final version of the
items test
Test Construction ▪ The test developer may write a large number
o Test Construction – stage in the process that of items from personal experience or
entails writing test items, revisions, formatting, academic acquaintance with the subject
setting scoring rules matter or experts
o Scaling – process of setting rules for assigning o Item Format – form, plan, structure,
numbers in measurement arrangement, and layout of individual test items
▪ Process by which a measuring device is ▪ Selected-Response Format – require
assigned and calibrated and by which testtakers to select response from a set of
numbersꟷscale valuesꟷare assigned to alternative responses
different amounts of the trait, attribute, or Multiple-Choice Format
characteristic being measured Has three elements: stem (question), a correct
▪ Age-Based – age is of critical interest option, and several incorrect alternatives
▪ Grade-Based – grade is of critical interest (distractors or foils)
▪ Stanine – if all raw score of the test are to be Should’ve one correct answer, has grammatically
transformed into scores that range from 1-9 parallel alternatives, similar length, alternatives
▪ Unidimensional – only one dimension is that fit grammatically with the stem, avoid
presumed to underlie the ratings
Psychological Assessment
Test Development
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
ridiculous distractors, not excessively long, “all of ▪ Floor Effects – occurs when there is some
the above”, “none of the above” lower limit on a survey or questionnaire and a
Probability of getting the correct answer is 25% large percentage of respondents score near
Matching Item this lower limit (testtakers have low scores)
Test taker is presented with two columns: ▪ Ceiling Effects – occurs when there is some
Premises and Responses upper limit on a survey or questionnaire and a
Should be fairly short and to the point and only one large percentage of respondents score near
premise would match to one response this upper limit (testtakers have high scores)
Binary Choice o Item Branching – ability of the computer to tailor
True-False Item the content and order of presentation of items on
Usually takes the form of a sentence that requires the basis of responses to previous items
the testtaker to indicate whether the statement is o Cumulative Scoring – the higher score one
or is not a fact achieved on the test, the higher the testtaker is
Contains single idea and not subject to debate on the ability that the test purports to measure
Probability of obtaining the correct answer is 50% o Class Scoring/Category Scoring – testtaker
▪ Constructed-Response Format – requires responses earn credit toward placement in a
testtakers to supply or to create the correct particular class or category with other
answer, not merely selecting it testtakers who pattern of responses is
presumably similar in some way
Completion Item
o Ipsative Scoring – comparing a testtaker’s score
Requires the examinee to provide a word or phrase
on one scale within a test to another scale within
that completes a sentence
that same test
Should be worded properly so that the correct
o Semantic Differential Rating Technique -
answer is specific
measures an individual's unique, perceived
Short-answer item
meaning of an object, a word, or an individual;
Should be written clearly enough that the testtaker usually essay type, open-ended format
can respond succinctly, with short answer Test Tryout
Essay Item o The test should be tried out on people who are
Respond by writing a composition similar in critical respects to the people for
Allows creative integration and expression of the whom the test was designed
material o An informal rule of thumb should be no fewer
Tends to focus on a more limited area than can be that 5 and preferably as many as 10 for each item
covered in the same amount of time when using a (the more, the better)
series of selected-response items or completion o Risk of using few subjects = phantom factors
items emerge
Subject to scoring and inter-scorer differences o Should be executed under conditions as
o Item Banks – relatively large and easily identical as possible
accessible collection of test questions o Pseudobulbar Affect – neurological disorder
o Computerized Adaptive Testing – refers to an characterized by frequent involuntary outburst
interactive, computer administered test-taking of laughing or crying that may or may not be
process wherein items presented to the appropriate to the situation
testtaker are based in part on the testtaker’s o A good test item is one that answered correctly
performance on previous items by high scorers as a whole
▪ The test administered may be different for o Empirical Criterion Keying - administering a
each testtaker, depending on the test large pool of test items to a sample of individuals
performance on the items presented who are known to differ on the construct being
▪ Reduce the number of test items that need to measured
be administered by 50% while simultaneously ▪ To retain items that discriminate two groups
reducing measurement error by 50% Item Analysis
▪ Reduces floor and ceiling effects o Statistical procedure used to analyze items
Psychological Assessment
Test Development
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Item Difficulty – defined by the number of people
who get a particular item correct
o Item-Difficulty Index – calculating the
proportion of the total number of testtakers who
answered the item correctly
▪ The larger, the easier the item
▪ For achievement testing
▪ Item-Endorsement Index for personality
testing
▪ The optimal average item difficulty is approx. o Item-Characteristic Curve – graphic
50% with items on the testing ranging in representation of item difficulty and
difficulty from about 30% to 80% discrimination
o Guessing – one that eluded any universally
accepted solutions
o Item analyses taken under speed conditions
yield misleading or uninterpretable results
▪ Restrict item analysis on a speed test only to
the items completed by the testtaker
▪ Test developer ideally should administer the
test to be item-analyzed with generous time
limits to complete the test
o Item-Reliability Index – provides an indication of o Qualitative Methods – techniques of data
the internal consistency of a test generation and analysis that rely primarily on
▪ The higher this index, the greater the test’s verbal rather than statistical procedures
internal consistency o Qualitative Item Analysis – various
o Item-Validity Index – designed to provide an nonstatistical procedures designed to explore
indication of the degree to which a test is how individual test items work
measure what it purports to measure o An innovative cognitive assessment entails
▪ The higher this index, the greater the test’s having respondents verbalize thoughts as they
criterion-related validity occur (Think Aloud Test Administration)
o Item-Discrimination Index – measure of item Test Revision
discrimination o Characterize each item according to its strength
▪ Measure of the difference between the and weaknesses
proportion of high scorers answering an item o As revision proceeds, the advantage of writing a
correctly and the proportion of low scorers large item pool becomes more apparent
answering the item correctly because some items were removed and must be
▪ Extreme Group Method – compares people replaced by the items in the item pool
who have done well with those who have o Administer the revised test under standardized
done poorly conditions to a second appropriate sample of
▪ Discrimination Index – difference between examinee
these proportion o Cross-Validation – revalidation of a test on a
▪ Point-Biserial Method – correlation between sample of testtakers other than those on who
a dichotomous variable and continuous test performance was originally found to be a
variable valid predictor of some criterion
▪ Often results to validity shrinkage
o Validity Shrinkage – decrease in item validities
that inevitably occurs after cross-validation
o Co-validation – conducted on two or more test
using the same sample of testtakers
Psychological Assessment
Test Development
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Co-norming – creation of norms or the revision
of existing norms
o Anchor Protocol – test protocol scored by highly
authoritative scorer that is designed as a model
for scoring and a mechanism for resolving
scoring discrepancies
o Scoring Drift – discrepancy between scoring in
an anchor protocol and the scoring of another
protocol
o Differential Item Functioning – item functions
differently in one group of testtakers known to
have the same level of the underlying trait
o DIF Analysis – test developers scrutinize group
by group item response curves looking for DIF
Items
o DIF Items – items that respondents from
different groups at the same level of underlying
trait have different probabilities of endorsing a
function of their group membership
end
Psychological Assessment
Intelligence
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Intelligence Factor Analytic Theories
o Intelligence – multifaceted capacity that o Factor Analysis – group of statistical techniques
manifests itself in different ways across the life designed to determine the existence of
span underlying relationships between sets of
o Includes the abilities to: variables
✓ Acquire and apply knowledge o Charles Spearman – found that measures of
✓ Logical reasoning intelligence tended to correlate to various
✓ Planning degrees with each other
✓ Inferences ▪ Two-Factor Theory of Intelligence –
✓ Sound Judgments and problem solving intelligence have two components: (g)
✓ Grasp and visualize concepts general intelligence and (s) specific
✓ Attentions intelligence
✓ Intuition ▪ Tests that exhibited high positive correlation
✓ Finding the right words and thoughts with with other intelligence tests were thought to
facility be highly saturated with g, whereas test with
✓ Cope with, adjust to, and make the most of low or moderate correlations with other
new situations intelligence tests were viewed as possible
o Young children may define intelligence in terms measures of specific factors
that emphasize positive interpersonal skills ▪ The greater the magnitude of g in a test of
o Older children are more emphasized with their intelligence, the better the test was thought to
academic skills predict overall intelligence
o Galton: most intelligent persons are those who ▪ G Factor – linked to general ability
are equipped with the best sensory abilities ▪ S Factor – linked to specific ability
o Weschler: there is an explicit reference to an ▪ Group Factors – neither as general as g nor
aggregate or global capacity as specific as s
o Piaget: intelligence may be conceived of as a kind o Guilford – sought to explain mental activities by
of evolving biological adaptation deemphasizing, if not eliminating, any reference
o Binet: Intelligence as the capacity to find and to g
maintain a definite direction or purpose, to make o Thurstone – conceived of intelligence as being
necessary adaptations, and to engage in self- composed of seven “primary abilities”
criticism so that necessary first step in ▪ verbal comprehension, word fluency,
developing a measure of intelligence number facility, spatial visualization,
▪ Age Differentiation – refers to the simple fact associative memory, perceptual speed, and
that one can differentiate older children from reasoning
younger children o Gardner – developed a theory of multiple
▪ General Mental Ability – total product of the intelligences: logical-mathematical, bodily-
various separate and distinct elements of kinesthetic, linguistic, musical, spatial,
intelligence interpersonal, and intrapersonal
o Interactionism – refers to the complex concept ▪ Interpersonal Intelligence – ability to
by which heredity and environment are understand other people
presumed to interact and influence the ▪ Intrapersonal – capacity to form an accurate,
development of intelligence veridical model of oneself and to be able to
o Thurstone: Primary Mental Abilities, developed use that model to operate effectively in life
Primary Mental Abilities Test ▪ Emotional Intelligence – Interpersonal and
o Factory-Analytic Theories – focus on identifying Intrapersonal Intelligence
the ability or groups of abilities deemed to o Raymond B. Cattell – postulated Crystallized and
constitute intelligence Fluid Intelligence
o Information-Processing Theories – identifying ▪ Crystallized – acquired skills and knowledge
specific mental processes that constitute that are dependent on exposure to a
intelligence
Psychological Assessment
Intelligence
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
particular culture as well as on formal and o Aleksandr Luria – two basic types of
informal education information-processing styles, simultaneous
▪ Fluid – nonverbal, relatively culture-free, and and successive, have been distinguished
independent of specific instruction ▪ Simultaneous (Parallel) – information is
o Horn – added Visual Processing, Auditory integrated all at one time
Processing, Quantitative Processing, Speed of ▪ Successive (Sequential) – each bit of
Processing, facility with reading, and writing, information is individually processed in
short-term memory, and long-term memory sequences
storage and retrieval o PASS model – Planning, Attention,
▪ Vulnerable Abilities – decline with age and Simultaneous, and Successive
tend not to return to preinjury levels ▪ Planning – refers to strategy development for
following brain damage problem solving
▪ Maintained Abilities – tend not to decline with ▪ Attention – refers to receptivity to
age and may return to preinjury levels information
following brain damage Measuring Intelligence
o Three-Stratum Theory of Cognitive Abilities Stanford-Binet Intelligence Scales: Fifth Edition
▪ Top stratum level is general intelligence (g) (SB5)
▪ The second stratum is composed of eight o Lewis Terman published an English translation
abilities: fluid intelligence, crystallized of the first Stanford-Binet Intelligence Scale,
intelligence, general memory and learning, which stimulated a worldwide appetite for
broad visual perception, broad auditory intelligence tests
perception, broad retrieval capacity, broad o The first version was the first published
cognitive speediness, and intelligence test that provide detailed
processing/decision speed administration and scoring
▪ Hierarchical Model o 1908: Introduced the concept of age scale and
mental age
o It was also the first test to employ the concept of
IQ and the first test to introduce the concept of
Alternate Item (item to be substituted for regular
item under specified conditions)
o 1916: Intelligence Quotient (recommended by
Stern) – subject’s mental age in conjunction with
his or her chronological age
o Also employed the concept of Mental Age (the
age level at which an individual appears to be
o Cattell-Horn-Carroll Model of Cognitive functioning intellectually as indicated by the
Abilities – a psychometric taxonomy designed to level of items responded correctly)
explain how and why individuals differ in ▪ Ratio IQ: ratio of the testtaker’s mental age
cognitive ability. It provides a common frame of over his chronological age, multiplied to 100
reference and nomenclature to organize o 1926: Terman revised the test with Maud Merrill
cognitive ability research which included the development of two
o E.L Thorndike – intelligence can be conceived in equivalent forms labeled (L) and (M), as well as
terms of three clusters of ability: social new types of tasks for use with pre-school level
intelligence, concrete intelligence, and abstract and adult-level testtakers
intelligence o 1956 (Terman’s Death): SB was again revised
▪ Also incorporated a general mental ability with only a single form (L-M) and it included the
factor into the theory items considered to be the best from the two
▪ One’s ability to learn is determined by the forms
number and speed of the bonds that can be ▪ Deviation IQ instead of ratio IQ tables
marshaled
Information-Processing View
Psychological Assessment
Intelligence
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
▪ Deviation IQ – reflects a comparison of the
performance with the performance of others
of the same age
o 4th edition introduced Point Scale (a test
organized into subtests by category of item, not
by age at which most testtakers are presumed
capable of responding in the way that is keyed as
correct
o Fifth Edition was designed for administration to
assess as young as 2 and as old as 85
▪ Based on the Cattell-Horn-Carrol theory
▪ Fluid Intelligence, Crystallized Intelligence,
Quantitative Knowledge, Visual processing,
Short-Term Memory
o Routing Test – a task used to direct or route the
Wechsler Tests
examinee to a particular level of questions to
o 1939: Wechsler-Bellevue 1 (W-B 1) was published
direct an examinee to test items that have a high
o Point scale rather than age scale
probability of being at an optimal level of
o Arranged in order of increasing difficulty
difficulty
o 1942: W-B 2 was created but never thoroughly
o Teaching Items – designed to illustrate the task
standardized
required and assure the examiner that the
o Standardization was restricted
examinee understands
o Some subtest lacked of inter-item reliability
o Basal Age – the highest year level at which the
o Some of the subtests were made up of easy
subject successfully passes all tests
items
o SB5 items are not timed to accommodate
o Scoring criteria for some items were too
testtakers with special needs and to fit them with
ambiguous
IRT model used to calibrate the difficulty of items
o 1955: Weschler Adult Intelligence Scale (WAIS)
o Exemplary for Adaptive Testing (testing
was published
individually)
▪ Yielded a verbal IQ, performance IQ, and a full
▪ Ensures that early test or subtest items are
scale IQ
not so difficult as to frustrate the testtaker
o 1981: WAIS-R
and not so easy as to lull the testtaker into a
o WAIS IV – the current Wechsler Adult Scale
false sense of security or a state of mind
▪ Made up of : Core and Supplemental Subtest
▪ Allows the test user to collect the maximum
▪ Core Subtests – administered to obtain
amount of information in minimum amount of
composite score
time
▪ Supplemental Subtest – for providing
▪ Facilitates rapport
additional clinical information or extending
▪ Minimizes potential for examinee fatigue
the number of abilities or processed
from being administered too many times
sampled
o Extra-test Behavior – the way examinee copes
▪ Core Subtests: Block design, similarities,
with frustration, reaction to stimulus, amount of
digit span, matrix reasoning, vocabulary,
support, etc. are observed
arithmetic, symbol search, visual puzzles,
information, and coding
▪ Supplemental Subtests: Letter-Number
Sequencing, Figure Weights,
Comprehension, Cancellation, and Picture
Completion
▪ more explicit administration instructions as
well as the expanded use of demonstration
and sample items
Psychological Assessment
Intelligence
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
▪ Practice Items (teaching items) are “disregarding, insofar as possible, the degree of
presumed to pay dividends in terms of instruction which the subject possesses
ensuring that low scores are actually due to o Culture-Free Intelligence Test – cultural factors
a deficit of some sort and not simply to a can be controlled then differences between
misunderstanding of directions cultural groups will be lessened
Short Forms of Intelligence Tests o Elimination of verbal items and the exclusive
o Short Form – a test that has been abbreviated in reliance on nonverbal, performance items can
length to reduce time needed for administration, control the effect of culture
scoring and interpretation o Nonverbal items were thought to represent the
o David Wechsler endorsed the used of short best available means for determining the
forms for screening purposes (not to make cognitive ability of minority group
placement or educational decisions) o Culture-Loading – the extent to which a test
o Wechsler Abbreviated Scale of Intelligence – incorporates the vocabulary, concepts,
designed to answer the need for a short traditions, knowledge, and feelings associated
instrument to screen intellectual ability of with a particular culture
testtakers from 6 to 89 yrs of age o Culture-Fair Intelligence Test – a test or
o 2011: WASI-2 assessment process designed to minimize the
Group Test of Intelligence influence of culture, with regard to various
o Robert Yerkes lead the development of Army aspects of the evaluation procedures
Alpha and Army Beta to measure the ability to be ▪ Include only those tasks that seemed to
a good soldier during WWI reflect experiences, knowledge, and skills
o WWII led to development of Army General common to all different cultures
Classification Test ▪ Lack the hallmark of traditional tests of
o School Ability Test – group intelligence Test intelligence: predictive validity
▪ Alert educators to students who might profit o Flynn Effect – progressive rise in the IQ scores
function of data from more extensive that is expected to occur on a normed
assessment with individually administered intelligence from the date when the test was first
ability tests normed
o Four terms common to many measures of ▪ More on fluid intelligence
creativity: end
a. Originality – the ability to produce something
that is innovative or nonobvious
b. Fluency – ease with which responses are
reproduced and is usually measured by the
total number of responses produced
c. Flexibility – variety of ideas presented and the
ability to shirt from one approach to another
d. Elaboration – richness of detail in a verbal
explanation or pictorial display
o Convergent Thinking – deductive reasoning
process that entails recall and consideration of
facts as well as a series of logical judgments to
narrow down solutions and eventually arrive at
one solution
o Divergent Thinking – free to move in many
different directions, making several solutions
possible
Issues in the Assessment of Intelligence
o Binet-Simon test was designed to separate
“natural intelligence from instruction” by
Psychological Assessment
Education
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Achievement Tests o Culture Fair Intelligence Test (CFIT) – provide
o Achievement Tests – designed to measure an estimate of intelligence relatively free of
accomplishment/past learnings cultural and language influences
o Designed to measure the degree of learning ▪ Covers three levels and randomly selected
that has taken place as a result of exposure to adults
a relatively defined experience Aptitude Tests
o Adequately samples the targeted subject o Aptitude Tests – tend to focus more on informal
matter and reliably gauges the extent to which learning or life experiences
the examinees have learned it o Also referred as Prognostic Tests – used to
o Help in making decisions about placements, predict future behavior
gauging the quality of instruction in particular o Checklist – questionnaire on which marks are
institution, screen for difficulties and etc. made to indicate the presence or absence of a
o Wechsler Individual Achievement Test – III behavior
Edition (WIAT-III) – designed for use in the o Rating Scales – form completed by an evaluator
schools as well as clinical and research to make a judgement of relative standing with
settings regard to specific variable or list of variables
▪ Potential to yield actionable data relating to o Informal Evaluation – nonsystematic, relatively
student achievement in academic areas brief and off the record assessment leading to
such as reading, writing, etc. the formation of an opinion or attitude
o The test most appropriate for use is the one conducted by any person, in any way, for any
most consistent with the educational objectives reason, in an unofficial context that is not
of the individual teacher or school system subject to ethics or other standards of an
o Curriculum-Based Assessment – used to refer evaluation by a professional
to assessment of information acquired from o Tests such as WPPSI-III (Wechsler Preschool
teaching at school and Primary Scale of Intelligence) and
o Curriculum-Based Measurement – Stanford-Binet 5 Edition may be used to gauge
characterized by the use of standardized developmental strengths and weaknesses by
measurement procedures to derive local sampling children’s performance in cognitive,
norms to be used in the evaluation of student motor, and social/behavior content areas
performance o The most obvious example of Aptitude
o Different types of Achievement Test: TestꟷScholastic Aptitude Test (SAT)
a. Fact-Based Items – one that draws Diagnostic Tests
primarily on facts and how to apply those o Evaluative – applied to tests or test data that
facts are used to make judgments
b. Conceptual Items – designed to measure o Diagnostic Information – used in educational
mastery of the material context is typically applied to tests or test data
o Raven Progressive Matrices (RPM) – one of the used to pinpoint a student’s difficulty
best known and most popular nonverbal group o Diagnostic Test – used to identify areas of
tests deficit to be targeted for intervention
▪ Suitable test anytime one needs an estimate o Woodcock Reading Mastery Tests-Revised
of an individual’s general intelligence (WRMT-III) – measure of reading readiness,
▪ Original RPM has 60 items, which were reading achievement, and reading difficulties
believed to be of increasing difficulty o Stanford Diagnostic Mathematic Test, Fourth
o Goodenough-Harris Drawing Test (G-HDT) – Edition (SDMT4)
one of the quickest, easiest, and least o KeyMath Diagnostic System
expensive to administer of all ability tests o Brazelton Neonatal Assessment Scale –
▪ Draw a picture of a whole man and to do the individual tests for infants between 3 days and
best job possible 4 weeks of age to provide an index of newborn’s
▪ Each detail is given one point competence
Psychological Assessment
Education
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
▪ Reflexes, response to stress, startle vocabulary, presumably providing a nonverbal
reactions, cuddliness, motor maturity, ability estimate of verbal intelligence
to habituate to sensory stimuli, and hand- o Leiter International Performance Scale-3rd
mouth-coordination Edition – provide a nonverbal measure of
o Gesell Developmental Schedule – one of the intelligence in individuals 3 to 75 years and
oldest and the most established infant older
intelligence measures o Porteus Maze Test – poorly standardized
▪ Gesell Maturity Scale, Gesell Developmental nonverbal performance measure of
Observation, Yale Tests of Child intelligence
Development o Illinois Test of Psycholinguistic Abilities (ITPA-
▪ Provide an appraisal of the developmental 3) – designed for use with children ages 5
status of the children from 2.3 months to 6.3 through 12 years old
years of age o Bender Visual Motor Gestalt Test – consist of
▪ Developmental Quotient – determined by a nine geometric figures that the subjects is
test score, evaluated by assessing the simply asked to copy
presence and absence of behavior ▪ Anyone older than 9 who cannot copy the
associated with maturation figures may suffer from some type of deficit
o Bayley Scales of Infant and Toddler end
Development – assesses on normative
maturational developmental data, designed for
infants between 1 and 42 months old and
assesses development across five domains:
cognitive, language, motor, socioemotional, and
adaptive
o Cattell Infant Intelligence Scale – designed as a
downward extension of Stanford-Binet Scale
for infants and preschoolers between 2 and 30
months of age, measure intelligence of infants
and young children
Psychoeducational Test Batteries
o Psychoeducational Test Batteries – generally
contain two types: those that measure abilities
related to academic success and those that
measure educational achievement
o Kaufman Assessment Battery for Children (K-
ABC) – designed for ages 2 ½ through 12 ½ that
measures both intelligence and achievement
o KABC-II was published in 2004 and the age
range was extended up to 18 years old
o Woodcock-Johnson IV (WJ IV) – consisting
three co-normed batteries: achievement,
cognitive, and oral language ability
o Columbia Mental Maturity Scale-Third Edition
(CMMS) – scale requires the subject to
discriminate similarities and differences by
indicating which drawing does not belong on a
6-by-9 inch card containing 3-5 drawings,
depending on the level of difficulty
o Peabody Picture Vocabulary Test-4th Edition
(PPVT-IV) – measure hearing or receptive
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
NOTE:
LEVEL A: Anyone under a direction of a supervisor or consultant
LEVEL B: Psychometricians and Psychologists only
LEVEL C: Psychologists only
Name of the Test Acronym Age Group Level Type of Test Developer Description
Individual Test Administration (Verbal and Nonverbal Tests)
Stanford-Binet SB-5 2yrs – 89 C Janzen, H., o Individually administered
Intelligence Scale- yrs old Obrzut, J., & o Norm-referenced of cognitive
Fifth Edition Marusiak, C. abilities
(2004) o Administered by clinicians
o Scales: Verbal, Nonverbal and Full
First translated Scale (FSIQ)
by Lewis o Nonverbal and Verbal Cognitive
Terman (1916) Factors: Fluid Reasoning, Knowledge,
Quantitative Reasoning, Visual-
First developed Spatial Processing, and Working
by Alfred Binet Memory
and Theodore o Early SB-5 provides lower cost
Intelligence
Simon version of the test for preschool
assessment
o Useful for measuring individuals with
scores in both giftedness and
intellectual disability ranges
o Useful for assessing persons with
limited English proficiency, deaf and
hard of hearing conditions, nonverbal
disabilities, ADHD, traumatic brain
injury, and ASD
o Provides only a single general
intelligence score
Weschler Intelligence WISC-V 6 yrs – 16 C o Individually administered
Scale for Children yrs old and o Provides subtest and composite
Fifth Edition 11 months scores that represent intellectual
Intelligence David Wechsler functioning in specific cognitive
domains, as well as a composite
score that represents the general
intellectual ability
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o 10 Primary subtests are
recommended for comprehensive
description of intellectual ability
o 6 Secondary subtests can be
administered in addition to primary
subtests to provide a broader
sampling of intellectual functioning
and to yield more information for
clinical decision making
o It is possible for intellectual abilities
to change over the course of
childhood due to motivation,
attention, interest, and opportunities
for learning
Wechsler Adult WAIS-IV 16 yrs – 90 C o Designed to measure intelligence and
Intelligence Scale yrs old and cognitive ability in adults and older
11 months adolescents
o Developed to address weaknesses in
Stanford-Binet
o Contains some time subtests
o Provides a number of different scores
o 10 core subjects, 5 supplemental
subtests
o 4 Major Scores: Perceptual
Reasoning, Processing Speed, Verbal
Intelligence Comprehension, Working Memory
o Provides two overall summary
scores including Full-Scale IQ and
General Ability Index
o Frequently used intelligence test
o Assess cognitive functioning in
people with psychiatric conditions
o Assess functioning in people with
brain injury
o Evaluating patterns of brain
dysfunction
o Diagnostic Purposes
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Wechsler Preschool WPPSI- 2 yrs old C o Individually administered, norm-
and Primary Scale of IV and 6 referenced instrument
Intelligence Fourth months – 7 Intelligence o Short, game-like tasks that engage
Edition yrs old and young children
7 months
Group Test Administration (Nonverbal Tests)
Colored Progressive CPM 5 yrs – 11 B o Used to assess the degree to which
Matrices yrs old children and adults can think clearly,
or the level to which their intellectual
Intelligence
abilities have deteriorated
o 12-items
o Individual or by group
Raven’s Progressive RPM 4 yrs -90 B o Provides clear-thinking ability and
Matrices yrs old intellectual capacity that minimizes
the impacts of language skills and
Intelligence John Raven cultural differences
(1936) o Multiple choice intelligence test of
abstract reasoning
o Group test
Standard Progressive SPM 6 yrs – 16 B o More difficult than CPM
Matrices yrs old, 17 o Appropriate for children and teens
Intelligence
yrs - older with each item becoming
progressively more difficult
Advanced Progressive APM 12+ yrs old B o Geared towards adults and teenagers
Matrices Intelligence of advanced intelligence
o 40-60 mins
Culture Fair CFIT Scale 1: 4-8 B Raymond o Nonverbal instrument to measure
Intelligence Test yrs old Cattell (1949) your analytical and reasoning ability
Scale 2: 8- in the abstract and novel situations
14 yrs old o Measures individual intelligence in a
Scale 3: 14+ manner designed to reduced, as
Intelligence
yrs old much as possible, the influence of
culture
o Individual or by group
o Aids in the identification of learning
problems and helps in making more
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
reliable and informed decisions in
relation to the special education
needs of children
o Other uses: selection of students
qualified for accelerated educational
programs, advising students,
increasing effective of vocational and
guidance decisions
o Level B
Purdue Non-Language PNLT 13+ yrs old B Joseph Tiffin, o Designed to measure mental ability,
Test Allen Gruder, since it consists entirely of geometric
Intelligence and Kay Inaba forms
(1958) o Culture-fair
o Self-Administering
Panukat ng PKP 16+ yrs old Aurora R. o Basis for screening, classifying, and
Katalinuhang Pilipino Palacio Ed.D. identifying needs that will enhance
(1991) the learning process
o In business, it is utilized as predictors
of occupational achievement by
gauging applicant’s ability and fitness
Intelligence for a particular job
o Essential for determining one’s
capacity to handle the challenges
associated with certain degree
programs
o Subtests: Vocabulary, Analogy,
Numerical Ability, Nonverbal Ability
Differential Aptitude DAT 16+ yrs old B George K. o Used to determine and measures
Test Bennet, individual’s ability to acquire, through
Grade 7 and Harrold G. future training, some specific set of
older Seashore, skills
Aptitude
Alexander G. o Verbal Reasoning, Numerical Ability,
Wesman (1947) Abstract Reasoning, Spelling, and
Language Use
o Individual or group
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Thurstone Test of TMA Adults Louis Leon o Four Job-Related Tasks assessed by
Mental Alertness Thurstone TMA Test: Adjusting to new situations,
Ability (1943) learning new skills quickly,
understanding complex or subtle
relationships, thinking flexibly
Wonderlic Cognitive WPT- Adults Eldon F. o Assessing cognitive ability and
Ability Tests R/WCAT Wonderlic problem-solving aptitude of
(Wonderlic Personnel Ability (1939) prospective employees
Test) o Multiple choice, answered in 12
minutes
Watson Glaser Critical W-GCTA 20 yrs – 64 Goodwin o Designed to assess a person’s critical
Thinking Test yrs old Watson and thinking abilities and is widely used
Edward Glaser across legal practices
(1964) o Helps law firms to create shortlist of
candidates deemed likely to have
what it takes for training
o Assesses ability for critical thinking,
creating conclusions, analyzing
strong and weak arguments,
recognizing assumptions, and
evaluating arguments
o Multiple choice
Flanagan Industrial Adults A Flanagan, J.C. o use for personnel selection
Tests (1960) programs, based on the identified job
elements
o Job Elements: Arithmetic, Assembly,
Components, Coordination,
Electronics, Expression, Ingenuity,
Aptitude Inspection, Judgment and
Comprehension, Mathematics and
Reasoning, Mechanics, Memory,
Patterns, Planning, Precision, Scales,
Tables, Vocabulary
o Individual and Group Administration
o Level A
Philippine Aptitude PACT 14 yrs – 15 Center for o Developed to measure student’s
Aptitude
Classification Test yrs old Educational abilities and help students decide on
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
System (Ma. the course they will take after high
Lourdes M. school
Franco) o Assumes that aptitudes are required
in different combinations and in
varying degrees for successful
performance in different post-
secondary courses
o 18 subtests
o Individual or by group
Otis-Lennon Ability OLSAT 5 yrs – 18 C Otis, A.S., & o designed to assess general mental
Test yrs old Lennon, R.T. ability or scholastic aptitude of pupils
Ability (1967) o Individual or by group
o Level C
o Identify gifted children
Armed Services ASVAB o Most widely used aptitude test in US
Vocational Aptitude o multiple-aptitude battery that
Battery Aptitude measures developed abilities and
helps predict future academic and
occupational success in the military
Personality Inventory Tests/Projective Tests
Minnesota Multiphasic MMPI-2 18 yrs old C Hathaway & o Multiphasic personality inventory
Personality Inventory and older Mckinley (1943) intended for used with both clinical
and normal populations to identify
sources of maladjustment and
personal strengths
o Self-report
o 567 true or false questions
o 60-90 minutes
Personality o Help in diagnosing mental health
disorders
o Clinical Scales: Hypochondriasis,
Depression, Hysteria, Psychopathic
Deviate, Masculinity/Femininity,
Paranoia, Psychasthenia (Anxiety,
Depression, OCD), Schizophrenia,
Hypomania, Social Introversion
o High in L scale = faking good
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o High in F scale = faking bad, severe
distress or psychopathology
o K Scale = reveals a person’s
defensiveness around certain
questions and traits
o “Cannot Say” (CNS) Scale = measures
how a person doesn’t answer a test
item
o True Response Inconsistency (TRIN) =
five true, then five false answers
o Varied Response Inconsistency (VRIN)
= random true or false
o Fp Scale = reveal intentional or
unintentional over-reporting
o FBS Scale = “symptom validity scale”
designed to detect intentional over-
reporting of symptoms
o S Scale = Superlative Self-
Presentation to see if you
intentionally distort answers to look
better
Basic Personality BPI Adults & C Douglas N. o Self-report measure of the general
Inventory Adolescents Jackson (1996) domain of psychopathology
o 240 true/false items, 11 substantive
clinical scales and one critical item
scale
o 35 mins
Personality
o Dimensions: Alienation, Anxiety,
Denial, Depression, Deviation,
Hypochondriasis, Impulse Expression,
Interpersonal Problems, Persecutory
Ideas, Self-Depreciation, Social
Introversion, Thinking Disorder
Myers-Briggs Type MBTI 13+ yrs old B Isabel Myers, o Self-report inventory designed to
Indicator Personality Katherine identify a person’s personality type,
Briggs (1943) strengths, and preferences
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Allow respondents to further explore
and understand their own
personalities
Edward’s Preference EPPS Adults B Edwards, A. L. o Objective, forced-choice inventory for
Personality Schedule (1959) assessing the relative importance
that an individual places on 15
Personality, personality variables
Attitude o Useful in personal counselling and
with non-clinical adults
o Individual
o Level 2 or B
NEO Personality NEO PI- 17 yrs – 89 B Paul T. Costa, o Standard questionnaire measure of
Inventory R yrs old Robert McCrae the Five Factor Model, provides
systematic assessment of emotional,
Personality interpersonal, experiential,
attitudinal, and motivational styles
o Self-Administered
o 30-40 minutes
NEO Five Factor NEO- 12+ yrs old B o Shorter, 60-item inventory that
Personality
Inventory FF-I measures 5 broad dimensions only
Pictorial Self-Concept PSC 3 yrs – 13 Jack Joseph o Allows clinician to measure self-
Scale for Children yrs old concept in children as young as 3
o Identifies children with negative self-
appraisal put them at risk for
academic and behavioral difficulties
o Let youngster respond using pictures
rather than words
o Children are shown pairs of
Personality
illustrations representing common
self-appraisal situations and are
asked to choose between a picture
representing negative and positive
self-concept
o Can be used to evaluate
psychological and educational
interventions
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Panukat ng Ugali at PUP 12 yrs – 18 Enriquez & o Indigenous personality test
Pagkatao/Panukat ng yrs old Guanzon (1993) o Tap specific values, traits and
Pagkataong Pilipino Personality behavioral dimensions related or
meaningful to the study of Filipinos
o 30 minutes
Millon Clinical MCMI-IV 18 yrs and C o Self-report assessment used to help
Multiaxial Inventory older diagnose and treat personality
Fourth Edition Personality disorders
o Measures: Typical Functioning,
Abnormal Types/Traits, Disordered
Sixteen Personality 16PF 16 yrs and Cattell (1946) o Evaluates a personality on two levels
Factor Questionnaire older of traits
o Primary Scales: Warmth, Reasoning,
Emotional Stability, Dominance,
Liveliness, Rule-Consciousness,
Social Boldness, Sensitivity,
Personality Vigilance, Abstractedness,
Privateness, Apprehension,
Openness to change, Self-Reliance,
Perfectionism, Tension
o Global Scales: Extraversion, Anxiety,
Tough-Mindedness, Independence,
Self-Control
Interest Tests
Strong Campbell SCII 13+ yrs old Edward Strong o Composed of 325 items with each
Interest Inventory Test Jr., David piece referring to an activity
Interest Campbell o Classifies testtakers into 6 general
(1927) occupational themes based on the
theory of John L. Holland (RIASEC)
Thurstone Interest Adults A Thurstone, L.L o Checklist by which a persona can
Schedule (1947) systematically clarify his
understanding of his vocational
interests
Interest/Aptitude
o Designed as a counseling instrument
to be used in situations in which the
client-counselor relationship is such
that straightforward and honest
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
expression of choices can be
expected
o Individual or by group
o Level A or 3
end
Psychological Assessment
Personality
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Personality and Personality Assessment administration or application of tools of
o McClelland: Personality is the most adequate assessment
conceptualization of a person’s behavior in all ▪ Personality Profile – the targeted
its detail characteristics are typically traits, states, or
o Menninger: Personality as the individual as a types
whole, his height and weight and love and hates o Personality State – inferred psychodynamic
and blood pressure and reflexes; his smiles disposition designed to convey the dynamic
and hopes and bowed legs and enlarged quality of id, ego, and superego in perceptual
tonsils. It means all that anyone is and that he conflict
is trying to become ▪ Indicative of relatively temporary
o Personality – individual’s unique constellation predisposition
of psychological traits that is relatively stable o Woodworth Personal Data Sheet – first
over time personality inventory every developed during
o Personality Assessment – the measurement WWI
and evaluation of psychological traits, states, Developing Instruments to Assess Personality
values, interests, attitudes, worldview, Personality Assessment: Basic Questions
acculturation, sense of humor, cognitive and o Who is being Assessed, and who is doing the
behavioral styles, and/or related individual assessment?
characteristics ▪ Self-Report – process wherein information
o Personality Traits – real physical entities that about the assessment is supplied by the
are bona fide mental structures in each examinee himself
personality (Allport, 1937) ▪ Commonly used to explore an assessee’s
▪ Trait is a generalized and focalized self-concept (one’s attitudes, beliefs,
neuropsychic system with the capacity to opinions, and related thoughts about
render many stimuli functionally equivalent, oneself)
and to initiate and guide consistent forms of ▪ Self-Concept Measure – an instrument
adaptive and expressive behavior designed to yield information relevant to
▪ Any distinguishable, relatively enduring way how an individual sees him or herself with
in which one individual varies from another regard to selected psychological variables
(Guilford, 1959) ▪ Self-Concept Differentiation – the degree to
o Personality Type – constellation of traits that is which a person has different self-concepts
similar in pattern to one identified category of in different roles
personality within a taxonomy of personalities ▪ People who are highly differentiated are
▪ Hippocrates: Melancholic, Phlegmatic, likely to perceive themselves quite
Sanguine, and Choleric differently in various roles
▪ Jung: basis for MBTI ▪ Low levels of self-concept differentiation
▪ Holland: Artistic, Enterprising, Investigative, tend to be healthier psychologically
Social, Realistic, or Conventional (Self- ▪ Unfortunately, some clients tend to paint
Directed Search Test for vocational distorted pictures of themselves
guidance) intentionally or unintentionally in self-report
▪ Friedman & Rosenman: Type A and Type B measures
▪ Type A – competitiveness, haste, ▪ Faking good or Faking Bad
restlessness, impatience, feelings of being ▪ Some testtakers may be impaired with
time pressured, and strong need for regard to their ability to respond accurately
achievement and dominance to self-report questions
▪ Type B – mellow and laid-back ▪ In some situations, the best available
▪ Profile – narrative description, graph, table, method for the assessment of personality
or other representation of the extent to involves reporting by a third party
which a person has demonstrated certain ▪ Leniency Error, Severity Error, Error of
targeted characteristics as a result of the Central Tendency, and Halo Effect
Psychological Assessment
Personality
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o What is assessed when a personality o Idiographic Approach – characterized by efforts
assessment is conducted? to learn about each individual’s unique
▪ Personality measures are tools used to gain constellation of personality traits
insight into array of thoughts, feelings, and o Revised NEO Personality Inventory (NEO PI-R)
behaviors is widely used in both clinical and research
o Response Style – tendency to respond to a test about personality assessment
item or interview question in some ▪ Measure five dimensions of personality and
characteristic manner regardless of the a total of 30 elements or facets that define
content of the item or the question each domain
▪ Impression Management – used to describe ▪ Neuroticism – taps the aspects of
the attempt to manipulate others’ adjustment and emotional stability
impressions through the selective exposure ▪ Extraversion – taps aspects of sociability
of some information ▪ Openness – openness to experience, active
▪ Social Desirable Response – present imagination, aesthetic sensitivity,
oneself in favorable light attentiveness to inner feelings, preference
▪ Acquiescence – agree with whatever is for variety, intellectual curiosity,
presented independence of judgment
▪ Nonacquiescence – disagree with whatever ▪ Agreeableness – altruism, sympathy,
is presented friendliness, and the belief that others are
▪ Deviance – make unusual or uncommon similarly inclined
response ▪ Conscientiousness – active process of
▪ Extreme – make extreme ratings on the planning, organizing and following through
scale ▪ Designed for use with persons 17 yrs of age
▪ Gambling/Cautiousness – guess or not and older, self-administered
guess when in doubt o Criterion – standard on which a judgment or
▪ Overly Positive – claim extreme virtue decision can be made
through self-presentation in a superlative o Criterion Group – reference group of testtakers
manner who share specific characteristics and whose
▪ Response style can affect the validity of the responses to test items serve as a standard
outcome according to which items will be included in or
o How are personality assessments structured discarded from the final version of a scale
and conducted? o Empirical Criterion Keying – process of using
▪ Scope of an evaluation may be very wide, criterion groups to develop test items
seeking to take a kind of general inventory o Minnesota Multiphasic Personality Inventory
of individual’s personality (MMPI) – collaboration of Starke R. Hathaway
▪ Locus of Control – person’s perception and John Charnley McKinley
about the source of things that happen to ▪ Contains 566 true-false items and was
him or her (external or internal) designed to aid to psychiatric diagnosis with
▪ Face-to-face, computer-administered, adolescents and adults
behavioral observation, paper-and-pencil, ▪ Scales: Hypochondriasis, Depression,
etc. Hysteria, Psychopathic Deviate, Masculinity-
▪ Structured or Unstructured Femininity, Paranoia, Psychasthenia,
▪ Frame of Reference – defined as the aspects Schizophrenia, Hypomania, Social
of the focus of exploration such as the time Introversion
frame, as well as other contextual issues ▪ MMPI-3 – latest version (2020)
o Nomothetic Approach – characterized by o California Psychological Inventory (CPI), 3rd
efforts to learn how a limited number of Edition – attempts to evaluate personality in
personality traits can be applied to all people normally adjusted individuals and thus finds
more use in counseling settings
Psychological Assessment
Personality
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
▪ Commonly used in research settings to o Positive and Negative Affect Schedules
examine everything from typologies of (PANAS) – developed by Watson, Clark and
sexual offenders Tellegen (1988) to measure two orthogonal
▪ Can be used in normal subjects dimensions of affect
o Factor Analysis – statistical procedure for o Cognitive Intervention for Stressful Situations –
reducing the redundancy in a set of developed Endler and Parker (1990), measures
intercorrelated scores coping styles by asking subjects how they
o Guilford-Zimmerman Temperament Survey – would respond to variety of stressful situation
reduces personality to 10 dimensions, each of Personality Assessment and Culture
which is measured by 30 different items o Acculturation – ongoing process by which an
o Sixteen Personality Factor Questionnaire individual’s thoughts, behaviors, values,
(16PF) – developed by Raymond Cattell worldview, and identity develop in relation to
Frequently Used Measures of Positive Personality general thinking, behavior, customs, and values
Traits of a particular cultural group
o Rosenberg Self-Esteem Scale – measures o Values – which an individual prizes or the
global feelings of self-worth using 10 simple ideals an individual believes in
and straightforward statements that o Instrumental Values – principles that help one
examinees rate on a 4-point likert scale attain some objective
o General Self-Efficacy Scale (GSE) – measure o Terminal Values – guiding principles and a
an individual’s belief in his or her ability to mode of behavior that is an endpoint objective
organize resources and manage situations, to o Identity – set of cognitive and behavioral
persist in the face of barriers, and to recover characteristics by which individuals define
from setbacks themselves as members of a particular group
o Ego Resiliency Scale Revised – developed by o Identification – process by which an individual
Block and Kremen (1996), consists of 14 items assumes a pattern of behavior characteristic of
and using4-point likert scale, rate statement other people, and referred to it as one of the
such as “I am regarded as a very energetic central issues that ethnic minority groups must
persion,” “I get over my anger at someone deal with
reasonably quick,” o Worldview – unique way people interpret and
o Dispositional Resiliency Scale (DRS) – make sense of their perceptions as a
developed by Bartone et. Al., (1989) to measure consequence of their learning experience,
“hardiness” – the ability to view stressful cultural background, and related variables
situations as meaningful, changeable Objective Methods
o Hope Scale – characterize hope as a goal o Usually administered by paper-and-pencil
driven energy (agency) in combination with the means or by computer
capacity to construct systems to meet goals o Characteristically contain short-answer items
(pathways) for which assessee’s task is to select one
▪ Also know as Dispositional Hope Scale response from the two or more provided
(Synder et al., 1991) o Scoring is done according to set procedures
o Life Orientation Test-Revised (LOT-R) – most involving little, if any, judgment on the part of
widely used self-report measure of the scorer
dispositional optimism, which is defined as an o Can usually scored quickly and reliably by
individual’s tendency to view the world and the varied means
future in positive ways o Objective personality tests typically contain no
o Satisfaction with Life Scale (SWLS) – developed one correct answer, but the selection from
as a multi-item scale for the overall multiple choice items provides info relevant to
assessment of life satisfaction as a cognitive- something about the testtaker
judgmental process, rather than for the o Self-report, though, can be subjective
measurement of specific satisfaction domains o Objective personality test are objective in a
sense that they employ short-answer format,
Psychological Assessment
Personality
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
one that provides little, if any, room for o Form: how accurately the individual’s
discretion in terms of scoring perception matches or fits the corresponding
Projective Methods part of the inkblot
o Projective Hypothesis – an individual supplies o Holtzman Inkblot Test – was created to meet
structure to unstructured stimuli in a manner these difficulties while maintaining the
consistent with the individual’s own unique advantages of inkblot methodology; alternative
pattern of conscious and unconscious needs, for Rorschach
fears, desires, impulses, conflicts, and ways of Thematic Apperception Test
perceiving and responding o Published by Christiana D. Morgan and Henry A.
o Projective Method – as a technique of Murray
personality assessment in which some o Originally designed as an aid to eliciting fantasy
judgment of the assessee’s personality is made material from patients in psychoanalysis
on the basis of performance on a task that o Consists of 31 cards, one of which is blank
involves supplying some sort of structure to o Should find the source of the examinee’s story
unstructured or incomplete stimuli o Apperceive – to perceive in terms of past
o Indirect methods of personality assessment perceptions
o Born in the spirit of rebellion against normative o The raw material used in deriving conclusions
data and through attempts by personality about the individual examined with TAT are the
researchers to break down the study of stories they were told by the examinee,
personality into the study of specific traits of clinicians notes about the way or the examiner
varying strengths in which the examinee responded to the cards,
Rorschach Inkblot Projective Test and the clinician’s notes about extra-test
o Developed by Hermann Rorschach behavior and verbalizations
o John Exner argued that inkblots are not o The last two categories of raw data are sources
completely ambiguous, the task does not of clinical interpretations for almost any
necessarily force projection, and that it has individually administered test
been mislabeled as projective test for far too o Many interpretive systems incorporate, or to
long some degree based on, Needs (determinants of
o Consists of 10 bilaterally symmetrical inkblots behavior arising from within the individual) and
printed on separate cards Thema (unit of interaction between needs and
o After the entire set of cards has been press)
administered once, a second administration o The guiding principle in interpreting TAT stories
(Inquiry) is conducted, wherein the examiner is that the testtaker is identifying with someone
attempts to determine what features of the in the story and that the needs, environmental
inkblot played a role in the formulation of demands, and conflicts of the protagonist in the
percept (perception of an image) story
o Testing the Limits – provide additional o Implicit Motive – nonconscious influence on
information concerning personality functioning behavior typically acquired on the basis of
o Rorschach protocols are scored according to experience
several categories, including location, Others
determinants, content, popularity, and form o Hand Test – consists of 9 cards with pictures of
o Location: part of the inkblot that was utilized in hand on them and tenth blank cards
forming the percept ▪ Responses are interpreted according to 24
o Determinants: qualities of the inkblot that categories such as affection, dependence,
determine what the individual perceives and aggression
o Content: content category of the response ▪ Edwin Wagner
o Popularity: refers to the frequency with which a o Rosenzweig Picture-Frustration Study –
certain response has been found to correspond employs cartoons depicting frustrating
with a particular inkblot or section of an inkblot situations and the testtaker is asked to fill in the
response of the cartoon figure being frustrated
Psychological Assessment
Personality
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
▪ Based on the assumption that the testtaker o Analogue Study – research investigation in
will identify with the person being which on or more variables are similar or
frustrated analogous to the real variable that the
o Apperceptive Personality Test – represent an investigator wishes to examine
attempt to address some long-standing o Analogue Behavioral Observation –
criticisms of the TAT as projective instrument observation of a person or persons in an
while introducing objectivity into the scoring environment designed to increase the chance
system that the assessor can observe the targeted
▪ Consists of 8 stimulus cards depicting behavior
recognizable people in everyday settings o Situational Performance Measure – allows for
o Word Association Tests – verbalized the first observation and evaluation of an individual
word that comes to mind; a Semistructured, under a standard set of circumstances
individually administered projective technique o Leaderless Group Technique – several people
of personality assessment that involves the are organized into a group for the purpose of
presentation of a list of stimulus words carrying out a task an observer records
▪ Kent-Rosanoff Free Association Test – one information related to individual group
of the earliest attempts to develop member’s initiative, cooperation, leadership,
standardized test using words as projective and related variables
stimuli o Role Play – acting an improvised part in a
▪ Carl Jung developed the first WAT simulated situation
o Sentence Completion Tests – semi-structured o Biofeedback – class of psychophysiological
projective technique of personality assessment assessment technique designed to gauge,
that involves presentation of a list of words that display, and record a continuous monitoring of
begin with a sentence and the assessee’s task selected biological processes
is to respond by finishing each sentence ▪ Plethysmograph – records changes in the
▪ Rotter Incomplete Sentences Blank (Rotter volume of a part of the body arising from
& Rafferty, 1950) variations in blood supply
▪ Sack’s Sentence Completion Test (Sacks & ▪ Penile Plethysmograph – designed to
Levy, 1950) measure changes in blood flow in penis
o Figure Drawing Test – produces drawing that is ▪ Polygraph – lie detector
analyzed on the basis of its content and related o Unobtrusive Measure – telling physical trace or
variables record without necessarily the presence of
▪ Draw-A-Person Test (Florence respondents
Goodenough, 1926) o Contrast effect – rating is affected by the
▪ House-Tree-Person Test (John Buck & previous or prior rating
Emmanuel Hammer, 1958) o Composite Judgment – averaging multiple
▪ Kinetic Family Drawing (Hulse, 1951, 1952) judgments
Behavioral Assessment Methods end
o Behavioral assessment – what a person does
in situation rather than on inferences about
what attributes he has more globally
o Timeline followback Methodology – designed
for use in the context of clinical interview for
the purpose of assessing alcohol abuse
o Behavioral Observation – watching the
activities of targeted clients or research
subjects
o Self-Monitoring – systematically observing and
recording aspects of one’s own behavior and/or
events related to that behavior
Psychological Assessment
Clinical and Counseling
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Overview goals, expectations, and mutual obligations
o Clinical Psychology – branch of psychology that with regard to a course of therapy
has its primary focus the prevention, diagnosis, ▪ Seasoned Interviewers – create a positive,
and treatment of severe abnormal behavior accepting climate in which to conduct the
o Counseling Psychology – concerned with the interview
prevention, diagnosis, and treatment of ▪ Effective interviewer conveys understand
abnormal behavior but more on everyday type to the interviewee verbally or nonverbally,
of concerns and problems includes attentive posture and facial
o Premorbid Functioning – level of psychological expression, as well as, acknowledging or
and physical performance prior to the summarizing what the interviewee is trying
development of a disorder, an illness, or a to say
disability ▪ A highly structured interview is one in
o Diagnostic and Statistical Manual – reference which all the questions asked are prepared
source for making clinical diagnoses for mental in advance; it provides uniformity in
disorders exploration and evaluation
▪ Latest ver: DSM-V and DSM-V-TR ▪ Stress Interview – any interview where one
▪ Lists all the criteria that have been met in objective is to place the interviewee in a
order to diagnose each of the disorder listed pressured state for some particular reason
▪ Conveys information about how extreme, ▪ Hypnotic Interview – interview under
problematic, troubling, odd, or abnormal the hypnosis; may more suggestible to leading
individual’s behavior is likely to be perceived questions and thus more vulnerable to
by others distortion of memories
o Incidence – rate of new occurrences of a ▪ Cognitive Interview – rapport is established
particular disorder or condition in a particular and the interviewee is encouraged to use
population imagery and focused retrieval to recall
o Prevalence – approximate proportion of information
individuals in a given population at a given point ▪ Collaborative Interview – allows the
in time who have been diagnosed or otherwise interviewee wide latitude to interact with
labeled with a particular disorder or condition the interviewer, collaborating on a common
o Biopsychosocial Assessment – mission of discovery, clarification, and
multidisciplinary approach to assessment that enlightenment
includes exploration of relevant biological, ▪ Types of Responses:
psychological, social, cultural, and a. Level-One – bear a little or no relationship
environmental variables for the purpose of to the interviewer’s response
evaluating how such variables may have b. Level-Two – communicates a superficial
contributed to the development and awareness of the meaning of a statement
maintenance of a presenting problems c. Level-Three – interchangeable with the
▪ Fatalism – belief that what happens in life is interviewee’s statement; minimum level of
largely beyond a person’s control responding that could help the interviewee
▪ Self-Efficacy – confidence in one’s own d. Level-Four and Level-Five – provide
ability to accomplish a task accurate empathy but also go beyond the
▪ Social Support – expressions of statement given
understanding, acceptance, empathy, love, e. Active Listening – power of understanding
advices, guidance, care, etc. response; foundation of good interviewing
o Interview – key tool of biopsychosocial skills for many different types of interviews
assessment ▪ Standard Questions during initial intake
▪ Guide decisions about what else needs to interviews:
be done to assess an individual 1. Demographic Data
▪ Therapeutic Contract – an agreement 2. Referral Question
between client and therapist setting forth 3. Medical History
Psychological Assessment
Clinical and Counseling
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
4. Present Medical Condition observed behavior is culture-specific and
5. Family Medical History arises from long-held family beliefs
6. Past Psychological history Special Applications of Clinical Measures
7. Past History with medical or psychological o MacAndrew Alcoholism Scale (MAC) and
professionals MacAndrew Alcoholism Scale-Revised –
8. Current psychological conditions personality and attitude variables thought to
▪ Mental Status Examination – a parallel to the underlie alcoholism
general physical examination conducted by o Addiction Potential Scale (APS) – personality
a physician is special clinical interview traits thought to underlie drug or alcohol abuse
conducted by a clinician for screening of o Addiction Acknowledgement Scale (AAS) –
intellectual, emotional, and neurological direct acknowledgement of substance abuse
deficits o Addiction Severity Index (ASI) – rater assess
▪ MSE starts the moment the interviewee severity of addiction in 7 problem areas:
enters the room by observing the medical condition, employment functioning,
appearance of the client and assessing their drug use, alcohol use, illegal activity,
Orientation family/social relations, and psychiatric
o Biographical and related data about an functioning
assessee may be obtained by interviewing the o Michigan Alcohol Screening Test (MAST) –
client and/or their significant others lifetime alcohol-related problems
o Clinicians and counselors may use different o Forensic Psychological Assessment – theory
tests that could help in the assessment and application of psychological evaluation and
▪ Millon Clinical Multiaxial Inventory-III measurement in a legal context
(MCMI-III) – Millon et al., 1994, yield scores ▪ In forensic situation, clinician may be the
related to enduring personality features as client of the third party (court) and not of the
well as acute symptoms assessee
▪ Beck Depression Inventory-II (BDI-II) – Beck ▪ Patient is compelled to undergo assessment
et al., 1996, tapping a specific symptom or ▪ It is imperative that the assessor rely not
attitude associated with depression only in the client’s representations but also
▪ Center for Epidemiological Studies on all available documentation
Depression Scale (CES-D) – self-report ▪ Determination of dangerousness is ideally
measure of depressive symptoms made on the basis of multiple data sources,
o Test Battery – group of tests administered including interview data, case history data,
together to gather information about an and formal testing
individual from a variety of instruments ▪ When dealing with potentially homicidal or
o Standard Battery – battery of tests suicidal assessees, the professional
Culturally Informed Psychological Assessment assessor must have knowledge of the risk
o Culturally Informed Psychological Assessment factors associated with such violent acts
– keenly perceptive of an responsive to issues ▪ The assessor has a legal duty to warn if ever
of acculturation, values, identity, worldview, he/she finds his/her client is about to do
language, and other culture-related variables homicide, a duty that overrides the
as they may impact the evaluation process privileged communication between
o Carefully read any existing case history data psychologist and client
which may provide answers to key questions ▪ Competence to stand Trial – defendant’s
regarding the assessee’s level of acculturation ability to understand the charges against
and other factors useful to know about in him and assist in his own defense
advance of any formal assessment ▪ The competency requirement protects an
o Shifting Cultural Lenses – tied to critical individual’s right to choose and assist
thinking and hypothesis testing, which permits counsel, the right to act as a witness on
the clinician to test another hypothesis that the one’s own behalf, and the right to confront
opposing witness
Psychological Assessment
Clinical and Counseling
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
▪ The person will be found to be incompetent o Anatomically Detailed Dolls – dolls that
if and only if she is unable to understand the accurately represent genitalia used for
charges against her and is unable to assist observation of children who suffered from child
in her own defense abuse
o Emotional Injury – term sometimes used o Elder Abuse – intentional affliction of physical,
synonymously with mental suffering, pain and emotional, financial, or other harm on older
suffering, and emotional harm individual
▪ Discrimination, harassment, malpractice, o Elder Neglect – failure to provide elder care
stalking, and unlawful termination of o Signs of suicidal ideation:
employment ✓ Talking about committing suicide
o Profiling – crime-solving process that draws ✓ Making reference to a plan for committing
upon psychological and criminological suicide
expertise applied to the study of crime scene ✓ One or more past suicide attempts
evidence
▪ Assuming that the perpetrators of serial The Psychological Report
crimes leave more than physical evidence at o Psychological reports may be as different as
a crime scene the reasons for undertaking assessment, in
o Custody Evaluation – psychological terms of no. of variables and etc.
assessment of parents or guardians and their o Barnum Effect – people tend to accept vague
parental capacity and/or of children and their personality descriptions as accurate
parental needs and preferencesꟷusually descriptions of themselves (Aunt Fanny Effect)
undertaken for the purpose of custody o Elements of Typical Psych Report:
Child Abuse and Neglect a. Demographics
o Abuse – refer to the creation of conditions that b. Reason for Referral
give rise to abuse of a child by an adult c. Test Administered
o Infliction of physical injury or emotional d. Findings
impairment that is nonaccidental e. Recommendations
o Creation or allowing the creation of substantial f. Summary
risk of physical injury or emotional impairment o Actuarial Assessment/Actuarial Prediction –
that is not accidental refer to the application of empirically
o Committing or allowing of sexual offence to be demonstrated statistical rules and probabilities
committed against a child as determining factor in clinical judgment and
o Neglect – failure on the part of adult to be actions
responsible of child care o Computerized Assessment – computerized
o Physical signs can be deceiving especially applications of clinical opinionꟷthat is, the
when they try to convince the panel it was from application of clinician’s judgments, opinions,
an accident, however, inappropriate clothing and expertise to a particular set of data as
for the season/event, poor hygiene, and lagging processed by the computer surface
physical development could be physical o Clinical Prediction – application of clinician’s
manifestations of abuse own training and clinical experience as
o Young sexually abused children could feel determining factor in clinical judgment and
discomfort in their sexual parts and older actions
children could manifest STDs or pregnancy o Mechanical Prediction – application of
o Emotional and behavioral indicators could empirically demonstrated statistical rules and
include fear of going hone or fear of adults, probabilities to the computer generation of
unusual reactions in response to other children findings and recommendations
crying, low self-esteem, social withdrawal, end
aggressiveness, etc.
o Tardiness in school, chronic fatigue, and
chronic hunger could be signs as well
Psychological Assessment
Clinical and Counseling
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Psychological Assessment
Neuropsychological Assessment
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
The Nervous System and Behavior o Ataxia – deficit in motor ability and muscular
o Neurology – focuses on nervous system and its coordination
disorders Neuropsychological Evaluation
o Neuropsychology – focus on the relationship o Hard Sign – defined as an indicator of definite
between brain functioning and behavior neurological deficit
o Neuropsychological Assessment – evaluation o Soft sign – indicator that is merely suggestive
of brain and nervous system functioning as it of neurological deficits
relates to behavior o Objective of the typical neuropsychological
o Behavioral Neurology – a subspecialty within evaluation is to draw inferences about the
the medical specialty of neurology that also structural and functional characteristics of a
focuses on brain-behavior relationships person’s brain by evaluating an individual’s
o Neurons – nerve cells behavior in defined stimulus-response
o Central Nervous System – consist of brain and situations
the spinal cord o Common to all thorough neuropsych exams are
o Peripheral Nervous System – consisting of the history taking, MSE, and administration of tests
neurons that convey messages to and from the and procedures designed to reveal problems of
rest of the body neuropsychological functioning
o Contralateral Control – each of the two o Neuropsychs must also have knowledge of the
cerebral hemisphere receives sensory possible effects of various prescription
information from the opposite side of the body medications taken by them assessees because
and also controls motor responses on the such medication can actually cause certain
opposite side of the body neurobehavioral deficits
o Neurological Damage – may take the form of a o Elements of Neuropsych Evaluation:
lesion in the brain or any other site within CNS 1. History Taking, Case History
or PNS 2. Interview
o Lesion – pathological alteration of tissue, such 3. Neuropsychological Mental Status
as that which could result from injury or Examination
infection o Noninvasive Procedures – procedures that do
o Focal – circumscribed at one site not involve any intrusion into the examinee’s
o Diffuse – scattered at various site body
o Brain Damage – general reference to any o May test for simple reflexes
physical or functional impairment in the CNS Neuropsychological Tests
o Acalculia – inability to perform arithmetic o Pattern Analysis – examiner looks beyond
calculations performance on individual tests to study of the
o Acopia – inability to copy geometric designs pattern of test scores
o Agnosia – deficit in recognizing sensory stimuli o Deterioration Quotient (DQ) – brain damage
o Agraphia – deficit in writing ability have devised various quotients based on
o Akinesia – deficit in motor movements patterns of subtest scores
o Alexia – inability to read o One symptom commonly associated with
o Amnesia – loss of memory neuropsychological deficit, regardless of the
o Amusia – deficit in ability to produce or site or exact cause of problem, is inability or
appreciate music lessened ability to think abstractly
o Anomia – deficit associated with finding words ▪ The Proverbs Test, Object Sorting Test,
to name things Color-Form Sorting Test, Wisconsin Card
o Anopia – deficit in sight Sorting Test-64 Card Version
o Anosmia – deficit in sense of smell o Executive Function – organizing, planning,
o Aphasia – deficit in communication due to cognitive flexibility, and inhibition of impulses
impaired speech or writing ability and related activities associated with the
o Apraxia – voluntary movement disorder in the frontal and prefrontal lobes of the brain
absence of paralysis
Psychological Assessment
Neuropsychological Assessment
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
▪ Towers of Hanoi, Mazes, Clock-Drawing
Tests
o Perceptual Test – general reference to any of
many instruments and procedures used to
evaluate varied aspects of sensory functioning
o Motor Test – reference to any of many
instruments and procedures used to evaluate
varied aspects of one’s ability and mobility
o Perceptual-Motor Test – any of many
instruments and procedures used to evaluate
the integration or coordination of perceptual
and motor abilities
▪ Bender Visual-Motor Gestalt Test
o Fixed Battery – group of test pre-modified
before the assessment
o Flexible Battery – consisting of an assortment
of instruments hand-picked for some purpose
relevant to the unique aspects of the patient
and the presenting problem
o Halstead-Reitan Neuropsychological Battery –
class neuropsychological test battery
o Severe Impairment Battery (SIB) – designed for
use with severely impaired assessees who
might otherwise perform at or near the floor of
existing tests
Other Tools of Neuropsychological Assessment
o fMRI – creates real-time moving images of
internal functioning
end
Psychological Assessment
Assessment, Careers, and Business
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
Career Choice and Career Transition ▪ Special Aptitude Test Battery – used to
Measure of Interest selectively measure aptitudes for a specific
o Interest Measure – instrument designed to line of work
evaluate testtakers’ likes, dislikes, leisure Measures of Personality
activities, curiosities, and involvements in o Guilford-Zimmerman Temperament Survey
various pursuits for the purpose of comparison (GZTS) and Edwards Personal Preference
with groups of members of various occupations Schedule (EPPS) may be preferred because the
and professions measurement they yield tend to be better
▪ To formulate job descriptions and attract related to the specific variables under study
new personnel o NEO PI-R and MBTI – most widely used
o Strong Interest Inventory – one of the first personality test in the workplace
measure of interest (G. Stanley Hall, 1907) o Integrity Test – specifically designed to predict
▪ Designed to assess children’s interest in employee theft, honesty, adherence to
various recreational pursuits established procedures, and/or potential for
o Strong Vocational Interest Blank – developed violence
by Edward K. Strong Jr. (1920s) o Applicant Potential Inventory (API) – can be
o Strong-Campbell Interest Inventory (SCII) – administered quickly and efficiently
developed under the direction of David P. o White (1984) suggested that preemployment
Campbell (1974) honesty testing may induce negative work-
o Self-Directed Search (SDS) – explores interest related attitudes
within Holland’s Theory “RIASEC” o Myers-Briggs Type Indicator – used to classify
o Minnesota Vocational Interest Inventory – assessees by psychological type and to shed
designed to compare respondents’ interest light on basic differences in the way human
patterns with those of persons employed in a being take in information and make decisions
variety of nonprofessional occupations o Issues about establishing relationship between
Measures of Ability and Aptitude personality and work performance:
o Wonderlic Personnel Test - measures mental a. How work performance is defined – there is
ability in general tests no single metric that can be used for all
▪ Includes items that assess spatial skill, occupations
abstract thought, and mathematical skill b. What aspect of personality to measure
▪ Useful in screening individuals for jobs that o High Conscientiousness = good work
require both fluid and crystallized performance
intellectual abilities o High Neuroticism = poor work performance
o Bennet Mechanical Comprehension Test – o High Extraversion = good work performance
widely used paper-and-pencil measure of Other Measures
testtakers’ ability to understand the o Checklist of Adaptive Living Skills (CALS) –
relationship between physical forces and survey the life skills needed to make a
various tools as well as other common objects successful transition from school to work
o Hand-Tool Dexterity Test – requires testtaker o Cross-Cultural Adaptability Inventory (CCAI) –
to actually take apart, reassemble or otherwise self-administered and self-scored instrument
manipulate materials, usually in prescribed designed to provide information on the
sequence and within time limit testtaker’s ability to adapt to other cultures
o General Aptitude Test Battery – available for o Career Transitions Inventory (CTI) – designed
use by state employment services as well as for use with people contemplating a career
other agencies and organizations change
▪ Identify aptitudes for occupations Screening, Selection, Classification, and Placement
▪ Consists of 12 timed tests that measure nine o Screening – relatively superficial process of
aptitudes, which in turn can be divided into evaluation based on certain minimal standards,
three components criteria, or requirements
Psychological Assessment
Assessment, Careers, and Business
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Selection – refers to a process whereby each o Assessment Center – widely used tool in
person evaluated for a position selection, classification, and placement
o Classification – does not imply acceptance or o Physical Test – measurement that entails
rejection but rather a rating, categorization, or evaluation of one’s somatic health and
“pigeonholing” with respect to two or more intactness
criteria ▪ Includes Drug Testing
o Placement – disposition, transfer, or Cognitive Ability, Productivity, and Motivation
assignment to a group or category that may be Measures
made on the basis of one criterion Measures of Cognitive Ability
o Resume – information related to one’s work o Cognitive-Based Tests are popular tools of
objectives, qualifications, education, and selection because they have been shown to be
experience valid predictors of future performance
o Letter of Application – attached with resume o It is in society’s interest to promote diversity in
which lets a job applicant demonstrate employment settings
motivation, businesslike writing skills, and his o Developers and users of cognitive tests in the
or her unique personality workplace to place demand for verbal skills
o Application Forms – biographical sketches that and abilities
supply employers with information pertinent to Productivity
the acceptability of job candidates, especially o Productivity – simply as output or value yielded
contact information which is useful for quick relative to work effort made
screening o Using techniques such as supervisor ratings,
o Letters of Recommendation – unique source of interviews with employees, and undercover
detailed information about the applicant’s past employees planted in the workshop,
performance, the quality of the applicant’s management might determine what is
relationship with peers, and so forth responsible for the unsatisfactory performance
o Interviews – whether individual or group in o Forced Distribution technique – involves
nature, provide an occasion for the face-to- distributing predetermined number or
face exchange of information percentage of assessees into various
▪ Factors that might affect the outcome of an categories that describe performance
employment: backgrounds, attitudes, o Critical Incidents Technique – involves the
motivations, perceptions, expectations, supervisor recording positive and negative
knowledge about the job, and interview employee behaviors
behavior of both the interviewer and the o Peer Ratings/Evaluations – valuable method of
interviewee identifying talent among employees
o Portfolio Assessment – entails evaluation of an Motivation
individual’s work sample for the purpose of o Maslow’s Theory argued that there is an
some screening, selection, classification, or hierarchy of human needs and after one
placement decision category is of need is met, people seek to
o Performance Tests – requires assessees to satisfy the next level
demonstrate certain skills or abilities under a o Alderfer proposed that once need is satisfied,
specified set of circumstances the organism may try to satisfy it further and
▪ To obtain a job-related performance sample frustrating one need might channel energy into
▪ Minnesota Clerical Test (MCT) – designed to satisfying a need at another level
measure clerical aptitude o McClelland described induvial with a high need
▪ Leaderless Group Technique – commonly for achievement as one who prefers a task that
used performance test in the assessment of is neither too simple nor extremely difficult
business leadership ability o Intrinsic Motivation – driving force comes from
▪ In-Basket Technique – used to assess within individual
managerial ability, organizational skills, and o Extrinsic Motivation – stems from rewards,
leadership potential external factors
Psychological Assessment
Assessment, Careers, and Business
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018)
o Work Preference Inventory – scale designed to development, advertising, and marketing of
assess aspects of intrinsic and extrinsic products and services
motivation o Attitudes formed about products, services, or
o Burnout – psychological syndrome of brand names are frequent focus of interest in
emotional exhaustion, depersonalization, and consumer attitude research
reduced personal accomplishment that can o Implicit Attitude – nonconscious, automatic
occur among individuals who work with other association in memory that produces a
people in some capacity disposition to react in some characteristic
▪ Emotional Exhaustion – inability to give of manner to a particular stimulus; “gut feeling”
oneself emotionally to others o Survey – fixed list of questions administered to
▪ Depersonalization – distancing from other a selected sample of persons for the purpose
people and even developing cynical attitudes of learning about consumer’s attitudes, beliefs,
toward them opinions, and/or behavior with regard to the
▪ Maslach Burnout Inventory, Third Edition targeted products, services or advertising
(MBI) – developed by Christina Maslach et o Consumer Panel – makes up a list of people or
al., 1996 families who have agreed to respond to
Job Satisfaction, Organizational Commitment, and questionnaires sent to them
Organization Culture ▪ Diary Panel – must keep detailed records of
o Attitude – presumably learned disposition to their behavior
react in some characteristic manner to a o Semantic Differential Technique – defining the
particular meaning and concepts of relating concepts to
Job Satisfaction one another in semantic space, the technique
o A pleasurable or positive emotional state entails graphically placing a pair of bipolar
resulting from the appraisal of one’s job or job adjectives on a seven point scale
experiences o Motivation Research Methods – typically
o Satisfied workers are more productive, more analyze motives for consumer behaviors and
consistent, in work output, and less likely to be attitudes
absent o Dimensional Qualitative Research – an
o Measures: Cognitive Evaluations, Work approach to qualitative research that seeks to
Schedule, Perceived Sources of stress, various ensure a study is comprehensive and
aspects of well-being, and mismatches systematic from psychological perspective by
between an employee’s cultural background guiding study design and proposed question for
and the prevailing organizational culture discussion on the basis of BASIC ID (Behavior,
Organizational Commitment Affect, Sensation, Imagery, Cognition,
o Refers to a person’s feelings of loyalty to and Interpersonal Relations and Drugs)
involvement in an organization end
o Organizational Commitment Questionnaire
Organizational Culture
o Totality of socially transmitted behavior
patterns characteristic of a particular
organization or company, including: the
structure of the organization and the roles
within it, leadership style, etc.
o Provides a way of coping with internal and
external challenges and demands
Other tools of Assessment for Business
Applications
o Consumer Psychology – branch of social
psychology that deals primarily with the
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
Terms to Remember:
o Descriptive Statistics – used to summarize, organize, and simplify data
o Inferential Statistics – measurement of the extent to which pairs of related values on 2 variables tend to change together; drawing conclusions or
inferences about population based on samples
▪ Allows to study samples and make generalizations to about the population selected
o Parametric Tests – assumed normal distributions; requires numerical scores for each participants
o Nonparametric Tests – little to no information about the population; participants are just classified into categories
o Independent Groups – sample values selected from one population are not related in any way to sample values selected from the other populations
o Dependent Groups – provide information about subjects in other groups
o ANOVA – used to evaluate mean differences between two or more treatments (populations)
o T-Test - used to test hypotheses about an unknown population mean and variance
o Chi-Square – measures the relationship between categorical/non-numerical data
Differences/Correlation Level of Data No. of Groups Description
Correlation
Pearson R Correlation Interval/Ratio + 2 o Quantifies a linear relation between two scale variables
Interval/Ratio o Single number is used to describe the direction and
(Continuous) strength of the relation between two variables
Spearman Rho’s Correlation Ordinal + 2 o Measure of the agreement between two rankings
Ordinal o More sensitive to error and discrepancies
(Ranking) o Larger values
o Calculations based on deviations
o Measures consistencyꟷwhen two variables are
consistently related, their ranks are linearly related
Kendall’s Coefficient of Correlation Ordinal + 3 or more o Measure of the agreement between two or more
Concordance W Ordinal + rankings
Ordinal o Has smaller gross error sensitivity and smaller
(Ranking) asymptotic variance
o Usually smaller values
o Calculations based on concordant and discordant pairs
o An index of interrater reliability of ordinal data
Point-Biserial Coefficient Correlation Nominal 2
(Dichotomous) +
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
Interval/Ratio
(Continuous)
Phi or Fourfold Coefficient Correlation Nominal + 2
Nominal
Rank Biserial Correlation Nominal 2
(Dichotomous) +
Ordinal
(Ranking)
Tetrachoric R Correlation Artificial 2
Dichotomous +
Artificial
Dichotomous
(supposed to be
interval/ratio
but labelled as
nominal)
Inferential Statistics
Z-Scores Differences Interval 1-2 o Describe the exact location of any specific sample mean
within distributions of sample means
T-test
1. Independent Samples Difference DV: 2 or more 2 o Also known as Independent T-Test, Independent
(Subjects) Measure T-Test, Independent Two-Sample T-Test,
Nominal Unpaired T-Test
o Between-Subjects design
IV: 2 o Compares the means of two independent groups in
(Treatments) order to determine whether there is a statistical
Interval/Ratio evidence that the associated population means
significantly differently
Experimental + o Must have equal variances, independent, and normally
Control distributed
o Parametric Test
2. Dependent Samples Difference DV: 1 (Subjects) 1 o Also known as Paired T-Test or Paired-Samples T-Test
Nominal o Within Subjects
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
IV: 2 o Compares means of two related groups to determine
(Treatment) whether there is a statistically difference between these
Interval/Ratio means
o Same participants are present in both groups
Before + After o Fewer subjects
o Suited for studying learning, development, or other
changes that takes place over time
o Reduces or eliminates problems caused by individual
differences
o Allow factors other than the treatment effect to cause a
participant’s score to change from one treatment to
next
o E.g., pre-test-post-test
3. Proportions/Percentages
4. Variances
5. 2 Correlation Coefficients Differences 2 o Used to assess the significance of the difference
between two correlation coefficients found in two
independent samples
6. One Sample T-Test Differences Interval/Ratio 1 o Whether the mean of a population is statistically
different from a known or hypothesized value
o Test variable’s mean is compared against a “test value”
which is a known or hypothesized value of the mean in
the population
Regression Equation
Linear Regression of Y on X Prediction Interval 2 o Y = a + bX
o A and B are unknown constants know as intercept and
Y – DV slope of the equation
X - IV o Used to predict the unknown value of variable Y when
value of variable X is known
o Y’s coefficient will change if there would be an
increase in X
Linear Regression of X on Y Prediction Interval 2 o X = c + dY
o Used to predict the unknown value of variable X using
Y – IV the known variable Y
X - DV
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
Coefficient of Determination
ANOVA
1.1 One-Way Within-Groups Differences DV: 1 1 o Compares means of 1-2 independent groups in order to
ANOVA determine whether there is statistical evidence that the
e.g., female associated population means are significantly different
o Analyze data from field studies, experiments, quasi
IV: 1 but more experiments
than 2 levels o If the grouping variable has two groups, then the
(Categorical) results of one-way ANOVA and independent samples
T-Test will be EQUIVALENT
e.g., 3pm (High o Denoted as F
Blood Pressure o Within-Subjects, Repeated Measures
or Low Blood o Determine whether the differences that are found
Pressure), 5pm between treatment conditions are significantly greater
(High Blood that would be expected if there is not treatment effect
Pressure or Low
Blood Pressure),
6pm (High
Blood Pressure
or Low Blood
Pressure)
1.2 One-Way Between- Differences DV: 2 or more 2 or more o Independent Groups, Between-Subjects
Groups ANOVA
E.g., Boys &
Girls
IV: 1 but more
than 2 levels
each
e.g., low,
medium, high
usage
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
2. Two-Way ANOVA Differences DV: 1 or more 1-2 o Hypothesis test that includes two nominal IV, regardless
(Factorial ANOVA) of their numbers of levels and a scale dependent
e.g., Age group variable
(Adolescence, o Mixed ANOVA
Young o Two-factor
Adulthood,
Middle
Adulthood)
IV: 1-2
regardless of
the no. levels
E.g., Therapy
Techniques &
Monthly report
of improvement
Multivariate Analysis of Differences o More than one dependent variable
Variance o Provides regression analysis and ANOVA for multiple
dependent variables by one or more factor variables or
covariates
Analysis of Covariance Differences o A covariate is included so that statistical findings reflect
effects after a scale variable has been statistically
removed
o Analyzes the differences between three or more groups
while controlling the effects of at least one continuous
covariate
Chi-Square
Goodness of Fit Differences Categorical o Measure of how well a statistical model fits a set of
observations
o High = values expected based on the model are close
to the observed values
o Allows to draw conclusions about the distribution of a
population based on a sample
Psychological Assessment
Intelligence, Developmental, Aptitude, Achievement Tests (Table)
Source: Cohen & Swerdlik (2018), Kaplan & Saccuzzo (2018), Gravetter & Wallnau (2013)
o Used when you want to test a hypothesis about
distribution one 1 categorical variable
o Data Binning – converting categorical variable to
continuous by separating variables into intervals
Independence Differences Categorical 2 o non-parametric
o used to determine whether your data are significantly
different from what you expected
o aka Chi-Square test of independence
o based on observed frequencies
Non-Parametric Tests
Median Test Differences Ordinal o used to test whether two (or more) independent groups
differ in central tendency (mean, median, mode)
Fischer’s Sign Test Differences Ordinal o alternative to paired t-test, chi-square
o knowing whether the proportions for one variable are
different among values of other variable
Wilcoxon Rank Sum Test Differences Ordinal o when the requirements for t-test for two independent
samples are not satisfied
o Paired T-test
Mann-Whitney (U) Test Differences Ordinal 2 o evaluates the difference between two groups of scores
(Independent T-Test) from an independent-measures design
o Unpaired T-test, Independent T-test
Wilcoxon Signed Ranks Tests Differences Ordinal 2 o evaluates the differences between two groups of scores
(T) (Paired T- from a repeated-measures design
Test/Dependent T-Test)
Kruskal-Wallis H Test (One- Differences Ordinal 3 or more o evaluates differences between three or more groups
Way Between Groups from an independent-measures design
ANOVA)
Friedman Rank (Repeated Differences Ordinal 3 or more o evaluates differences among three or more groups from
Measures ANOVA) a repeated-measures design
Spearman Rho (Pearson R) Ordinal