Industrial and
Organizational
Psychology
Performance
Appraisal
Copyright Paul E. Spector,
All rights reserved, March
15, 2005
Introduction
• Task of determining how
well your subordinates are
doing their jobs
• If you were a manager, how
would you do it?
• How would you appraise
their job performance?
• Would you watch them as
they perform their job?
• If you watch them, how
would you know what to
watch for?
• List the uses of job performance
information.
• Discuss the importance of criteria for
Learning performance appraisal.
• Describe the various methods of
Objectives performance appraisal, as well as their
advantages and limitations.
• Discuss how to conduct a legally
defensible performance appraisal.
Administrative
decisions
• Administrative decisions are based in part on job
performance of employees; both contracts and law
• Punishments; demotion, termination
(firing)
• Rewards; promotion, pay raise, develop
pay merit systems that tie raises to the
level of job performance
• E.g., In US, civil service (government)
employees only be fired for unsatisfactory
job performance or violation of work rules
Employee Development and
Feedback
• Information regarding their job performance in order to improve or maintain their
job skills.
– Supervisor’s role; what is expected of subordinates and how well they are
meeting these expectations
• High performing employees; feedback about how to perform better, keep good
performance, and enhance skills to get promoted
• Cannot occur if standards are not clear, or if the employees not given feedback.
• In many organizations supervisors provide annual performance reviews of their
subordinates (conducting an appraisal interview).
– Evaluating the performance of the subordinates
– Planning development activities for the subordinates
– Sharing knowledge about the evaluation system
– Solving performance problems
– Planning future actions
• New trend for companies; performance
management
– Go beyond the once per year evaluation by
designing a comprehensive performance
management system
Employee • Can include goal setting and periodic
coaching and feedback sessions between
development employee and supervisor
and feedback – Annual review for administrative purposes,
the interim reviews used only for feedback,
• Reduces anxiety and defensiveness
employees experience when being
evaluated for raises and promotions
Performance Appraisal versus
Performance Management
Performance Appraisal Performance Management
Once a year Frequent
Developed by HR Developed by managers &
employees
Feedback once a year Feedback at any time
Understanding employee’s Understanding the
level of effectiveness performance criteria in
relation to goals and how
employee’s behavior fits
those criteria
Why Do • Much of I/O activities and effort for
improving performance
We – E.g., designing better equipment, hiring
better people, motivating employees, and
Appraise training
Employees? • Performance data is the criterion/measuring
stick for evaluating such activities
Criteria – E.g., comparing employee performance
before and after the implementation of a
new program designed to enhance it
for – Better design; one group receives a new
Research procedure whereas the other does not;
compare their performances
• Before we can evaluate performance, must:
1. Set up a criterion or standard of
comparison that defines good performance.
2. Develop a procedure to compare
employees to this standard.
Performance • Cannot evaluate adequately someone’s
Criteria performance until know what performance
should be
– A criterion is a standard against which
you can judge the performance of
anything, including a person.
– Distinguish good from bad performance
Characteristics
of Criteria
• Criteria can be classified as either
theoretical or actual.
a. Theoretical criterion is the
definition of what good
performance is.
– It is a theoretical construct.
b. Actual criterion is the way in
which theoretical criterion is
assessed.
– Operational definition of the
theoretical criterion;
– OR the performance appraisal
technique used.
• Criteria can be classified as either theoretical
or actual.
– A close correspondence between
• For some jobs; e.g., insurance
salesman:
– Theoretical criterion; to sell;
– Actual criterion; count the sales
Characteristics the person made
of Criteria • But not for others; e.g., artist
– Theoretical criterion; produce
great works
– Actual criterion; asking art
experts for an opinion about the
person’s work
Characteristics of Criteria
Theoretical
Job Actual Criterion
Criterion
Artist Create great works Judgments of art experts
of art
Insurance Sell insurance Monthly sales
salesperson
Store clerk Provide good Survey the customer
service to satisfaction with service
customers
Teacher Impart knowledge Student achievement test
to students scores
Weather Accurately predict Compare predictions to
forecaster the weather actual weather
Actual criteria
• Only rough estimates/imperfect indicators of
theoretical criteria it intends to measure
– Might assess a piece of the intended
theoretical criterion but likely some of
theoretical criterion left out
– Can be biased and can assess something
other than the theoretical criterion
• This imperfection may result from
– Contamination: Actual measures
something other than the theoretical
– Deficiency: Actual fails to capture the
theoretical
– Relevance: Actual assesses the
theoretical
• Refers to the part of the actual criterion that reflects
something other than the theoretical criterion.
– Biases and unreliability
• Biases
– Common when judgments/opinions are used as the
actual criterion
• E.g., using the art experts’ judgments as the
actual criterion to appraise an artwork reveal
biases of the judges as it does about the work
Criterion • Because there is no objective standard for
quality of art, experts will likely disagree with
Contamination one another
• Unreliability
– Arises from random errors in measurement
• Reflected in inconsistency in
measurement over time
– If assessed performance repeatedly over
time, measurement will vary even if the
theoretical criterion is the same
Criterion • Actual performance criterion with less
than perfect reliabilities
Contamination
Criterion Deficiency
• Actual criterion not adequately cover entire
theoretical criterion.
– Refers to content validity
– E.g., using mathematical achievement test
scores as actual performance criterion for
elementary school teachers
• Deficient because elementary school
teachers teach more than just mathematics
• Less deficient; student scores on a
comprehensive test battery including math,
reading, science, and writing
• If university evaluates professors only by
student evaluations; criterion deficiency,
because ignores significiant factors such as
research output, service to university etc.
Criterion relevance
• The extent to which the actual criterion
assesses the theoretical criterion/overlap
between them.
– Construct validity
– Closer the correspondence, the
greater the relevance of the actual
criterion.
– When relevance based on inferences
and interpretations about meaning of
measurements of performance
• If theoretical criteria is abstract,
difficult to determine the
relevance of a criterion; e.g.,
quality of artwork
Criteria: Levels of specificity
• Criteria can be developed for individual tasks or for
entire jobs.
– Most jobs complex, involve many different
functions/tasks
– Desired level of specificity depends on the
purpose of the assessment information.
• For developing employee’s skills, better to
focus on individual task level so that
feedback be specific (task level).
– E.g., making arrests for PO, selling
products for a salesperson
– E.g., employee be told that s/he types too
slowly/makes too many errors
– For administrative purposes, overall job
performance might be of more concern (general
evaluation of performance over period).
Criterion Complexity
Criterion complexity
• For most, more than one performance dimension,
• Nature of job and purpose of assessment determine
criterion measures.
• Multiple criterion measures
Most jobs even for single tasks have more
than one criterion
• Quality (how well a person may perform a job) vs.
quantity (how much or how quickly a person does a job)
• Other potential criteria such as maintaining professional
appearance, motivating others, following directions
Criterion complexity
• Multiple ways to combine performance
dimensions
– Composite criterion approach-individual
criteria are combined into a single score.
• For comparing individual employee
performance
• If a score for each dimension, a
composite would be the average score
of them
– Multidimensional approach-separate score
for each criterion.
• For giving employees feedback about
the dimensions of performance;
feedback for each dimension
• If 4 dimensions, there will be 4 scores
for each employee
Dynamic Criterion
Mixed results:
Does performance The standard stays sometimes
vary over time? the same performance is stable
and sometimes not.
• Dynamic criteria; variability of
performance over time
Characteristics – Performance not the standard that
changes
– Deadrick & Madigan (1990);
of Criteria performance of sewing machine
operators was stable over short
periods of time (weeks) but not
(Cont.) very consistent over long periods of
time
– Vinchur et al. (1991) found the
performance of employees were
stable over a 5-year time span
– Deadrick et al. (1997); performance
improve through time
• Factors for improvement for
new not similar to those for
later
• Depend on nature of the job and job
setting
– Person’s performance over time can
be variable
– People differ in their patterns of
performance variability
Characteristics of
Criteria (Cont.)
• Contextual performance
• Not specifically required but extra,
voluntary tasks benefit coworkers
& organization (e.g., helping/
assisting with coworkers and
volunteering to carry out extra
tasks).
• Although not specifically required,
contextual performance noticed and
appreciated by managers, and their
ratings of subordinate performance
will be affected by it.
Performance
• Objective
Appraisal Methods
– Counts of behaviors
• Days absent
– Results of behaviors
• Units sold
• Subjective
– Human judgments of
performance
– Ratings by supervisors
• Usually do not agree for the same
employee; may reflect different
aspects of performance
Methods for Assessing Performance:
Objective Measures of Job Performance
Advantages Disadvantages
• Easy to interpret • Not always appropriate when there is no
• Easy to compare individuals in countable output.
the same job. • No clear cut-off (what number is
• Not biased by judgment considered satisfactory performance).
• Can be tied directly to • Records can be contaminated (e.g.,
organizational objectives, such employees might fail to report accidents
as making a product or providing and injuries).
a service. • Focus on specific behaviors, which may
• No need for special performance be only part of the criterion (focus on
appraisal systems; data are often quantity, omit quality)
are collected and stored, • What is measured may not be under the
frequently in computers. employee’s control (e.g., the quality of
the machine)
Subjective Methods
• People’s judgments about performance
– Ratings by people knowledgeable about
the person’s job performance
– Usually, supervisors rate subordinates
annually.
– Most frequently used means of assessing
job performance.
• Many rating forms are in use
– Graphic rating form
– Behavior-focused rating form
• Behaviorally Anchored Rating Scale
• Mixed Standard Scale
• Behavior Observation Scale
•
Subjective Measures
of Job Performonce:
Graphic Rating
Form
– Most popular measure
– Asks for ratings on several
dimensions of performance,
including
• Work quality, work Dimension/Rating Poor Fair Adequate Good Outstanding
quantity, and relevant Attendance
personal traits such as Work quantity
appearance, motivation, Following
and dependability. instructions
Motivating others
– For each dimension, the
supervisor checks off his rating Professional
appearance
(e.g., “poor” to “outstanding”) Communication
of the employee for each
Work quality
dimension
Graphic rating form
• Graphically display
performance scores
running from high to low
– Needs to have:
• well-defined dimensions,
• understandable and
appropriately placed
anchors,
• unambiguous ratings
– Ratees prefer more rating
scale units (> 3)
Subjective Methods
• Behavior-focused rating form concentrate
on specific instances of behavior.
– Behaviors representing good and poor
levels of performance for various
dimensions.
• The rater indicates which
behaviors are typical of the person
being rated.
– Several different such forms; provides
description of behavior but differ in
presenting description or responses
• Behaviorally Anchored Rating
Scale
• Mixed Standard Scale
• Behavior Observation Scale
Behaviorally Anchored
Rating Scale (BARS)
– Rating scale in which response choices are
defined in behavioral terms along a continuum
of performance effectiveness
• Whereas graphic rating form (GRF) for a
rating how well employee performs along
the dimension in question without any
behavior description
– Contains several scales similar to GRF,
• Important dimensions of job performance.
• For each scale, the rater picks the behavior
that comes closest to describing the
performance of the person being rated.
Mixed
Standard
Scale (MSS)
• Provides the rater with a list of
behaviors that vary in their
effectiveness on performance
dimensions
– For each behavior, the rater
indicates if the ratee is
• Better than the
statement,
• The statement fits the
ratee,
• Or the ratee is worse
than the statement.
Behavior- • Mixed Standard Scale (MSS)
Focused – Several performance dimensions although
nature of dimension is not told
Rating • Several behaviors for each dimension.
Form – The statements for dimensions are presented
in random order.
• No differences if the dimensions are
identified/statements are mixed up
(Dickinson & Glebocki, 1990)
Behavior Observation Scale (BOS)
Rater indicates the percent
Contains items based on
of time the employee
critical incidents like MSS.
engaged in each behavior.
• Critical Incident; an • How often employees
event reflecting either engage in performance
effective or ineffective relevant behavior.
behavior by the employee • Different from MSS, in
• E.g., a poor incident for which the employee
a teacher; “slapping a behavior is compared with
child who made a the one represented by the
disrespectful comment” item
Behavior Observation Scale
1 2 3
List of behaviors that Rater indicates how Weaknesses:
vary in performance often employee does Behaviors vary in
from poor to each behavior using criticality, so
excellent percentages frequency is not
equally important
Behavioral Observation Scale
• Example – Student’s class engagement
Never Seldom Sometimes Generally Always
Critical Incidents
1. Attendance 1 2 3 4 5
2. Takes Notes 1 2 3 4 5
3. Asks Questions 1 2 3 4 5
4. Is asleep 1 2 3 4 5
Steps for Developing Behavior-
Focused Forms
Requires developing form for specific job or similar group of jobs
1. Job analysis identifies dimensions of performance
2. Develop descriptions of effective and ineffective behavior on these
dimensions using critical incidents obtained from SMEs.
3. Have people knowledgeable about the job sort the behaviors into
dimensions
– To verify that the behaviors reflect the intended dimensions.
4. Have people knowledgeable about the job rate the behaviors on a
continuum of effectiveness
– Select a subset of behaviors that represent the desired levels
• E.g., for the BARS, placing the behaviors along the scale; For the
MSS, grouping the behaviors into good, satisfactory, or poor
Subjective Methods; Cognitive
Processes Underlying Ratings
• To use ratings in performance appraisal, need to
know
– how people make these ratings and
– what factors influence them
• Some models focus on how
– people utilize info to make judgments
– people’s views of job performance influence
their employee evaluations
• Models influencing performance ratings suggest
rating process includes several steps
– Observing the employee’s performance.
– Storing information about performance in
memory.
– Retrieving information about performance
from memory.
– Translating the retrieved information into
ratings.
Models in Cognitive
Processes
• Models vary in descriptions of how information is
processed at each step
• Schemata to interpret and organize experience;
categories or frames of reference
• Stereotypes; a belief about characteristics of
the members of a group
– Can be favorable or unfavorable; e.g.,
private sector managers are hard working
• Prototypes; a model of some characteristic or
type of person
– Used as a standard to assign people to
category of good manager,
– E.g., Bill Gate (the founder and head of
Microsoft) can be considered as a
prototype of a good corporate manager
» If someone looks like him, may be
subsumed under the category of good
manager
Models in Cognitive
Processes Underlying Ratings
• Schemata can influence all four steps in the
evaluation process.
– Might affect the behaviors observed;
– How the behaviors are organized and stored
in memory, how they are retrieved; and
– How the behaviors are used to decide on
ratings.
– The evaluations that result are not necessarily
inaccurate although schemata are used
• Use of schemata can simplify experience
so that it can be more easily interpreted
Models in Cognitive Processes
Underlying Ratings
Can the use of cognitive models help raters Study by Jelley & Goffin (2001); college
do a more accurate job of evaluating job students were asked to rate the
performance? performance of a videotaped college
instructor using a BOS
Results; although inconsistent, they found
some accuracy increases after priming the
rater’s memory
• By making raters to do some preliminary
global ratings designed to stimulate recall
of performance observed
Content of subordinate
effectiveness
• Appraisal techniques might be improved by designing them to use the
dimensions in the supervisors’ schemata
– Appraisal form dimensions match the dimensions in supervisor schemata
about performance; easier for the supervisor do their ratings
• E.g., Borman’s (1987) study of Army officers.
– The army officers described the effective and ineffective soldiers
– Generated 189 descriptive items subsumed under 6 dimensions
– Include working hard, being responsible, being organized, knowing the
technical parts of the job, being in control of subordinates, and displaying
concern for subordinates.
– Result; the experienced supervisors might have a schemata that accurately
represents effective performance
– There was agreement between officers about what is a good performance
• Werner (1994); supervisors rated the secretaries in a series of incidents in terms
of importance
Rater Error: Imperfect Human
Judgement
• Rater error & bias; Patterns
– Halo errors
– Distributional errors
– Research shows that behavior-focused ratings alone do not eliminate
errors
• Rater training
– Rater error training: Instructs raters in how to avoid errors
• Reduces halo and leniency error
• Less accuracy in some studies
– Frame of reference training: Give raters examples of performance and
correct ratings
• Initial research promising in reducing errors (Day & Sulsky, 1995)
Rater Error & Bias
• Halo errors
– When rated individual same on all
dimensions despite differences in
performance across dimensions
• E.g., One employee gets all l’s,
another all 5’s
– One problem with halo error is
distinguishing it from true halo
• In which an employee actually
performs at the same level on all
dimensions.
– Might occur if raters form general
impression of individual’s performance
and use it to make ratings on dimensions
• Use whatever is salient
• Study: When halo error was controlled, an
underlying performance dimension that creates this
pattern.
– Organizational citizenship behavior (behavior
that goes beyond what is expected).
– Result; OCB influences all aspects of
performance including technical performance.
Halo • Rater may
– Believe that one particular dimension is key to
other dimensions.
errors – Use it habitually because believes that there is
only one performance factor; people either bad
or good performers.
– Believe that the most essential dimension is
not in the form; adaptability or OCB.
• Downside; impossible to give feedback to
employee by identifying the strengths and
weaknesses.
Rater Error & Bias
• Distributional errors
– When rater tends to rate everyone the same;
pattern seen in rating forms of different
individuals.
– Leniency errors—the rater rates everyone at the
top end of the rating scale.
– Severity errors—the rater rates everyone at the
bottom of the rating scale.
– Central tendency errors—the rater rates
everyone in the middle of the scale.
• Again, what appears to be error may reflect reality—
all employees might have performed the same,
leading to similar ratings
• [Link]
Examples of Rater Errors
• Two approaches
– Research shows that one approach alone do
not eliminate errors
How to 1. Error resistant forms to assess
•
performance
Behavior-focused rating scales an
Control attempt and may be effective; more
concrete with less idiosyncratic
judgments
Rater • Rater be more accurate if s/he
focused on specific behaviors rather
than traits.
Error & – How frequently one is absent vs.
dependability
Bias • Evidence: Behavior-focused scales
have little advantage in rating error
resistance than graphic rating scale
2. Rater training to reduce errors
– Rater error training (RET)
– Observation training
– Frame of reference training
• Rater error training (RET) is the
most popular.
• Familiarizes raters with rating
errors and teaches to avoid those
patterns.
• It seems to reduce errors but
reduces accuracy as well.
• Avoiding certain patterns may
Rater Training lead raters to create less accurate
patterns
– I.e., rater might reduce
number of halo and leniency
patterns in their ratings, but
result (ratings) less accurate
in reflecting the true level of
performance
Rater error training
(RET)
• Why avoiding certain patterns may lead lower
accuracy?
• Possibility; performance of individuals similar
across different performance dimensions; true halo
– Or all individuals in supervisor’s department
perform their jobs equally well
• Training raters increase the focus on avoiding the
same rating across dimensions or avoiding certain
patterns
• Other explanation; Nathan & Tippins (1990);
raters with less halo in their ratings might have
given too much weight to inconsequential negative
events
– E.g., rating the performance of an employee
as low due to being absent for a week
previous year
Rater Training
• Observation training
• Involves teaching raters how to observe
performance behavior and how to make
judgments using those observations.
• Seems to increase accuracy but not to
reduce rating errors
• Frame of reference training
• Has demonstrated some promising
results.
• It attempts to provide a common
understanding of the rating task
• Raters are given specific examples that
represent various levels of performance
for each dimension
More Performance Appraisal
• Liking of subordinates
Other factors that • Expectations
influence ratings: • Rater mood
characteristics of the • Perceptions of motivation
supervisor and the • Cultural factors/Race
rating situation
360 degree feedback
• Monitoring of objective productivity
• Performance management systems
Technology
• Web-based
• Automates process and notifies raters
More Performance Appraisal
– Supervisors give better ratings to subordinates they like; favoritism
• Though also tend to like people who perform well
– Expectancies;
• Who have done well in the past tend to get good ratings
– People tend to forget the instances that do not fit their view of the
person they are evaluating
• Can produce biased ratings when performance changes through time
– Rater mood
• Laboratory study: the mood of the subjects were manipulated; induced to
feel either good or bad
– Read a description of a professor's performance and rated it
– Results; showed that depressed subjects gave lower ratings than
elated subjects,
» Also, more accurate ratings with less halo
More Performance Appraisal
• But such views can be subject to
Perceptions of cultural factors.
motivation; • E.g., American managers
Manager’s views considered intrinsic motivation
(i.e., wanting to do a good job for
of subordinate its own sake) to be more important,
motivation can be and Asian managers consider both
a factor in ratings intrinsic and extrinsic (i.e., working
hard for rewards) to be important.
• Black and White raters give similar
Race; Black ratings to whites but give lower
employees ratings to blacks
receive worse • However, difference greater for
appraisals than white raters than black raters
White employees • Unclear why
on average, • One possibility; white raters more
regardless of the biased against black employees
race of the rater. • OR, black raters biased in favor of
blacks and overrate them
Ratings by all relevant parties
If for managers; 3600 feedback
Ratings by peers, subordinates, supervisor
360 (modest agreement), and own ratings
Degree Enables manager to see how different groups
view him/her
Feedback
Favoritism by the immediate supervisor is
diminished; increases trust in system
Technology used to deal with volume of data in
large companies
360-degree
feedback
• Have been shown to have positive effects
for some individuals but not all
• Contrary to intended purpose, good
performers seem to benefit the most
from the 360-degree feedback.
• Atwater & Brett (2005); those
individuals who receive low ratings
from others and rated themselves low
as well had the worst reactions to
feedback
– Conclusion; if one knows that one’s
performance is poor, having this
belief corroborated by others is not
helpful
[Link]
Use of Technology
Monitoring of objective Performance management
productivity systems
Phone bank monitoring Keep track of employee data using computers
• Compute analytics • 360 data; feasible economically through
• Number of calls automation
• Average time of calls • Track knowledge and skill levels and training
Performance appraisals can be legally
challenged
• Subjective methods are especially vulnerable to
legal attack; allow room for supervisors to express
prejudices against certain groups
• Overall, organizations lost 41% of cases whereas
those that used multiple raters only lost 11 %
• Should create perceptions of fairness
• When perceived fair, it even reduces employee
Legally intentions of quitting the job
Defensible
Performance Practices that minimize legal challenges
Appraisal • Job analysis to define dimensions of performance
• Develop rating form to assess dimensions from prior
point
• Train raters in how to assess performance
• Management review ratings and allow employee
appeal
• Document performance and maintain detailed
records-easier to take action when necessary if take
for a long time
• Provide assistance and counseling
• can lead to better attitudes
Job analysis to
define dimensions
4 of performance;
• A job analysis will
Use multiple
raters
ensure that the
Components dimensions are job
relevant
that
Minimize Train raters in
how to assess
Management
review ratings
Legal performance
• The raters should
and allow
employee appeal
Challenges learn how to use the
rating form
• In order to minimize
potential bias by
raters