Module 2
Human-Centred AI Framework
Introduction: Rising above the Levels of Automation
The success of AI algorithms has brought many new possibilities for improving
design of widely used technologies. Researchers, developers, managers, and
policy-makers who embrace human-centred approaches will accelerate the
design and adoption of advanced technologies. The new synthesis that combines
AI algorithms with human centred design will take time.
This module shows how human-centred AI ideas open up new
possibilities for design of systems that offer high levels of human control
and high levels of computer automation.
It will take time to refine and disseminate this idea and respond to
resistance.
I believe that these designs can increase human performance by providing
just the right automated features, which can be accomplished reliably,
while giving users control of the features that are important to them.
The human-centred artificial intelligence (HCAI) framework clarifies how
to:
(1) design for high levels of human control and high levels of computer
automation so as to increase human performance,
(2) understand the situations in which full human control or full computer
control are necessary, and
(3) avoid the dangers of excessive human control or excessive computer control.
Achieving these goals will also support human self-efficacy, creativity,
responsibility, and social connections.
The new guideline is to seek both high levels of human control and high levels
of automation, which is more likely to produce computer applications that are
reliable, safe, and trustworthy.
This chapter focuses on the reliable, safe, and trustworthy goals, which may
help to achieve other important goals such as privacy, cybersecurity, economic
development, and environmental preservation.
I share the belief that computer autonomy is compelling for many applications.
Who wouldn’t want an autonomous device that would do a repetitive or difficult
task reliably and safely? But what if it does that task correctly only 98% of the
time? Or if it catches fire once in a hundred times?
Autonomy may be fine for some tasks, but dangerous for others. In consumer
applications, like making recommendations, unusual films, or offbeat
restaurants might lead to a welcome discovery, so the dangers of errors are low.
For other tasks, like sensor-based pattern recognition for flushing a toilet in
a bathroom, incorrect activations or failure to activate are minor annoyances.
However, if the task is driving you to work, would you accept a car that once a
month took an hour to start or a car that occasionally parked itself a mile from
your workplace? For lightweight tasks like toilet flushing, failures are just an
annoyance, but for transportation, health, financial, and military applications
higher levels of reliability, safety, and trustworthiness are needed to gain the
high levels of user acceptance that lead to commercial success.
But may be there are designs of recommender systems, toilet flushing, or car
driving in which users can take more control of what the machine does,
prevent failures, and take over when user intent is not fulfilled. In a study of a
news recommender system, users were given three sliders to allow them to
indicate their interest in reading politics, sports, or entertainment articles. As
users move the sliders, the list of recommended articles changes immediately so
users can explore alternatives, leading to users clicking on recommendations
more often and expressing “a strong desire for having more control.” A second
study of two different control panels for a news recommender showed similar
strong results favouring user control.
Despite many good designs, there is room for improvement in existing systems.
I had an annoying experience at a swimming pool I go to regularly. While
I was getting dressed, the changing room toilet flushed unnecessarily
eight times, but I could see no way of stopping the automated device that
used pattern recognition technologies from AI’s early days.
By contrast, I often had trouble activating the sink and the hand dryer,
which became a nuisance. These pattern recognition devices were neither
reliable or trustworthy. A remedy would be a standard icon, approved by
all manufacturers, placed just at the spot where the pattern recognition
sensor would be sure to spot it or a button to turn the sink or dryer on.
The black mats in front of automatic doors are usually not the sensor, but
they are a clear indicator that stepping on to the mat will put you in the
place where the camera sensor will pick up your image. Override buttons for
users with handicaps enable them to activate the door on demand and keep the
door open for longer, so they can get through safely.
Greater user control, even in these relatively low-tech applications, make
for a more reliable, safe, and trustworthy system. In summary, automated
systems can be wonderful, but strategies to ensure activation when
needed and prevent inadvertent activations contribute to reliable, safe,
and trustworthy operation that increases user acceptance.
Good design also includes ways for maintenance staff to gather
performance data to know whether the operation is correct and adjust
controls to make sure the device functions correctly.
These issues become vital in life-critical applications such as airbag
deployment.
Airbags explosively deploy in two-tenths of a second to save about 2500
lives a year in the United States. However, in the early years more than 100
infants and elders were killed by inadvertent deployments, until design changes
reduced the frequency of these tragedies.
The important lesson is that collecting data on unsuccessful and
inadvertent deployments provides the information necessary for
refinements that make systems reliable, safe, and trustworthy.
When autonomous systems’ failures and near misses are tracked to
support improved designs and when public websites enable reporting of
incidents, quality and acceptance increase.
For consequential applications such as car driving, safety comes first. I
might be willing to buy a car that would not let me drive if my breath
alcohol was above legal levels.
I would be even more eager to require that all cars had such a system so
as to prevent others from endangering me.
On the other hand, if that system made an incorrect measurement when I
was excitedly getting in my car to drive my injured child to the hospital, I
would be very angry, because the car was not trustworthy.
The first lesson about autonomous systems, which we will return to later,
is that user controls to activate, operate, and override can make for more
reliable, safe, and trustworthy systems.
The second lesson is that performance histories, called audit trails or
product logs, provide valuable data that lead to continuous design
refinements.
Maybe the most important lesson is that humility in an important quality
for designers, who will be more successful if they think carefully about
the possibilities of failures.
Tom Sheridan and William Verplank ten levels of automation from
human control to computer autonomy
Sheridan and Verplank’s ten levels of autonomy have been widely influential,
but critics noticed that it was incomplete, missing important aspects of what
users do. Over the years, there have been many refinements such as recognizing
that there were at least four stages of automation:
(1) information acquisition,
(2) analysis of information,
(3) decision or choice of action, and
(4) execution of action.
The four stages were an important refinement that helped keep the levels of
autonomy alive, even as critics continued to be troubled by the simple one-
dimensional framework, which assumed that more automation meant less
human control. Shifting to the two-dimensional framework for HCAI, could
liberate design thinking so as to produce computer applications that increase
automation, while amplifying, augmenting, empowering, and enhancing
people to apply systems innovatively and creatively refine them.
The US Society of Automotive Engineers adopted the unnecessary
trade-off in its six levels of autonomy for self-driving cars.
A better approach would be to clarify what features could be automated, like
collision avoidance or parking assist, and what activities required human
control, like dealing with snow-covered roads, sensor failures, or verbal
instructions from police officers.
Defining Reliable, Safe, and Trustworthy
Systems
While machine autonomy remains a popular goal, the goal of human autonomy
should remain equally strong in designers’ minds. Machine and human
autonomy are both valuable in certain contexts, but a combined strategy uses
automation when it is reliable and human control when it is necessary. To guide
design improvements it will be helpful to focus on the attributes that make
HCAI systems reliable, safe, and trustworthy.
These terms are complex, but I choose to define them with four levels of
recommendations:
(1) reliable systems based on sound software engineering practices,
(2) safety culture through business management strategies,
(3) trustworthy certification by independent oversight, and
(4) regulation by government bodies.
These four communities are eager to accelerate development of reliable, safe,
and trustworthy systems. Software engineers, business managers, independent
oversight committees, and government regulators all care about the three goals,
but I suggest practices that each community might be most effective in
supporting.
Reliable systems produce expected responses when needed. Reliability comes
from appropriate technical practices for software engineering teams. When
failures occur, investigators can review detailed audit trails much like the logs
of flight data recorders, which have been so effective in civilaviation. The
technical practices that support human responsibility, fairness, and
explainability include:
Audit trails and analysis tools
• Software engineering workflows
• Verification and validation testing
• Bias testing to enhance fairness
• Explainable user interfaces
Cultures of safety are created by managers who focus on these strategies
Leadership commitment to safety
• Hiring and training oriented to safety
• Extensive reporting of failures and near misses
• Internal review boards for problems and future plans
• Alignment with industry standard practices
These strategies guide continuous refinement of training, operational practices,
and root-cause failure analyses.
Trustworthy systems are discussed more frequently than ever, but there is a
long history of discussions to define trust. An influential contribution was
political scientist Francis Fukuyama’s book Trust: The Social Virtues and the
Creation of Prosperity, which focused on social trust “within a community of
regular, honest, and cooperative behaviour, based on commonly shared norms,
on the part of the members of that community.” He focused on human
communities, but there are lessons for design of trust in technology.
The list of independent oversight organizations includes:
• Accounting firms conducting external audits
• Insurance companies compensating for failures
• Non-governmental and civil society organizations
• Professional organizations and research institutes
Government intervention and regulation also play an important role as
protectors of the public interest, especially when large corporations elevate their
agendas above the needs of residents and citizens.
Governments can encourage innovation as much as they can limit it, so learning
from successes and failures will help policy-makers to make better decisions.
I chose to focus on reliable, safe, and trustworthy to simplify the discussion, but
the rich literature on these topics advocates other attributes of systems, their
performance, and user perceptions.
Successful technologies enable humans to work in interdisciplinary teams so as
to coordinate and collaborate with managers, peers, and subordinates. Since
humans are responsible for the actions of the technologies they use, they are
more likely to adopt technologies that show current status on a control panel,
provide users with a mental model to predict future actions, and allow users to
stop actions that they can’t understand. Well-designed user interfaces provide
vital support for human activities in ways that reduce workload, raise
performance, and increase safety.
Two-Dimensional HCAI Framework
The HCAI framework steers designers and researchers to ask fresh questions
and rethink the nature of autonomy. As designers get beyond thinking of
computers as our teammates, partners, or collaborators, they are more likely to
develop technologies that dramatically increase human performance. Novel
designs will take stronger advantage of unique computer features such as
sophisticated algorithms, voluminous databases, advanced sensors, information
abundant displays, and powerful effectors, such as claws, drills, or welding
machines.
Clarifying human responsibility also guides designers to support the human
capacity to invent creative solutions in novel contexts with incomplete
knowledge. The HCAI framework clarifies how design excellence promotes
human self-efficacy, creativity, and responsibility, which is what managers and
users seek. The goals of reliable, safe, and trustworthy are useful for lighter
weight recommender systems for consumers and consequential applications
for professionals, but they are most relevant for life-critical systems:
Recommender systems: These are widely used in consumer services, social
media platforms, and search engines, and have brought strong benefits to
consumers. Consequences of the frequent mistakes by recommender systems
are usually less serious, possibly even giving consumers interesting suggestions
of movies, books, or restaurants. However, malicious actors can manipulate
these systems to influence buying habits, change election outcomes, spread
hateful messages, and reshape attitudes about climate, vaccinations, gun control
etc.
Consequential applications: Getting correct outcomes is more important in
medical, legal, environmental, or financial systems that can bring substantial
benefits but also have risks of serious damage. A well-documented case is the
flawed Google Flu Trends, which was designed to predict flu outbreaks based
on user searches on terms such as “flu,” “tissues,” or “cold remedies.” It was
intended to enable public health officials to assign resources more effectively .
Life-critical systems: Moving up to the challenges of life-critical applications,
we find the physical devices like self-driving cars, pacemakers, and implantable
defibrillators, as well as complex systems such as military, medical, industrial,
and transportation applications. These applications sometimes require rapid
actions and may have irreversible consequences. Designing life critical systems
is a serious challenge, but a necessary one as these systems can avoid dangers
and save lives.
Designers of recommender systems, consequential applications, and life
critical systems were guided by the one-dimensional levels of autonomy.
I shared the belief that more automation was better and that to increase
automation, designers had to reduce human control. In short, designers had to
choose a point on the one-dimensional line from human control to computer
automation (Figure 8.1).
The mistaken message was that more automation necessarily meant less user
control. Over the years, this idea began to trouble me so I eventually shifted to
the belief that it was possible to ensure human control while increasing
computer automation. Even I wrestled with this puzzling notion, till I began to
see examples of designs that had high levels of human control for some features
and high levels of automation for other features. The decoupling of these
concepts leads to a two-dimensional HCAI framework, which suggests that
achieving high levels of human control and high levels of computer automation
is possible (yellow triangle in Figure 8.2). The two axes go from low to high
human control and from low to high computer automation. This simple
expansion to two dimensions is already helping designers imagine fresh
possibilities.
The lower right quadrant (Figure 8.3), with high computer automation andlow
human control, is the home of computer autonomy requiring rapid action, for
example, airbag deployment, anti-lock brakes, pacemakers, implantable
defibrillators, or defensive weapons systems. In these applications, there is no
time for human intervention or control. Because the price of failure is so high,
these applications require extremely careful design, extensive testing, and
monitoring during usage to refine designs for different use contexts. As systems
mature, users appreciate the effective and proven designs, paving the way for
higher levels of automation and less human supervision.
The upper left quadrant, with high human control and low automation, is the
home of human autonomy where human mastery is desired to enable
competence building, free exploration, and creativity. Examples include bicycle
riding, piano playing, baking, or playing with children where errors are part of
the experience. During these activities, humans generally want to derive
pleasure from seeking mastery, improving their skills, and feeling fully
engaged. A safe bike ride or a flawless violin performance are events to
celebrate. They may elect to use computer-based systems for training, review,
or guidance, but many humans desire independent action to achieve mastery
that builds self-efficacy.
In these actions, the goal is in the doing and the personal satisfaction that it
provides, plus the potential for creative exploration. The lower left quadrant is
the home of simple devices such as clocks, music boxes, or mousetraps, as well
as deadly devices such as land mines.
Two other implementation aspects greatly influence reliability, safety, and
trustworthiness: the accuracy of the sensors and the fairness of the data. When
sensors are unstable or data sources incomplete, error-laden, or biased, then
human controls become more important. On the other hand, stable sensors
and complete, accurate, and unbiased data favor higher levels of automation.
The take-away message for designers is that, for certain tasks, there is value
in full computer control or full human mastery. However, the challenge is to
develop effective and proven designs, supported by reliable practices, safety
cultures, and trustworthy oversight.
In addition to the extreme cases of full computer and human autonomy, there
are two other extreme cases that signal danger—excessive automation and
excessive human control. On the far right of Figure 8.4 is the region of
excessive automation, where there are dangers from designs such as the Boeing
737 MAX’s MCAS system, which led to two crashes with 346 deaths in
October 2018 and March 2019. There are many aspects to this failure, but some
basic design principles were violated. The automated control system took the
readings from only one of the two angle of attack sensors which indicate
whether the plane is ascending or descending.
When this sensor failed, the control system forced the plane’s nose downwards,
but the pilots did not know why, so they tried to pull the nose up more than
twenty times in the minutes before the crash, which killed everyone on board.
The aircraft designers made the terrible mistake in believing that their
autonomous system for controlling the plane could not fail. Therefore, its
existence was not described in the user manual and the pilots were not trained in
how to switch to manual override. The unnecessary tragedies were entirely
avoidable. The IBM AI Guidelines wisely warns that “imperceptible AI is not
ethical AI.”
Figure 8.5 shows the relative positions of 1940 and 1980 cars, 2020 self driving
cars, and the proposed goal of reliable, safe, and trustworthy cars in 2040 by
way of mature automation and supervisory control.
The four quadrants may be helpful in suggesting differing designs for a product
or service, as in this example. Patient-controlled analgesia (PCA) devices allow
post-surgical, severe cancer, or hospice patients to select the amount and
frequency of pain control medication. There are dangers and problems with
young and old patients, but with good design and management, PCA devices
deliver safe and effective pain control. A simple morphine drip bag design
for the lower left quadrant (low computer automation and low human control)
delivers a fixed amount of pain control medication (Figure 8.6). A more
automated design for the lower right quadrant (increased computer automation,
but little human control) provides machine-selected doses that vary by time of
day, patient activity, and data from body sign sensors, although these do not
assess perceived pain.
A human-centered design for the upper left quadrant (higher human control,
with low computer automation) allows patients to squeeze a trigger to control
the dosing, frequency, and total amount of pain control medication. However,
the dangers of overdosing are controlled by an interlock that prevents frequent
doses, typically with lockout periods of 6–10 minutes, and total dose limits over
one- to four-hour periods. Finally, a reliable, safe, and trustworthy design for
the upper right quadrant allows users to squeeze a trigger to get more pain
medicine but uses machine learning to choose appropriate doses based on
patient and disease variables, while preventing overdosing. Patients are able to
get information on why limiting pain medication is important with explanations
of how they operate the PCA device (Figure 8.7). The design includes a hospital
control center (Figure 8.8) to monitor usage of hundreds of PCA devices so as
to ensure safe practices, deal with power or other failures, review audit trails,
and collect data to improve the next generation of PCA devices.
Design Guidelines and Examples
Google’s complementary website gives guidelines for responsible AI that are
well-aligned with my principles.
Use a human-centered design approach
• Identify multiple metrics to assess training and monitoring
• When possible, directly examine your raw data
• Understand the limitations of your data set and model
• Test, Test, Test
• Continue to monitor and update the system after deployment.
The design decisions to craft user interfaces based on the Eight Golden Rules
typically involve trade-offs,13 so careful study, creative design, and rigorous
testing will help designers to produce high-quality user interfaces.
Example 9.1 Simple thermostats allow residents to take better control of the
temperature in their homes. They can see the room temperature and the current
thermostat setting, clarifying what they can do to raise or lower the setting.
Then they can hear the heating system turn on/off or see a light to indicate
that their action has produced a response. The basic idea is to give residents an
awareness of the current state, allow them to reset the control, and then give
informative feedback (the third Golden Rule) that the computer is acting on
their intent. There may be further feedback as the residents see the thermometer
rise in response to their action, and maybe an indication when their desired
goal has been achieved. Thermostats offer still further benefits—they continue
to keep the room temperature at the new setting automatically. In summary,
while some thermostats may lack the features necessary that give clear
feedback, well-designed thermostats give users an understanding of how they
have controlled the automation to get the temperature they desire in their homes
(Figure 9.1). Newer programmable thermostats, such as Google’s Nest, apply
machine learning to allow residents to accommodate their schedules better and
save energy. However, human behavior can be variable, undermining the utility
of machine learning, as when residents change their schedules, adopt new
hobbies such as baking, or have visitors with different needs. Getting the
balance between human and machine control remains a challenge.
Example 9.2 Home appliances, such as dishwashers, clothes washers/dryers,
and ovens allow users to choose settings that specify what they want, and then
turn control over to sensors and timers to govern the process. When well
designed, these appliances offer users comprehensible and predictable user
interfaces. They allow users to express their intent, with controls that let them
monitor the current status and they can stop dishwashers to put in another plate
or change from baking to broiling to brown their chicken (the seventh Golden
Rule). These automations give users better control of these active appliances to
ensure that they get what they want (Figure 9.2).
Example 9.3 Well-designed user interfaces in elevators enable substantial
automation while providing appropriate human control. Elevator users approach
a simple two-button control panel and press up or down. The button lights to
indicate that the users’ intent has been recognized. A display panel indicates
the elevator’s current floor so users can see progress, which lets them know
how long they will have to wait. The elevator doors open, a tone sounds, and
the users can step in to press a button to indicate their choice of floors. The
button lights up to indicate their intent is recognized and the door closes. The
floor display shows progress towards the goal. On arrival, a tone sounds and
the door opens. The automated design which replaced the human operator
ensures that doors will only open while on a floor, while triply redundant safety
systems prevent elevators from falling, even under loss of power or broken
cables. Machine learning algorithms coordinate multiple elevators,
automatically adjusting their placement based on time of day, usage patterns,
and changing passenger loads. Override controls allow firefighters or moving
crews to achieve their goals. The carefully designed experience provides users
with a high level of control over relevant features supported by a high level of
automation. There are many refinements, such as detectors to prevent doors
from closing on passengers (the fifth Golden Rule), so that the overall design is
reliable, safe, and trustworthy.
Example 9.4 Digital cameras in most cell phones display an image of what the
users would get if they clicked on the large button (Figure 9.3). The image is
updated smoothly as users adjust their composition or zoom in. At the same
time, the camera makes automatic adjustments to the aperture and focus, while
compensating for shaking hands, a wide range of lighting conditions (high
dynamic range), and many other factors. Flash can be set to be on or off, or
automatically set by the camera. Users can choose portrait modes, panorama, or
video, including slow motion and time lapse. Users also can set various filters
and once the image is taken they can make further adjustments such as
brightness, contrast, saturation, vibrance, shadows, cropping, and red-eye
elimination. These designs give users a high degree of control while also
providing a high level of automation. Of course, there are mistakes, such as
when the automatic focus feature puts the focus on a nearby bush, rather than
the person standing just behind it. However, knowledgeable users can touch the
desired focus point to override this mistake.
Example 9.5 Auto-completion is a form of recommender system that suggests
commonly used ways to complete phrases or search terms, as in the Google
search system. In Figure 9.4, the user begins a term and gets popular
suggestions based on searches by other users. This auto-completion method
does more than speed performance; it reminds users of possible search topics
that they might find useful. Users can ignore the recommendation or choose
one of the suggestions.
Example 9.6 Spelling and grammar checkers as well as natural language
translation systems offer useful suggestions in subtle ways, such as a red wiggly
line under the error, as in Figure 9.5, or the list of possible translations in
Figure 9.6. These are examples of good interaction design that offers help to
users while letting them maintain control. At the same time, autocorrecting text
messaging systems regularly produce mistakes, alternately amusing and
annoying senders and recipients alike. Still, these applications are widely
appreciated when implemented in a non-intrusive and optional way so they can
be easily ignored.
There is room to build on these Eight Golden Rules with an HCAI pattern
language (Table 9.2). Pattern languages have been developed for many design
challenges from architecture to social media systems. They are brief expressions
of important ideas that suggest solutions to common design problems.
Patterns remind designers of vital ideas in compact phrases meant to provoke
more careful thinking.
1) Overview first, zoom and filter, then details-on-demand: The first one
will be a familiar information visualization mantra that suggests users should be
able to get an overview of all the data, possibly as a scattergram, geographic
map, or a network diagram. This overview shows the scope and context for
individual items, and allows users to zoom in on what they want, filter out what
they don’t want, and then click for details-on demand.
2) Preview first, select and initiate, then manage execution: For temporal
sequences or robot operations the second pattern is to show a preview of the
entire process; it allows users to select their plan, initiate the activity, and then
manage the execution. This is what navigation tools and digital cameras do so
successfully.
3) Steer by way of interactive control panels: Enable users to steer the
process or device by way of interactive control panels. This is what users do
when driving cars, flying drones, or playing video games. The control panel can
include joysticks, buttons, or onscreen buttons, sliders, and other controls, often
placed on maps, rooms, or imaginary spaces. Augmented and virtual reality
extend the possibilities.
4) Capture history and audit trails from powerful sensors: Aircraft sensors
record engine temperature, fuel flow, and dozens of other values, which are
saved on the flight data recorder, but are also useful for pilots who want to
review what happened 10 minutes ago. Cars and trucks record many items for
maintenance reviews; so should applications, websites, data exploration tools,
and machine learning models, so users can review their history easily.
5) People thrive on human-to-human communication: Applications are
improved when users can easily share content, ask for help, and jointly edit
documents. Remember the bumper sticker: Humans in the Group; Computers in
the Loop.
6) Be cautious when outcomes are consequential: When applications can
affect people’s lives, violate privacy, create physical damage, or cause injury,
thorough evaluations and continuous monitoring become vital. Independent
oversight helps limit damage. Humility is a good attribute for designers to have.
7) Prevent adversarial attacks: Failures can come not only from biased data
and flawed software but also from attacks by malicious actors or vandals who
would put technology to harmful purposes or simply disrupt normal use.
8) Incident reporting websites accelerate refinement: Openness to feedback
from users and stakeholders will bring designers information about failures and
near misses, which will enable them to continuously improve their technology
products and services.