COMP4045/7450
Chapter 2: User Interface
Design Process – Part 3
Dr. Li Chen
October 2018
Outline
Evaluation
Introduction
Human-computer method
Cognitive walkthrough
Usability testing
Heuristic evaluation
Why evaluate?
How would you find out whether your
product would appeal to the targeted
users, and if they will use it?
Evaluation is integral to the design
process
Evaluation focuses on both the usability
of the system and on the users’
experience when interacting with it
What to evaluate?
Ranges from low-fidelity prototypes to complete
systems; a particular screen function to the
whole workflow; from aesthetic design to safety
features
For example,
If users find items faster with a new web browser
How engaging and fun a new game app is and how
long users will play it
If a computerized system for controlling traffic lights
results in fewer accidents
Where to evaluate?
Depends on what is being evaluated
For example,
Web accessibility and design choices (e.g.,
choosing the size and layout of keys for a
smartphone) are often evaluated in a lab
User experience aspects can be evaluated more
effectively in natural settings
Remote studies of online behavior can be
conducted for social networking to evaluate
natural interactions
Three types of evaluation
Controlled settings involving users
Enable evaluators to control what users do,
when they do it, and for how long
To test hypotheses and measure or observe
certain behavior
To reduce outside influences and distractions
Extensively and successfully used to evaluate
software applications running on PCs
Main methods: human-compute method,
usability testing
Cont.
Natural settings involving users
Little or no control of users’ activities in order to
determine how the product would be used in the
real world
Examples are online communities and products
that are used in public places
Main method: field studies
Cont.
Any settings NOT involving users
Consultants and researchers critique, predict,
and model aspects of the interface in order to
identify the most obvious usability problems
Imagine or model how an interface is likely to
be used
Main methods: cognitive walkthrough,
heuristic evaluation
Outline
Evaluation
Introduction
Human-computer method
Cognitive walkthrough
Usability testing
Heuristic evaluation
Human-compute method
Formalized method of doing Paper Prototype
testing
One person plays the “computer” and moves the paper prototypes
around in response to the participant’s actions
One person plays the “facilitator” who is in charge of making sure
the study runs smoothly
Users indicate where they want to click to find the information, and
“computer” changes the page to show that screen
When to use:
When you need more formal or in-depth feedback than just
showing someone your designs
Ref: [Link] 11
Cont.
Examples:
[Link]
[Link]
ature=[Link]
Preparing for a Test
Select your users
understand background of intended users
use a questionnaire to get the people you need
don’t use friends or family
Prepare scenarios that are
typical of the product during actual use
make prototype support these (small, yet broad)
Practice to avoid “bugs”
Ref: [Link] 13
Conducting a Test
Three roles
facilitator – only team member who speaks
Gives instructions & encourages thoughts, opinions
Must be clear & detailed
Keeps getting “output” from participant, e.g., “What are you
thinking right now?” (Think aloud)
computer – knows application logic & controls it
always simulates the response, w/o explanation
observers – notes down users’ feedbacks and
comments
Typical session is 1 hour
Ref: [Link] 14
Evaluating Results
Sort & prioritize observations
what was important?
lots of problems in the same area?
Create a written report on findings
gives agenda for meeting on design changes
Make changes & iterate
Ref: [Link] 15
Outline
Evaluation
Introduction
Human-computer method
Cognitive walkthrough
Usability testing
Heuristic evaluation
Cognitive walkthrough
A method that evaluates whether the order of cues and
prompts in a system supports the way people process
tasks and anticipate the “next steps” of a system.
When to use it:
Initial evaluation of a system
Low budget
Walk-up-and-use systems or first-use situations
Access to HCI experts
When to not use it:
Formal evaluation of your own system with you as an evaluator
Example:
[Link]
Ref: [Link]
Cont.
Designer presents an aspect of the
design & usage scenarios
Evaluator is told the assumptions about
user population, context of use, task
details
One or more evaluators walk through
the design prototype with the
scenarios
Evaluators are guided by four questions
18
The Four Questions
1. Will users want to produce whatever effect the
action has?
2. Will users see the control (button, menu, label,
etc.) for the action?
3. Once users find the control, will they recognize
that it will produce the effect they want?
4. After the action is taken, will users understand
the feedback they get, so they can confidently
continue on to the next action?
Ref: [Link]
Stages of a Cognitive Walkthrough
Briefing session to tell experts what to do
Evaluation period of 1-2 hours where:
Each expert works separately
Take one pass to get a feel for the product
Take a second pass to focus on specific
features
Debrief session in which experts work
together to prioritize problems
Ref: [Link]
Outline
Evaluation
Introduction
Human-computer method
Cognitive walkthrough
Usability testing
Heuristic evaluation
Content
1. What is usability testing?
2. Who runs the experiment?
3. Where to do the usability testing?
4. Who will join the test?
5. What will users do?
6. What are to be measured, and how?
7. What to report?
22
1. What is usability testing?
To determine whether an interface is usable by the
intended user population to carry out the tasks for
which it was designed
To identify any usability problems, collect qualitative
and quantitative data and determine the participant's
satisfaction with the product
To recommend improvements to designers
“Usability testing is one of the least glamorous, but
most important aspects of user experience
research.” Madrigal and McClain (2010)
23
Usability Test Proposal
A report that contains
objective
description of system being testing
task environment & materials
participants
methodology
tasks
test measures
Get approved & then reuse for final report
Seems tedious, but writing this will help
“debug” your test
Ref: [Link] 24
2. Who runs the experiment?
Trained usability engineers know how
to run a valid study
Instructor: instruct the user (explain the test
scripts), encourage users to speak aloud their
thoughts, and help users whenever it is
necessary (note: do not interfere and do not
help early)
Observer (note-taker): note down problems
that the user encounters, record the time and
paths that users have taken
Useful for developers & designers to
watch
Available if system crashes or user gets
completely stuck
But have to keep them from interfering 25
3. Where to do the usability testing?
Fixed laboratory:
Environmental settings
Involves recording
performance of typical users
doing typical tasks
Data is recorded on video &
key presses are logged
E.g., if participants repeatedly
pick the wrong menu item,
observer realizes that the label or
placement needs to be changed
26
4. Who will join the test? Getting Users
Participants should be chosen to represent the
intended user communities
Example: for our university’s website
([Link])
Three typical user groups
Current students, age from 18 to 25, f/m
Prospective students, age below 22, f/m
Faculties, age from 30 to 65, f/m
Try to include representatives of all these groups
27
Cont.
Number of users required
Jared Spool claims, for large and complete web sites
Only found 35% of problems after 5 users
Needed about 25 users to get 85% of the problems
Participation should always be voluntary; ethical
approval and informed consent should be
obtained
Use incentives to get participants
T-shirt, mug, free coffee/pizza, etc.
28
5. What will users do? Task Scenarios
Good scenarios are concise but answer the following key
questions:
Who is the user?
Why does the user come to the system? Note what motivates
the user to come to the system and their expectations upon
arrival, if any
What goals does he/she have?
Goal-based scenarios
State only what the user wants to do. Do not include any
information on how the user would complete the scenario
Realistic and compelling so users are motivated to finish
Short enough to be finished
Can let users create their own tasks if relevant
29
Instructions to Participants
Have an explicit script of what will say
Describe the purpose of the evaluation
“I’m testing the product; I’m not testing you.”
Tell them they can quit at any time
Demonstrate the interface
Explain how to think aloud
Explain that you will not provide help
Check if your friend
Describe the task has called, find out
what time he will be
going o the club.
30
Ref: [Link]
6. What are to be collected?
Two Types of Data to Collect
Process data
observations of what users
are doing & thinking
Qualitative [Link]
Bottom-line data
summary of what happened
• time, errors, success
Quantitative
[Link]
31
Ref: [Link]
The “Thinking Aloud” Method
Need to know what users are thinking, not
just what they are doing
Ask users to talk while performing tasks
tell us what they are reading
tell us what they are thinking
tell us what they are trying to do
tell us questions that arise as they work I’ll click on
the
checkout
button
Prompt the user to keep talking
“tell me what you are thinking”
32
Ref: [Link]
Quantitative data
Time to learn: how long does it take for typical members of
the community to learn how to perform relevant tasks with the
interface? (learnability)
Success: percent of users who have completed task
successfully (effectiveness)
Speed of performance: How long does it take to perform
relevant tasks? (efficiency)
Rate of errors by users: How many and what kinds of errors
are made by users during performing the tasks? (safety)
Retention over time: How well do users maintain their
knowledge after an hour, a day, or a week? (memorability)
User preference: How much did users like using various
aspects of the interface?
Note: It does not mean we need to cover all of them in a test
33
Measuring User Preference
n-point Likert scale
Propose something and let people agree or disagree:
strongly disagree strongly agree
The system was easy to use: 1 .. 2 .. 3 .. 4 .. 5
Two opposite feelings:
difficult easy
Finding the right information was: 1 .. 2 .. 3 .. 4 .. 5
If multiple choices, rank them:
Rank the choices in order of preference (with 1 being most preferred and 4 being
least):
Interface #1 Interface #2 Interface #3 Interface #4
34
Questionnaire: System Usability Scale (SUS)
SUS has become an industry standard, with
references in over 1300 articles and publications
It allows to evaluate a wide variety of products and
services, including hardware, software, mobile
devices, websites and applications
5-point Likert
scale
35
[Link]
Questionnaire for User Interface Satisfaction (QUIS)
36
QUIS (cont.)
37
7. What to report?
Quantitative data: success rates, task time, error rates,
and satisfaction questionnaire ratings
Qualitative data: observations about pathways
participants took, problems experienced,
comments/answers to open-ended questions, and
recommendations
Images & graphs help people get it!
Video clips can be quite convincing
Template: [Link]
tools/resources/templates/report-template-usability-
[Link]
38
Example
“Overall, I am satisfied “The interface of this
with this website” website is pleasant”
19 16
20
14
18 14
12
16 12
14
10 9
12
10 9 8
8
8 6
6
4 3
4
2 2
2 1 1
0 0
strongly agree not sure disagree strongly strongly agree not sure disagree strongly
agree disagree agree disagree
Mean rating = Mean rating = 3.03
(2*5+19*4+8*3+9*2+1*1)/39 = 3.31 Percent agree = (1+14)/39 = 38.5%
Percent agree = (2+19)/39 = 54%
39
Cont.
Summary of usability problems
Organize problems by scope and impact
How widespread is the problem? (scope)
• How many users have encountered similar problem?
How critical is the problem? (impact)
• Did the problem impede users from completing the task?
Proportion of users experiencing the problem
Few Many
Impact of the Small Low Severity Medium Severity
problem on the
users who
experience it Large Medium Severity High Severity
40
Example
1. Users complain that the login procedure requires too
many steps: Eight dialog boxes
few scope, small impact
2. Can’t copy info from one window to another
few scope, large impact
3. The interface used the string "Save" on the first screen
for saving the user's file, but used the string "Write file"
on the second screen. Users were confused by this
different terminology for the same function.
large scope, large impact
4. Spelling mistakes, e.g., Contact us -> Contatc us
large scope, small impact
41
Example, cont. Proportion of users experiencing the problem
Few Many
Low severity Medium severity
Impact of the Small
Problem 1 Problem 4
problem on the
users who
Medium severity High severity
experience it Large
Problem 2 Problem 3
Rank all problems by their severity levels in the descending order
So, for the above example, the four problems are ranked as:
Problem 3 (high severity)
Problem 2 (medium severity)
Problem 4 (medium severity)
Problem 1 (low severity)
42
How to recommend improvements?
Guidelines: Making usability
recommendations useful
Ensure the recommendations improves the
overall usability of the application
Make recommendations specific and clear
Be aware of the business or technical
constraints
Avoid vagueness by including specific
examples in your recommendations
43
So
Bad examples
“Improve the layout and interface.”
“Make the reserve function more clear and easy to
use.”
“To ease the search for books.”
“To show all the functions of the library clearly.”
Good example
“Automatically show the publications by its publication
years (in descending order thus the newest book can
show first), then add a button for the user to select by
ascending or descending order.”
44
Stages of a Usability Testing
Stage 1: Preparation
Make sure test ready to go before user arrives
Stage 2: Introduction
Say purpose is to test software, not the user!
Consent form explains procedures and deals with ethical
issues
Give instructions
Pre-test questionnaire (personal info, background, etc.)
Stage 3: Running the test
Stage 4: Debriefing after the test
Post-test questionnaire, open-ended questions, thanks
45
Best practices
Respect the participants
You are testing the interfaces NOT the users
Remain neutral
Do NOT direct the participants
Decide how much of a hint you will give and the maximum
time for a scenario
Take good notes, make analysis easier
Measure both objective and subjective data
Reference:
[Link]
46
Outline
Evaluation
Introduction
Human-computer method
Cognitive walkthrough
Usability testing
Heuristic evaluation
Content
Heuristic Evaluation Overview
Heuristic Evaluation vs. Usability Testing
Ten Heuristics
How to Perform Heuristic Evaluation
48
Heuristic Evaluation (HE)
Developed by Jakob Nielsen
Evaluation is performed by experts, whose
expertise is in the application or user-
interface domain and capable of helping
find usability problems in a UI design
Ref: [Link] 49
Small set (3-5) of experts (evaluators) to
examine UI by role-playing typical users
Independently check for compliance with usability
principles (“heuristics”)
Different evaluators will find different problems
Evaluators only communicate afterwards
Findings are then aggregated
Redesign/fix problems
Can perform on working (high-fidelity) UI or on
sketches
No. of evaluators & problems
Empirical evidence suggests that on average 5 evaluators identify 75-
80% of usability problems 51
Benefits and Cost Balance
problems found benefits / cost
So Jakob Nielsen suggests that the optimal number of
evaluators might be between 3 and 5
52
Experts End users
Usability testing
Heuristic evaluation
Vs Evaluate the UI by
Evaluate the UI based
performing a set of real
on a set of heuristics
tasks
More realistic
Much faster
More accurate
Less participants
Find more problems
53
Heuristic Evaluation (HE) vs. Usability
Testing (UT)
Heuristic Evaluation
is much faster
1-2 hours each evaluator (HE) vs. days-weeks (UT)
HE doesn’t require interpreting user’s actions
HE may miss problems & find“false positives”
Usability Testing
Usability testing is far more accurate (by definition)
Takes into account actual users and tasks
Know how typical users, especially first-time users, will
behave
Good to alternate between HE & UT
Find different problems
Don’t waste participants
54
Heuristic Evaluation Process
Usability heuristics
Nielsen’s “10 heuristics”
These guidelines closely resemble the high-level design
principles, e.g., reduce users’ memory load, make design
consistent, and use terms users understand
Evaluators go through UI several times (normally
twice)
Inspect various dialogue/interface elements
Check with the list of usability heuristics
Consider other principles that come to mind
55
Nielsen’s 10 heuristics
1. Visibility of system status
2. Match between system and real world
3. User control and freedom
4. Consistency and standards
5. Error prevention
6. Recognition rather than recall
7. Flexibility and efficiency of use
8. Aesthetic and minimalist design
9. Help users recognize, diagnose, recover from errors
10. Help and documentation
[Link]
56
Heuristics 1
1: Provide feedback (visibility of system
status)
Keep users informed about what is going on
Example
Changing the cursor to show whether a map interface is
in zoom-in or select mode
Pay attention to response time
• Less than 1.0 sec: no special indicators needed
• 1.0 sec to 10 sec: indicate max. duration if user stays
focused on action
• for longer delays, use percent-done progress bars
57
Provide feedback
User should always be aware of what is going on
So continuously inform the user about
What the system is doing
How it is interpreting the user’s input
What’s it
Time for
doing?
coffee.
> Doit
> Doit
This will take
5 minutes...
58
Bad Example – The title does not
contain the progress status
The title should take the form of
"[XX]% Completed - [Task
Description] - [Product Name]”
Bad example – this progress bar looks like it is stuck at 99%.
Ideally the progress bar should be hidden when completed
and replaced by a green tick 59
Good Example – users may change
Good Example - Microsoft Update uses 3 their mind about what they are doing
icons to indicate different status
60
Heuristics 2
2: Match between system & real world
Terminology in user’s language
Not computer terminology, not “techno-jargon”
Language from user’s perspective
“You have bought…” not “We have sold you…”
61
Heuristics 3
3: User control and freedom
Easy to abort: with “Cancel” buttons
E.g., cancel an order, cancel changing a profile
Easy to undo
“Back” button permits easy reversal of actions. If users know
that errors can be undone, they could be encouraged to
explore new and unfamiliar options
62
Heuristics 4
4: Consistency and standards
Same commands always have the same effect
Following standards/conventions will help
Locations for information, names of commands
Web: use templates or CSS, style guides
Inconsistent,
why?
63
1st screen
Standard Message text in
icon set Arial 14, left
2nd screen
Good:
adjusted
?
Do you really want
to delete the file
“[Link]” from
No Ok the folder “junk”?
Bad:
No Ok
2nd screen
Apply
The file was
destroyed
Cancel
64
Heuristics 5
5: Error prevention (revisit “safety”
usability goal)
Gray out menu items that are not
appropriate
Modern widgets: only “legal commands”
selected, or “legal data” entered
Provide reasonable checks for users to re-
consider their intentions
65
Heuristics 6
6: Recognition rather than recall
Because our short memory has limited size
Avoid interfaces in which users must remember
information from one screen and then use that
information on another screen
Make objects, actions, options, & directions
visible/easily recognizable
66
Which way is easier?
Type the URL by yourself
Select from the list
67
Heuristics 7
7: Flexibility and efficiency of use
Add explanations for novice users
Provide shortcuts and fast pacing for experienced users
Reuse users’ previously entered information
Give good default values
68
Heuristics 8
8: Aesthetic and minimalist design
“KISS” principle: Keep it Simple and Stupid
“Less is More”: Identify what is really needed
Extra options can confuse users
Avoid excessive use of color (magic number: 4)
Avoid excessive use of animation and blinking effects
69
Principle: KISS
Keep It Simple and
Stupid!
Variations
“Keep it short and simple",
"Keep it simple and
stimulating”
“Keep it simple and
straightforward"
70
Can you find how to check in?
[Link]
71
Can you find how to check in in the new page?
[Link]
Card-based user interface
72
Heuristics 9
9: Help users recognize, diagnose, & recover
from errors (revisit “safety” usability goal)
Error messages in plain language
Precisely indicate the problem
Constructively suggest a solution
Provides immediate
feedback with specific
instructions
73
Good Error Messages
Clearly indicate what
has gone wrong
Human readable
Polite
Describe the problem
Explain how to fix it
Highly noticeable
74
Ref: [Link]
Cont.
Be as specific and precise as possible
Not too simple, or too general, e.g., “syntax error” (not
enough information), “FAC RJCT 004004400400” (too
obscure)
Use positive tone
Avoid “Fatal error”, “catastrophic error”, “disastrous error”;
negative terms, such as illegal, invalid, bad, error, should be
used infrequently
Example:
Poor: “Disastrous string overflow. Job abandoned.”
Better: “Revise program to use shorter strings or expand
string space.”
Poor: “Undefined labels.”
Better: “Define statement labels before use.”
75
Heuristics 10
10: Help and documentation
Easy to search
Focused on the user’s task
List concrete steps to carry out
Not too large
Help tips are displayed on hover, answering the Embedded videos can be used to showcase
most likely questions about a field or instructions features as well as get people started using the
76
product
Stages of Heuristic Evaluation
Stage 1: Pre-evaluation training
Give evaluators needed domain knowledge
Stage 2: Evaluation
Have evaluators go through the UI and try different user-interface
elements, such as dialog boxes, menus, navigation structure, online help,
etc.
Ask them to see if it complies with heuristics
Note where it doesn’t & say why
Make the list of problems
Stage 3: Severity rating
Determine how severe each problem is
Have evaluators independently rate severity
Stage 4: Aggregation
Combine the findings from 3 to 5 evaluators
Group meets & aggregates problems (with ratings)
Stage 5: Debriefing
Discuss the outcome with designer team
77
Severity Ratings
0 – not a real usability problem
1 – need not be fixed unless extra time is
available on project
2 – minor usability problem – low priority
3 – major usability problem – important to
fix
4 – usability catastrophe – imperative to fix
before releasing product – high priority
Ref: [Link]
78
Ref: [Link]
Ref: [Link]
Best practices
Have evaluators go through the UI twice
Ask them to see if it complies with heuristics
note where it doesn’t & say why
Have evaluators independently rate severity
Combine the findings from 3 to 5 evaluators
Alternate with usability testing
Supplementary materials:
[Link]
Shneiderman’s 8 Golden Rules ([Link]
[Link]/literature/article/shneiderman-s-eight-golden-rules-
will-help-you-design-better-interfaces)
81
Ref: [Link]
Summary of Chapter 2
Getting requirements right is crucial
Scenarios and use cases are used to articulate
existing and envisioned work practices
Task analysis techniques such as hierarchical task
analysis (HTA) help to investigate existing systems
and practices
Two aspects of design: conceptual and physical
Different kinds of prototyping are used for different
purposes and at different stages
Different evaluation methods: human-computer
method, cognitive walkthrough, usability testing,
heuristic evaluation
82
Reference
Chapters 9-11, 13-15 of “Interaction
Design: Beyond Human-Computer
Interaction”, 4th edition, Wiley, 2015.
83