0% found this document useful (0 votes)
19 views22 pages

Design Principles For Generative AI Applications

This document presents six design principles for generative AI applications that address unique challenges and characteristics of generative AI user experiences. The principles aim to guide the effective and safe design of user interactions with generative AI technologies, supported by practical strategies for implementation. Developed through iterative validation and feedback from design practitioners, these principles are intended to inform actionable design recommendations in the field of human-computer interaction.

Uploaded by

puzachvladis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views22 pages

Design Principles For Generative AI Applications

This document presents six design principles for generative AI applications that address unique challenges and characteristics of generative AI user experiences. The principles aim to guide the effective and safe design of user interactions with generative AI technologies, supported by practical strategies for implementation. Developed through iterative validation and feedback from design practitioners, these principles are intended to inform actionable design recommendations in the field of human-computer interaction.

Uploaded by

puzachvladis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Design Principles for Generative AI Applications

Justin D. Weisz Jessica He Michael Muller


jweisz@[Link] jessicahe@[Link] michael_muller@[Link]
IBM Research AI IBM Research AI IBM Research AI
Yorktown Heights, NY, USA Seattle, WA, USA Cambridge, MA, USA

Gabriela Hoefer Rachel Miles Werner Geyer


ghoefer@[Link] [Link]@[Link] [Link]@[Link]
IBM IBM IBM Research AI
New York, NY, USA San Jose, CA, USA Cambridge, MA, USA

New interpretations

Design
Design for Design for

Design Principles

Responsibly Mental Models Appropriate Trust

& Reliance

New characteristics

Design for
Design for
Design for

Generative Variability Co-Creation Imperfection


User Goals

Exploration

Optimization

Figure 1: Six principles for the design of generative AI applications. Three principles ofer new interpretations of known issues
with AI systems through the lens of generative AI, and three principles identify unique characteristics of generative AI systems.
The principles support two user goals: optimizing a generated artifact to satisfy task-specifc criteria, and exploring diferent
possibilities within a domain.
ABSTRACT We present six principles for the design of generative AI applica-
Generative AI applications present unique design challenges. As tions that address unique characteristics of generative AI UX and
generative AI technologies are increasingly being incorporated into ofer new interpretations and extensions of known issues in the
mainstream applications, there is an urgent need for guidance on design of AI applications. Each principle is coupled with a set of
how to design user experiences that foster efective and safe use. design strategies for implementing that principle via UX capabil-
ities or through the design process. The principles and strategies
were developed through an iterative process involving literature
review, feedback from design practitioners, validation against real-
This work is licensed under a Creative Commons Attribution International world generative AI applications, and incorporation into the design
4.0 License. process of two generative AI applications. We anticipate the princi-
ples to usefully inform the design of generative AI applications by
CHI ’24, May 11–16, 2024, Honolulu, HI, USA
© 2024 Copyright held by the owner/author(s). driving actionable design recommendations.
ACM ISBN 979-8-4007-0330-0/24/05
[Link]
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

CCS CONCEPTS line interfaces, and graphical user interfaces), because it shifts con-
• Human-centered computing → HCI design and evaluation trol over how computation is performed away from the user and
methods; Interaction paradigms; HCI theory, concepts and models. toward generative AI models. With this shift in control, how are we
to design user experiences that help people interact with generative
KEYWORDS AI applications in efective and safe ways?
Over at least the past four decades, researchers and practitioners
Generative AI, design principles, human-centered AI, foundation
within human-computer interaction (HCI) have produced numer-
models
ous guidelines, principles, practices, and frameworks for the design
ACM Reference Format: of efective and safe computing systems. Some guidelines are pre-
Justin D. Weisz, Jessica He, Michael Muller, Gabriela Hoefer, Rachel Miles, sented as generally applicable to most kinds of interactive comput-
and Werner Geyer. 2024. Design Principles for Generative AI Applications. ing systems, such as Nielsen and Molich’s heuristics [128] and Shnei-
In Proceedings of the CHI Conference on Human Factors in Computing Systems derman et al.’s strategies for designing efective human-computer
(CHI ’24), May 11–16, 2024, Honolulu, HI, USA. ACM, New York, NY, USA,
interaction [155]. Other design guidelines are technology-specifc,
22 pages. [Link]
such as Bevan’s guidelines for web usability [16] and various guide-
lines for AI systems and human-AI interaction [7, 66, 133]. However,
1 INTRODUCTION none of the guidelines developed within the HCI community have
Generative AI technologies have reached an infection point in yet addressed the specifc nuances present with generative AI.
consumer adoption and enterprise value, sparked by technological In this paper, we build upon the HCI tradition of distilling cu-
advancements in machine learning architectures such as GANs [56, mulative knowledge into a practical set of principles, specifcally
79], VAEs [86], and transformers [38, 170]. Models such as Style- for the design of generative AI user experiences (UX). We opted to
GAN [79], GPT [20, 130, 137, 138], and Codex [28] have demon- produce a set of principles, rather than guidelines, to highlight their
strated that powerful generative models can produce works at a fundamental nature to the design of generative AI UX.
human-like level of fdelity. Today, consumer applications such as Our paper makes the following contributions to the CHI com-
ChatGPT1 , DreamStudio2 , and DALL-E3 are making these technolo- munity:
gies widely available and setting the bar for people’s expectations
of what generative AI can do. Startups such as Cohere4 and An- • We introduce a set of six principles for the design of gener-
thropic5 are reducing the friction of embedding large language ative AI applications (Table 1). Three principles identify new
models in consumer applications. Enterprises such as IBM, Mi- considerations for generative AI, and three ofer new inter-
crosoft, Amazon, and Google are creating platforms for businesses pretations of known issues with AI systems. Each principle
to infuse generative technologies into their products and services. is coupled with a set of practical strategies and examples
This commercialization of generative AI technologies is fueled by (Appendix A) for how to put the principle into practice –
the ultra-rapid development of large-scale foundation models [19] either through the design process itself or through specifc
that reduce the time and costs for developing generative AI sys- UX capabilities – in order to aid design practitioners6 in
tems. However, much attention in machine learning research com- applying the principles to real-world UX.
munities has focused on developing advancements to the technol- • We establish the practical value of the design principles
ogy: scaling model parameter counts [91, 158], evaluating model by conducting a rigorous and systematic validation that in-
performance [97, 163, 194], tuning models efciently to perform cluded multiple rounds of testing, feedback collection, it-
new tasks [27, 178], and aligning models [132, 196] to reduce their eration, and use by two generative AI application design
propensity to produce speech that is hateful, abusive, profane, or teams.
otherwise toxic [60, 70]. Although these advancements serve to • We provide insight on how we made decisions regarding
improve the state of the art, they do not recognize an important the organization of the principles, navigated issues of con-
half of what Ehsan et al. [43] call the “human-AI assemblage” – the ceptual redundancy and overlap, and fostered adoption of
human. the principles within our organization.
Generative models have enabled a radically new way for people
to interact with computing technologies. People are now able to 2 RELATED WORK
craft specifcations for the kinds of outputs they desire, such as Over the past four decades, the HCI community has conducted
via natural language prompts, and generative models are able to numerous examinations of user interface design for computing
produce outputs that conform to those specifcations. Nielsen [127] technologies and distilled best practices into various forms of de-
recently identifed this form of interaction as intent-based outcome sign guidelines. We identify and review two categories of these
specifcation and argued that it is the frst new UI interaction para- guidelines: those that apply to the general UX design of computing
digm in 60 years. This form of interaction is fundamentally diferent systems and those that apply specifcally to AI-infused systems.
from previous interaction paradigms (e.g. punchcards, command

1 ChatGPT. [Link]
2 DreamStudio. [Link]
3 DALL-E. [Link] 6We consider a design practitioner to be anyone involved in making deliberate deci-
4 Cohere. [Link]
sions that impact the design of an application, including user researchers, interaction
5 Anthropic. [Link] designers, visual designers, product managers, and more.
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

2.1 Guidelines for human-computer interaction The advent and widespread commercial adoption of smartphones,
and user interface design mobile apps, app stores, and mobile web sites necessitated yet
another new design language, optimized for smaller screens and
Design guidelines have a rich history in HCI research. We provide
touch-based interactions. Some work in this space focused on de-
a brief historical sketch of human-computer interaction design
veloping design frameworks and guidelines for mobile apps and
guidelines and their evolution as a result of new paradigms in user
workfows (e.g., [61, 78, 109, 124, 136, 151]). Guidelines also emerged
interface technologies.
covering mobile design for more specialized populations, including
Licklider and Clark provide an early example of guidelines for the
older users [5, 30], users of courseware [71], users from diverse
design of console interfaces in, “On-line man-computer communica-
cultures [5], users with disabilities [134], and users with diverse
tion” [99] (refecting the gendered language of their time). Although
literacies [162]. Guidelines were also developed for specifc mobile
their guidelines are not presented in a modernly-recognizable form7 ,
app domains such as health care [4, 75] and fnance [3, 64, 118],
their work does introduce ideas about humans’ needs when working
as well as ethical concerns around privacy [94, 96] and the use of
with computers. Other early forms of human-computer interaction
mobile apps for research purposes [115, 142].
guidelines include those for the Spacelab Experiment Computer
Application Software (ECAS) [39], and for the command terminals
of battlefeld automated systems [156, 157]. Even in these early 2.2 Guidelines for human-AI interaction
days of user interface design, it was recognized that, “[l]acking con- Within the past few years, the emergence of AI as a design ma-
sistent design principles, current practice results in a fragmented terial [41, 46, 62, 191] has necessitated guidelines that inform its
and unsystematic approach to system design, especially where the use. A growing body of work within the human-centered AI re-
user/operator-system interaction is concerned.” [156, p. v]. search community has proposed best practices for human-AI inter-
A signifcant infection point in the design of user interfaces for action in the form of design guidelines (e.g., [7, 11, 104, 186, 192]),
computing technologies came with the rise of personal computing formal studies (e.g., [24, 102]), toolkits (e.g., [110]), and reviews
and the concurrent emergence of graphical user interfaces (GUIs) (e.g., [59, 72, 180, 188]).
as a new interaction paradigm. New guidelines were needed for Some of these guidelines include claims of universal applicabil-
the new visual metaphors of windows, icons, menus, and point- ity8 or being of a general nature to AI-infused systems (e.g., [7, 153]).
ers. Apple’s Human Interface Guidelines [8] provides a prominent Other guidelines focus on specifc types of AI technologies (e.g.
example of practical guidelines for GUI design. Smith and Mosier text-to-image models [104]), specifc domains of use (e.g. creative
[159] developed more formal guidelines for GUIs in which they writing [24]), or specifc issues regarding the use of AI, including
described six functional areas: data entry, data display, sequence ethics [11, 59, 69, 72], fairness [110], human rights [50], explainabil-
control, user guidance, data transmission, and data protection. ity [119], and user trust [186]. Finally, as more consumer products in-
As Grudin described in an infuential retrospective analysis [58], corporate AI technologies, industry leaders including Google [133],
the “site” of human-computer interaction began to move away from Microsoft [7, 95] and Apple [9] have developed and published their
a terminal in a lab and “reached out” into other contexts such as own guidelines; Wright et al. [188] provide a comparative analysis
home and ofce environments. New methods were required to of these guidelines.
understand and design for these changing circumstances. With Guidelines that focus on the design of AI systems, and specif-
heuristic evaluation [128], Nielsen provided a set of methods for cally on the ethics of those systems, are critically important. Various
“discount usability engineering” [126] that helped designers more attempts have been made to assist design practitioners in the pro-
easily assess their interfaces, while Lewis and Wharton’s cognitive cess of operationalizing guidelines for AI systems, including guide-
walkthrough method [93] provided a way for designers to conduct books [133], toolkits [36], and checklists [110]. When design guide-
more detailed and tailored analysis. lines are successfully applied, they make a positive impact, such
The next technological infection point that necessitated a shift as in assisting cross-functional development teams in improving
in design guidelines occurred with the rise of the Web. Unlike user experiences [95] and addressing ethical challenges [11]. How-
the previous decades, web design involved diverse and competing ever, several studies have critiqued their comprehensiveness [59],
hardware and software. These complexities led to what Mariage the extent to which they can be operationalized [50], and the lack
et al. describe as a “jungle of guidelines [intended to] address many of consequences when they are not followed [59]. Additionally,
diferent issues” [113, Introduction]. One result was that, out of 11 Madaio et al. [110] argues that the adoption of an AI ethics process
generic web design guidelines, Cappel and Huang found that a mean within an organization, “would only happen if leadership changed
of only 5.5 guidelines were followed across 500 companies’ websites organizational culture to make AI fairness a priority, similar to pri-
[25]. Adding to the diversity, technologically-literate advocates orities and associated organizational changes made by leadership
emerged for people with disabilities [92, 141], older users [14, 90], to support security, accessibility, and privacy” [110, p. 8].
and users from diverse cultures [2]. Bevan [16] acknowledged that Despite the preponderance of guidelines for human-AI interac-
the available wealth of guidelines addressed diferent issues in tion and AI ethics, there is a gap in the technologies on which they
diferent ways for diferent constituencies, and that this situation focus. To date, many of the AI guidelines developed within the HCI
was likely to continue. community primarily focus on discriminative AI [7, 65, 66, 133], the
class of algorithms that identifes boundaries that separate diferent
7 Forexample, they provide guidelines on allocating tasks between humans and com-
puters that “exploit the complementation that exists between human capabilities and 8 Multiple scholars have critiqued the concept of “universal” design as privileging
present computer capabilities” [99, p.114]. certain assumed “normal” populations [12, 168].
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

classes or groups in a data set. These guidelines do not take into use cases in which false positives and false negatives are important
account generative AI algorithms that produce artifacts, rather than outcome metrics; in generative use cases, these metrics have no
decision boundaries, as outputs. Because generative AI ofers new meaning.
ways for users to interact with technology, and raises new issues Guidelines from Amershi et al. [7] may be more readily adapted
regarding the ethics of AI systems, a new set of design guidelines to generative AI applications, although their coverage of generative-
are needed. specifc considerations is limited and design practitioners may en-
counter difculties in making such adaptations. For example, the
3 WHY GENERATIVE AI NEEDS DESIGN recommendation to, “Make clear why the system did what it did” is
PRINCIPLES potentially less important when a user’s goal is to simply generate
a desirable artifact10 . The recommendation to, “Make clear what
Generative AI technologies have introduced a new paradigm of
the system can do” may be difcult to implement in light of the
human-computer interaction, what Nielsen refers to as “intent-
emergent and unanticipated behaviors of generative foundation
based outcome specifcation” [127]. In this paradigm, users specify
models [19], as well as the trial-and-error methods by which users
what they want, often using natural language9 , but not how it should
iterate toward a desired outcome [104, 131, 165, 190].
be produced. One challenge of this paradigm stems from the distin-
Finally, alongside their tremendous potential to augment peo-
guishing characteristic of generative AI: it generates artifacts as
ple’s creative capabilities, generative technologies also introduce
outputs and those outputs may vary in character or quality, even
new risks and potential user harms. These risks include issues of
when a user’s input does not change. This characteristic has been
copyright and intellectual property [47, 107], the circumvention or
described by Weisz et al. [182] as generative variability, and it pro-
reverse-engineering of prompts through attacks [35], the produc-
vides what Alvarado and Waern [6] describe as an “algorithmic
tion of hateful, toxic, or profane language [60], the disclosure of
experience,” raising questions on appropriate types of user control,
sensitive or personal information [83], the production of malicious
levels of algorithmic transparency, and user awareness of how the
source code [26, 28], and a lack of representation of minority groups
algorithms work and how to efectively interact with them.
due to underrepresentation in the training data [51, 108, 167, 171].
With generative AI applications, users will need to develop a
Work by Houde et al. [63] takes concerns such as these to an extreme
new set of skills to work with (not against) generative variability
by envisioning realistic, malicious uses of generative AI technolo-
by learning how to create specifcations that result in artifacts that
gies. Although it cannot be a designer’s responsibility to curb all
match their desired intent. One emerging skill revolves around
potentially-harmful usage, existing design guidelines for AI sys-
crafting efective natural language prompts, known as in-context
tems fall short in addressing these unique issues stemming from the
learning [20, 40, 189] or prompt engineering [185, 193]. This process
generative nature of generative AI, and AI ethics frameworks are
is typically informal and relies on trial-and-error [104, 131, 165, 190].
only just starting to appear to provide designers with the language
The use of open-ended natural language, rather than a fxed vocab-
they need to begin discussing these important issues [36, 66, 180].
ulary of commands, leads to new design challenges. For example,
We therefore conclude that there is a pressing need for a set
Nielsen argues, “users should not have to wonder whether diferent
of general design guidelines that help practitioners develop appli-
words, situations, or actions mean the same thing” [125, p.156];
cations that utilize generative AI technologies in safer and more
given the innumerable ways that users can express their intent in a
efective ways – safer because of the new risks introduced by gener-
natural language prompt, how can generative AI applications help
ative AI, and more efective because of the control that users have
users achieve desired results? Is it necessarily a “mistake” or “error”
lost over the computational process. Although recent work has be-
when a user’s prompt results in an output that they didn’t anticipate
gun to probe at design considerations for generative AI, this work
or like? Does it violate the consistency heuristic when it is difcult
has been limited to specifc application domains or technologies.
for users to achieve replicable results (e.g. [116, 135, 150]), because
For example, guidelines of various maturity levels exist for GAN-
each click of the “generate” button results in diferent outputs, even
based interfaces [57, 195], image creation [102, 104, 173], prompt
for the same input?
engineering [102, 104], virtual reality [169], collaborative story-
Existing human-AI design guidelines fail to address the unique
telling [146], and workfows with co-creative systems [57, 123, 173].
design challenges of generative AI because they do not cover gen-
Our work seeks to extend these studies toward principles that can
erative use cases or new considerations stemming from generative
be used across generative AI domains and technologies.
variability, and they do not cover new or amplifed ethical issues
stemming from the models’ generative nature. For example, guide-
lines published by PAIR make recommendations such as, “Design
for labelers & labeling” and “Design & evaluate the reward func-
tion” [133]. The former of these recommendations will not apply to
generative use cases that do not require data labeling, as the foun-
dation models often used to implement generative capabilities are
pre-trained and may not require additional labeled data for tuning.
In addition, the latter recommendation is tailored to classifcation

9 Some user interfaces to generative AI systems allow users to specify their intent
via sketches and gestures [32], UI controls [103, 105], video-scanned body move-
ments [174], and even various forms of improv performance [15, 68].
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

Design Responsibly Design for Generative Variability


Ensure the AI system solves real user issues and minimizes user Help the user manage the ability of generative models to produce
harms multiple outputs that are distinct and varied
• Use a human-centered approach*. Design for the user by under- • Leverage multiple outputs. Generate multiple outputs that are
standing their needs and pain points, and not for the technology either hidden or visible to the user in order to increase the chance
or its capabilities. of producing one that fts their need.
• Identify & resolve value tensions*. Consider and balance diferent • Visualize the user’s journey. Show the user the outputs they have
values across people involved in the creation, adoption, and usage created and guide them to new output possibilities.
of the AI system. • Enable curation & annotation. Design user-driven or automated
• Expose or limit emergent behaviors. Determine whether generative mechanisms for organizing, labeling, fltering, and/or sorting
capabilities beyond the intended use case should be surfaced to outputs.
the user or restricted. • Draw attention to diferences or variations across outputs. Help the
• Test & monitor for user harms*. Identify relevant user harms (e.g. user identify how outputs generated from the same prompt difer
bias, toxic content, misinformation) and include mechanisms that from each other.
test and monitor for them.

Design for Mental Models Design for Co-Creation


Communicate how to work efectively with the AI system, consid- Enable the user to infuence the generative process and work col-
ering the user’s background and goals laboratively with the AI system
• Orient the user to generative variability. Help the user understand • Help the user craft efective outcome specifcations. Assist the user
the AI system’s behavior and that it may produce multiple, varied in prompting efectively to produce outputs that ft their needs.
outputs for the same input. • Provide generic input parameters. Let the user control generic
• Teach efective use. Help the user learn how to efectively use the aspects of the generative process such as the number of outputs
AI system by providing explanations of features and examples and the random seed used to produce those outputs.
through in-context mechanisms and documentation. • Provide controls relevant to the use case & technology. Let the
• Understand the user’s mental model*. Build upon the user’s ex- user control parameters specifc to their use case, domain, or the
isting mental models and evaluate how they think about your generative AI’s model architecture.
application: its capabilities, limitations, and how to work with it • Support co-editing of generated outputs. Allow both the user and
efectively. the AI system to improve generated outputs.
• Teach the AI system about the user. Capture the user’s expectations,
behaviors, and preferences to improve the AI system’s interac-
tions with them.

Design for Appropriate Trust & Reliance Design for Imperfection


Help the user determine when they should or should not rely on Help the user understand and work with outputs that may not align
the AI system’s outputs by teaching them to be skeptical of quality with their expectations
issues, inaccuracies, biases, underrepresentation, and other issues • Make uncertainty visible. Caution the user that outputs may not
• Calibrate trust using explanations. Be clear and upfront about align with their expectations and identify detectable uncertainties
how well the AI system performs diferent tasks by explaining its or faws.
capabilities and limitations. • Evaluate outputs using domain-specifc metrics. Help the user iden-
• Provide rationales for outputs. Show the user why a particular tify outputs that satisfy measurable quality criteria.
output was generated by identifying the source materials used to • Ofer ways to improve outputs. Provide ways for the user to fx
generate it. faws and improve output quality, such as editing, regenerating,
• Use friction to avoid overreliance. Encourage the user to review or providing alternatives.
and think critically about outputs by designing mechanisms that • Provide feedback mechanisms. Collect user feedback to improve
slow them down at key decision-making points. the training of the AI system.
• Signify the role of the AI . Determine the role the AI system will
take within the user’s workfow.
Table 1: Design principles and strategies for generative AI applications. The left column contains principles that ofer new
interpretations of existing issues in the development of AI applications. The right column contains principles that focus on
new issues that stem from generative AI technologies. Strategies that involve following a design process are indicated with an
asterisk (*).
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

4 DESIGN PRINCIPLES FOR GENERATIVE AI that involve trade-ofs between model capabilities and user
APPLICATIONS needs, motivated by work that focuses simultaneously on
end-users of systems [55] and on designers as strategic and
We begin by presenting our fnal set of six design principles and
collaborative end-users of guidelines [82, 87]; and
their corresponding strategies in Table 1, along with our overall de-
• Sensitize designers to the possible risks of generative AI
sign framework in Figure 1. We also provide extended descriptions
applications and their potential to cause a variety of harms
and examples of each principle and strategy in Appendix A. In the
(inadvertent or intentional), and outline processes that could
rest of this paper, we describe the process we used to develop and
be used to avoid or mitigate those harms (e.g. [66, 180]).
validate these principles and strategies.
The principles are generally presented as high-level “design We used an iterative process to develop and refne the design
for...” statements that indicate the characteristics that are impor- principles, inspired by the process used by Amershi et al. [7] in
tant to consider when making design decisions. Three principles developing their guidelines for human-AI interaction. We crafted
focus on aspects of existing AI systems that have new interpre- an initial set of design principles via a literature search (Section 6),
tations through the lens of generative AI: Design Responsibly, refned those principles via multiple feedback channels (Section 7),
Design for Mental Models, and Design for Appropriate Trust conducted a modifed heuristic evaluation exercise to assess their
& Reliance. Three principles identify unique aspects of gener- clarity and relevance and identify any remaining gaps (Section 8),
ative AI UX: Design for Generative Variability, Design for and fnally applied the principles to two generative AI applications
Co-Creation, and Design for Imperfection. under design to demonstrate their applicability to design practice
Each design principle is coupled with a set of four design strate- (Section 9).
gies for how to implement that principle. In some cases, implement- In each iteration, we engaged in signifcant discussion and re-
ing the principle involves following a design process; in other cases, fection on the feedback gathered from the previous iteration to
it is implemented through the inclusion of specifc types of features produce a new version of the design principles and strategies. In
or functionality. some cases, principles or strategies moved to the next iteration
These principles and strategies can be employed to support two unchanged; in many cases, we made organizational and wording
user goals: (1) optimization, in which the user seeks to produce an changes. We summarize our iterative process and the outcomes of
output that satisfes some task-specifc criteria; and (2) exploration, each iteration in Table 2 and we show how the principles evolved
in which the user uses the generative process to explore a domain, over the iterations in Figure 2.
seek inspiration, and discover alternate possibilities in support of
their own ideation. The ways each principle and strategy are applied Iteration Goal Key Outcomes
may difer by user goal, and we elaborate on these diferences in Iteration 1: Identify relevant re- Observed hierarchy of
Section 10.2. Literature search and examples of high-level design prin-
We note that these principles are just that – principles – and Review generative AI applica- ciples implemented
not hard rules that must be followed in all design processes. Our tion design by specifc UX design
view is that it is up to design practitioners to exercise their best strategies; identifed 7
judgement in deciding whether a principle applies to their particular initial design principles
use case, and whether any particular strategy should (or should Iteration 2: Collect feedback from Developed clearer un-
not) be applied. Feedback conference workshop derstanding of Explo-
and designers within ration vs. Optimization
5 METHODOLOGY our institution purposes of use
Our goal is to produce a set of clear, concise, and relevant design Iteration 3: Test principles for clar- Recognized uniqueness
principles that can be readily applied by design practitioners in the Modifed ity, relevance, and cov- of Generative Variabil-
design of applications that incorporate generative AI technologies. Heuristic erage by having design- ity, Co-Creation, and
We aim for the principles to satisfy the following desiderata: Evaluation ers evaluate commer- Imperfection to gener-
cial generative AI appli- ative AI; re-categorized
• Provide designers with language to discuss UX issues unique cations Exploration and Opti-
to generative AI applications, motivated by work that pro- mization separately as
vides designers with specialized vocabulary for domains user goals
such as video games [23, 81] and IoT [31]; Iteration 4: Demonstrate real- Observed utility of prin-
• Provide designers with concrete strategies and examples that Application world applicability and ciples across ideation
are useful for making difcult design decisions, such as those utility by having two and evaluation design
product teams adopt phases; made clarity im-
10 Research by Sun et al. [166] explores the kinds of questions that users have when
the principles in their provements to strategy
working with a generative AI system, which include questions about how an artifact
was produced. We posit that the utility of a generated artifact need not depend upon own design work descriptions
the mechanics of how that artifact was generated in the same way that a user’s
trust in a decision recommendation is often predicated on an explanation for how
Table 2: Summary of the iterative process we used to develop
that recommendation was produced (e.g. [10, 98, 172]). Further, some applications of the design principles and strategies.
generative AI concern the exploration of a space of multiple possibilities (e.g. [88, 144]),
indicating that in some use cases, how an artifact was generated may be of lesser
importance than the generated artifacts themselves.
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

Iteration 1
Design
Design for
Design for
Design for
Design for
Design for
Design for

Literature Responsibly Mental Models Explanations Multiple Outputs H uman ontrol


C Imperfection Exploration
Review
Iteration 2
Design for

Appropriate
Design for

Feedback Trust & Reliance Optimization

Iteration 3

Modified Design for


Design for

enerati e User goal User goal


Heuristic G

V
v

ariability o reation
C -C

Evaluation

Iteration 4

Application

Design
Design for
Design for
Design for
Design for
Design for
Design for
Design for

Final Responsibly Mental Models Appropriate Trust


Generati e
v
o reation Imperfection Exploration Optimization
& Reliance V ariability C -C

Revisions to Revisions to principle


Principle (new interpretations Principle (new characteristics
User goal
associated strategies and associated strategies of existing AI issues) of generative AI)

Figure 2: Evolution of the design principles across four iterations. During Iteration 3, we recognized that some principles
ofered new interpretations of existing AI system characteristics whereas others identifed new characteristics of generative AI.

6 ITERATION 1: CRAFTING INITIAL DESIGN representative set of commercial generative applications (listed in
PRINCIPLES Table 3) to identify common design patterns.
One characteristic that stood out to us in our review was the
We began our process of identifying design guidelines suitable for
diference between work that identifed important user needs and
generative AI applications by examining recent research in the HCI
the specifc kinds of UX design that supported those needs. For
and AI communities. We conducted a literature review of research
example, one set of papers examined requirements for explainable
studies, guidelines, and analytic frameworks from these communi-
AI (XAI) and human-centered explainable AI (HCXAI) through
ties by searching the ACM Digital Library and Google Scholar for
experimental and heuristic methods [43, 98, 166], motivating “ex-
terms including “generative AI, ” “design guidelines,” and “human-
plainability” as an important high-level concept. Then, when exam-
centered AI.” These searches identifed a set of relevant publications,
ining a commercial generative AI system (ChatGPT), we observed
as well as several recent workshops covering human-AI interaction
how explanations of the system’s capabilities and limitations were
with generative AI: Human-AI Co-Creation with Generative Mod-
provided on the home screen. Observations such as these motivated
els [52, 112, 181], Generative AI and HCI [122], Human-Centered
our development of a two-tier principle/strategy structure in which
AI [120], and Human-Centered Explainable AI [44]. We then con-
a principle articulates an important characteristic or consideration
ducted additional searches for terms found within those workshops’
for a generative AI application and the strategies identify how to
proceedings, including “co-creation,” “human-AI collaboration,” “ex-
implement that principle in the UX.
plainability,” and “creative interfaces.” Our searches yielded a rep-
Our analysis helped us identify several characteristics unique to
resentative sample of work that included new advancements and
generative AI that have implications on the user experience: the
issues in generative AI11 , design guidelines12 , studies of design
models’ capability of producing multiple outputs [116, 135, 150], the
guideline implementation13 , and studies of human interaction with
possibility of faws or imperfections15 within those outputs [183,
AI (and generative AI) systems14 . Finally, to incorporate recent
184], and the various ways that people can control or infuence those
industry developments around generative AI, we also examined a
outputs [85, 103, 105]. We also identifed how generative AI could
11 Papers describing advancements and issues in generative AI, including demon- enable people to explore a space of possibilities [88] as a byproduct
strations of new user interaction capabilities, included Brown et al. [21], Johnson of the generative process. In addition, we identifed several existing
et al. [74], Kaiser et al. [76], Kim et al. [85], Liu and Chilton [103], Louie et al.
[105], Metz [117], Perez et al. [135], Ramesh et al. [139], Rombach et al. [140], Ross considerations of AI systems as being particularly important to the
et al. [143], Sharma [150]. generative case, such as using participatory methods [67] to design
12 Papers ofering design guidelines included Amershi et al. [7], Frijns and Schmidbauer
[49], Fukuda-Parr and Gibbons [50], Jobin et al. [72], Liu and Chilton [104], Mohseni
et al. [119], PAIR [133], Shneiderman [152], Srivastava et al. [162], Urban Davis et al. [148], Spoto and Oleynik [161], Sun et al. [166], Verheijden and Funk [173], Wan et al.
[169], Wickramasinghe et al. [186], Wright et al. [188]. [175], Weidinger et al. [180], Weisz et al. [183, 184].
13 Papers examining design guideline implementation included Alnanih and Ormand- 15We identifed two types of imperfection: (a) faws present in the model’s outputs,
jieva [4], Li et al. [95], Yildirim et al. [191]. such as bugs in source code or hallucinations in Q&A, and (b) misalignments between a
14 Papers examining human interactions with AI (and generative AI) systems in- user’s intent, how that intent is expressed in a prompt, and what the generative model
cluded Deterding et al. [37], Ehsan et al. [43], Gmeiner et al. [54], Grabe et al. [57], Inie produces as a result. Even if category (a) would disappear (e.g. due to technological
et al. [67], Kreminski et al. [88], Liao et al. [98], Lim et al. [101], Liu [102], Louie et al. improvements), category (b) would remain and imperfection would still be a critical
[105], Lubart [106], Maher [111], Megahed et al. [116], Muller et al. [123], Seeber et al. design issue.
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

for real user needs like explainability [44, 166], and understanding 8.1 Method
the role of the AI in the co-creative process [37, 57, 106, 123, 148, Heuristic evaluation is a discount usability method for identifying
161]. violations of usability guidelines in a user interface [128]. Amershi
At this stage, we identifed 7 high-level principles and 22 specifc et al. [7] developed a modifed heuristic evaluation in which eval-
strategies for implementing them. Some strategies were related uators reviewed an AI-infused user experience with the purpose
to multiple principles, and at this stage we allowed the overlap; of evaluating the heuristics themselves. We similarly developed a
in subsequent iterations, we eliminated these redundancies (we modifed heuristic evaluation to evaluate our design principles for
discuss this point further in Section 10.1). generative AI applications. We asked evaluators to examine a range
of commercial generative AI applications and identify examples
7 ITERATION 2: EXTERNAL AND INTERNAL that demonstrate the use of the principles and strategies, as well
FEEDBACK as examples of generative AI-specifc design choices that were not
covered by the principles and strategies. This exercise helped us
We published the frst iteration of the design principles at the evaluate the relevance, clarity, and coverage of the design principles
Human-AI Co-Creation with Generative Models (HAI-GEN) work- and strategies.
shop at IUI [182], attended by approximately 50 researchers from We identifed 9 commercial generative AI applications to use in
academia and industry. At this workshop, we received informal the evaluation, listed in Table 3. We selected these applications due
feedback through discussion sessions and follow-up conversations. to their popularity in consumer or enterprise markets, their ability
We also published this version within our organization as part of to be used within our organization without incurring costs, and the
a design guide on generative AI, which was viewed by over 1,000 range of use cases and output modalities they supported. We also
design practitioners. We created an internal discussion channel on considered applications that incorporated generative AI features in
this guide to receive additional feedback, including points of confu- one of two distinct ways16 : either as the core user experience or as
sion and gaps in our framework. Both sources of informal feedback a component within an existing user experience.
helped us craft the second iteration of the design principles, which
introduced the following major changes:
Application Description AI Incorporation
• We identifed how users’ goals in using a generative AI sys-
ChatGPT Conversational Q&A Core (Web app)
tem can difer, leading us to include two task-specifc princi-
ples: the existing Design for Exploration principle, in sup- Google Bard Conversational Q&A Core (Web app)
port of use cases around ideation, exploration, and learning; DALL-E Text-to-image Core (Web app)
and a new principle, Design for Optimization, in support generator
of use cases for which the production of a singular artifact
is desired. DreamStudio Text-to-image Core (Web app)
• We recognized that explainability needs for generative AI generator
systems, while important, were not necessarily an “end” in Midjourney Text-to-image Component
and of themselves. Rather, explainability is one way to De- generator (Discord)
sign for Appropriate Trust & Reliance, leading us to
Adobe Firefy Text-to-image Component
incorporate existing explainability strategies into this new
Generative Fill generator (Adobe Photoshop)
principle.
• We re-articulated all of the design strategies as rules of action IBM [Link] Prompt playground for Core (Web app)
(e.g. a verb followed by 2-6 words), akin to how Amershi Prompt Lab large language models
et al. phrased their guidelines. Github Copilot Natural language to Component (Visual
• We identifed that fve design strategies were about the de- source code Studio Code)
sign process itself rather than specifc UX capabilities.
AIVA Music generation Core (Web app)
At the end of Iteration 2, we had a set of 8 high-level principles Table 3: Commercial generative AI applications used in our
implemented by 29 specifc strategies. modifed heuristic evaluation. AI capabilities were present
either as the core user experience or embedded as a compo-
8 ITERATION 3: MODIFIED HEURISTIC nent within an existing application.
EVALUATION
Following Iteration 2, we sought to conduct a more rigorous evalua-
tion of the design principles and strategies. Given the potential gap
between research literature and real-world practice, we specifcally
wanted to determine their clarity to our target audience of design
practitioners, understand their relevance to commercial generative
16 Beavers et al. [13] refer to diferent “altitudes” at which AI capabilities can be
AI applications, and identify any additional gaps in our framework.
embedded in a UX, including “above” (the AI capability is the central focus of the
In support of these goals, we drew inspiration from Amershi et al. application), “beside” (the AI capability sits in a side panel next to the main UI), and
by creating a modifed heuristic evaluation exercise. “inside” (the AI capability is embedded within the existing UI).
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

We recruited 18 design practitioners within our organization 8.2 Results


and outside of our immediate team to perform the modifed heuris- Our evaluators produced 18 heuristic evaluation canvases laden
tic evaluation. We sought evaluators with varied design roles and with screenshots and sticky notes that identifed real-world in-
levels of experience to ensure the principles were clear and relevant stances of the principles and strategies. They also left notes about
across diferent specialties and expertise levels. Of the 18 evalua- points of difculty or confusion.
tors, 11 (61.1%) identifed as male, 6 (33.3%) identifed as female, Figure 3 shows a portion of one evaluator’s canvas in which they
and 1 preferred not to disclose. The majority of evaluators were evaluated AIVA for Design for Optimization. This example shows
User Experience Designers (16, 88.9%), one evaluator was a Design how the evaluator found examples of various kinds of controls in the
Researcher, and one was a Research Software Developer17 . Four tool, along with feedback on the repetitiveness between Leverage
evaluators (22.2%) reported having 1-4 years of experience, four multiple outputs, Show multiple outputs, and Design for Multiple
(22.2%) had 5-9 years, two (11.1%) had 10-14 years, two (11.1%) had Outputs: “Again?? I’m not copying my examples another time.”
15-19 years, and fve (27.7%) had 20+ years18 . Most evaluators had To analyze the evaluation data, two authors frst individually
some experience with discount usability testing methods: three examined the completed canvases for examples and comments
evaluators (16.6%) reported low or very low experience, fve (27.7%) that indicated the relevance, clarity, and coverage of the principles.
reported medium experience, and nine (50%) reported high or very They also reviewed participants’ ratings of relevance and clarity and
high experience. Given that generative AI design is an emerging delved into their examples and comments to understand instances
feld, our evaluators tended not to have high levels of experience in of lower ratings. They then converged with the other authors to
this area: eight evaluators (44.4%) reported very low to low expe- review their fndings and discuss potential ways to improve the
rience, eight (44.4%) reported medium experience, and one (5.5%) principles and strategies.
reported a high level of experience.
Evaluators self-selected an application familiar to them and com- 8.2.1 Relevance. To assess the relevance of the design principles
pleted their evaluation individually (as is standard practice [128]) and strategies to commercial generative AI applications, we counted
and remotely. As our evaluators were not involved in the design of the number of examples evaluators found. Evaluators identifed 286
these products, they were unable to evaluate the process-oriented total examples across all principles; they found a collective average
strategies (all strategies within Design Responsibly plus Evaluate of 11.9 examples for each strategy, and every strategy had at least
users’ mental models). Thus, these strategies were excluded from one example. The wealth of examples found suggests the principles
Iteration 3, and we made it a point to evaluate them in Iteration and strategies were relevant to a range of commercial generative
4 (Section 9). Participants recorded their evaluations of all other AI applications. Evaluators also generally rated each principle as
principles in a Mural19 template. Two evaluators examined each being relevant (Figure 4a).
application and each evaluation took approximately one hour. The relatively lower relevance ratings for Design for Appropri-
We crafted short descriptions20 for each principle and strategy ate Trust & Reliance, Design for Human Control, and Design
to orient our evaluators. For each principle, we asked evaluators for Optimization stemmed from diferences in application domain
to begin by capturing examples in the Mural canvas of how their and output modality. In some cases, we accepted that relevance
application applied the principle. At this stage, specifc strategies in may vary by use case; in other cases, we addressed issues raised by
the Mural were covered with an overlay to encourage evaluators to participants to clarify or expand relevance. For example, the four
fnd examples without being biased by our strategies, in hopes that evaluators who rated Design for Appropriate Trust & Reliance
they might identify new ones. After capturing examples, evaluators as “not relevant” had examined image or music generation applica-
were instructed to remove the overlay, then label each example tions and felt that overreliance was less of a concern for creative
with a strategy we provided, “not sure”, or a write-in for a new applications. In response to this observation, we added examples
strategy. After fnding and labeling examples, evaluators rated the of risks to be wary of in creative outputs (e.g. quality issues, bias,
relevance of each design principle and its strategies on a 4-point and underrepresentation [17, 18, 42]) to clarify its relevance to such
scale: “Yes, they were clearly relevant,” “Yes, they were relevant but applications. We made similar modifcations to strategies that were
I struggled to fnd examples,” “No, they were clearly not relevant,” too narrowly focused on specifc domains or output modalities.
and “Not sure.” They also rated the clarity of the principle as a
whole on a 5-point scale from “Very unclear” to “Very clear” and 8.2.2 Clarity. To assess the clarity of the design principles and
provided suggestions for improvement. Finally, after reviewing all strategies, we identifed instances where evaluators noted overlap
of the principles, evaluators were asked to identify any additional or redundancy between diferent principles or strategies, expressed
design features in their application that were not covered by the confusion, or interpreted a principle or strategy diferently from
principles. how we intended. We also asked evaluators to rate the clarity of
each principle (Figure 4b), and they were generally rated as being
clear.
17 Despite not having a design role, this evaluator worked on an HCI research team and Participants identifed eight overlap issues. Notably, fve evalua-
possessed over 20 years of professional design experience, which we felt sufciently tors found that nearly all strategies in Design for Exploration and
qualifed them to participate in this exercise. Design for Optimization overlapped in some way with strategies
18 Asone evaluator did not complete the demographic survey, percentages do not add in other principles. This observation led us to reconsider how to
up to 100%.
19 Mural is a collaborative, graphical canvas application. [Link] incorporate exploration and optimization within our framework.
20 These descriptions were the frst iteration of those shown in Appendix A. Ultimately, we recognized that exploration and optimization are
Describe when and why a user might use the product to create an "ideal"
output that meets some kind of criteria.

Skip this principle if your product does not support optimization use cases.

Be able to Change the Change


Be able to instrumentation
change the notes
"try again of a song or part
sections of after it's played in a
but similar"
a song generated given part
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

How does the product design for optimization?


Paste in screenshots of examples that you find.

You'll take two passes at this section - one open-ended pass, followed by
another pass after seeing the strategies in the next section.

First pass Second pass

Again?? I'm
not copying
Leverage multiple
my examples
outputs another time
Enable human-AI co- haha
creation

Evaluate outputs Don't really


Can take
using domain- feel like it
something you
specific metrics does this
already made
and "use as
influence"

Once in the
editor, you can re-
work the sections
it's divided the Can switch out
song up into the MIDI
instruments used
for playback of
each part
You can drag
individual notes in
the piano roll editor
(tho it's clunkier than
using a proper DAW
to move MIDI notes)

Figure 3: Portion of an evaluation of AIVA for the principle of Design for Optimization.
Consider the following strategies
The following strategies are examples of ways to design for optimization.

Drag and drop sticky notes to label your examples with relevant strategies.
If an example doesn't correspond to any of these strategies, use a blank sticky
note to come up with a new strategy OR label with "not sure".
Take another pass to look for strategies that don't have any examples. If you
don't find any, flag with

Leverage multiple outputs Evaluate outputs using domain-specific metrics Enable human-AI co-creation
Generate multiple outputs that are either hidden or visible Help the user find a generated artifact that satisfies Ensure the user can edit generated artifacts to
to the user in order to increase the chance that one of some objective criteria fix flaws and improve their quality
them fits the user’s need

Evaluate outputs
Leverage Evaluate outputs
Leveragemultiple
multiple Evaluate
Evaluate
using outputs
outputs
domain-
Enable human-AI co-
Enable human-AI co-
Leverage
outputs multiple using domain- creation
Enable human-AI co-
outputs using
using
specific domain-
domain-
metrics
outputs specific metrics creation
specific
specific metrics
metrics creation

New strategy Not sure


If you think an example falls under a strategy that If you're really not sure which strategy an example
isn't listed, write it in as a new strategy. falls under, label it with this sticky note.

Write in new strategy Not sure


Write
Writeininnew
newstrategy
strategy Not sure
Not
Not sure
sure

Drag this flag to any


strategies without examples
Figure 4: Evaluators’ ratings of the (a) relevance and (b) clarity of each principle and its strategies to their application in the
modifed heuristic evaluation.

Relevance
user goals rather than characteristics of a generative AI application, Clarity
We observed 16 instances in which an evaluator’s use of a strat-
andDidhence should be communicated as such (we discuss this point
you feel that this design principle and its strategies were relevant to the product?
egy label mismatched our intention for what the strategy repre-
How clear to you is the description of this principle?
further in Section 10.2). Other overlap issues were reconciled by
Please bold your selection. sented. We made two major changes in response to these mis-
Please bold your selection.

merging redundant
Yes, they were clearly relevant strategies. matches.
Very clear First, we reframed Design for Human Control as De-
Yes, they were relevant but I struggled to find examples Clear
sign for Co-Creation in response to frequent misinterpretations of
No, they were clearly not relevant Neutral
“controls” as afordances unrelated to the generative process (such
Not sure Unclear
as Photoshop’s editing tools). Design for Co-Creation provides
Very unclear

Reflect
Was there anything confusing about this principle or its strategies?
Do you have any feedback on how to improve them?

Multiple outputs again.

I think for something less subjective than music being generated, the one about evaluation using domain specific
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

greater specifcity to generative AI’s unique capabilities for human- Workshop 1: Evaluate existing LLM prompt tool
AI co-creation, which has been examined extensively within HCI
P1-1 Design Lead P1-3 Design Manager
communities (e.g., [34, 53, 77, 121, 129]). The second change was to
P1-2 Software Designer P1-4 UX Researcher
rename Design for Multiple Outputs to Design for Generative
Variability to better characterize its purpose after observing that Workshop 2: Ideate on a future LLM-based conversational tool
many evaluators narrowly interpreted this principle as solely being P2-1 UX Researcher P2-6 Design Research Lead
about the display of multiple outputs. We made additional wording P2-2 UX Researcher P2-7 Design Manager
changes and clarifcations to other principles and their associated P2-3 UX Researcher P2-8 UX Researcher
strategies in response to participants’ feedback. P2-4 Program Manager P2-9 Technical Program
8.2.3 Coverage. Evaluators found new examples that refected Lead
gaps in our framework, resulting in three new strategies: Teach the P2-5 UX Researcher P2-10 Design Research Lead
AI system about the user, Help the user craft efective outcome speci- Table 4: Participants in each of the workshops to assess the
fcations, and Support co-editing of generated outputs. We included actionability of the design principles.
Teach the AI system about the user in Design for Mental Models as
it addresses recent research in Mutual Theory of Mind [33, 177, 187].
We included Help the user craft efective outcome specifcations and
Support co-editing of generated outputs in Design for Co-Creation For each principle, participants were frst asked to identify ways
as they are most closely related to the co-creative process. they were already “designing for” or considering the principle. Next,
they identifed relevant strategies that they had not yet considered
9 ITERATION 4: APPLICATION TO and brainstormed ways to leverage them to improve their product.
GENERATIVE AI UX DESIGN This brainstorming session produced new design ideas that team
members shared and discussed with each other. At the end of the
Design guidelines can be difcult to put into practice [59, 113, 160,
session, participants refected on the actionability of the principles
164, 176, 192], often because they describe goals rather than ac-
within their design process. The moderators and note-takers of
tions [73]. Our strategies were meant to capture “actions” that
each workshop reviewed participants’ design ideas and recording
practitioners could take to apply the principles to their work. Af-
transcripts to identify insights on the usefulness of the principles
ter refning the principles and strategies for relevance and clarity,
and recommendations for improvement.
we evaluated their utility within the design process by conducting
structured, exploratory workshops with design practitioners within
9.2 Results
our organization who work on generative AI applications. Our pri-
mary goal was to understand how efectively the design principles 9.2.1 Applicability to practice. Participants in Workshop 1 brain-
could be applied in practice, but we remained open to identifying stormed a total of 46 design ideas and participants in Workshop 2
additional issues regarding relevance, clarity, and coverage. brainstormed 56 design ideas. These design ideas included feature
requirements, new afordances, design processes to try, and ques-
tions to consider when making design decisions. Groups generated
9.1 Method
between 5 and 14 ideas per principle, and participants generated
We held two workshops with two diferent teams (Table 4) to evalu- multiple, varied ideas for all of the principles. For example, when
ate the design principles and strategies in practice. Workshop 1 was considering the strategy, Help the user craft efective outcome speci-
held with an internal team comprised of four design practitioners fcations, P1-3 thought of an idea to “provide diferent ‘efects’ that
working on the IBM [Link] Prompt Lab21 , a prompt testing bake in some prompt content, (e.g. in the style of a famous au-
environment for large language models. Workshop 2 involved a thor),” and P1-4 proposed a system to “reward ‘best in class’ prompt
separate internal team of ten design practitioners in the early, for- authors and celebrate and share” their work.
mative stages of designing an internal LLM-based conversational When asked about the actionability of their brainstormed ideas,
tool that provides UX research support. We selected these two P1-1 responded, “I defnitely think a lot of these could go on a fu-
teams as they provided a broader view on the actionability of the ture roadmap.” P2-1 commented that the workshop, “made it clear
design principles in diferent phases of design: a later, evaluative crucial blind spots that could put the [application] idea at risk if
stage (Workshop 1) and an earlier, ideation phase (Workshop 2). not addressed,” and that it helped their team, “quickly generate new
Workshops took place remotely via video conferencing and were requirements.” Participants’ breadth of design ideas and comments
recorded with participants’ consent. Each session lasted 90 minutes on workshop outcomes indicate that practitioners are able to lever-
and included two moderators and two note-takers. One moderator age the principles and strategies to inform useful, actionable design
began each workshop by presenting an overview of the design prin- improvements.
ciples and strategies. To minimize the time required of participants, Participants also shared ideas on how to improve the actionabil-
we split each session into two groups and assigned three principles ity of the principles. As their understanding of the principles was
to each group. Each break-out group contained one moderator and limited to the brief overview provided at the start of the workshop,
one note-taker. they felt that having more details and resources to learn about the
principles and strategies would make them easier to apply. P1-3
[Link]. [Link] commented, “having some examples of these concepts out in the
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

world or in other tools might be a useful way to get a grasp of how to “draw the lines,” either between diferent principles and their
the concept works.” In support of this need, we include a library strategies when we identifed overlap or redundancy, or between
of examples in Appendix A that provide richer detail on how each what we later identifed as a diference between user goals and
strategy has been applied within existing applications. characteristics of generative AI. We also discuss our strategies for
Participants also shared insights on how the principles would putting the principles into action within our organization, as well
be incorporated into their design process. P1-4 commented that as limitations and opportunities for future work.
involving more roles in a workshop, such as developers and product
managers, would add value. P2-8 felt that Design for Imperfec- 10.1 Guideline organization
tion “can’t be applied unless user research is done” due to a lack of
Early in our frst iteration, we observed a hierarchical relation-
understanding of the user’s expectations for model outputs. P2-6
ship emerge between high-level design principles that identifed
also spoke about the value of user research since “the idea or solu-
unique or diferentiated aspects of generative AI and lower-level
tion... may look diferently depending on the user.” These comments
strategies for implementing those principles in a user experience.
indicate that a baseline understanding of users is needed to identify
However, the relationships between which strategies applied to
concrete design ideas from the principles and strategies, in line
which principles were not always clear, as some strategies could
with our recommendation to Use a human-centered approach.
be used to support multiple principles. For example, in Iteration
The outcomes from these workshops demonstrated that the de-
1, the strategy Visualizing diferences was included in Design for
sign principles can be applied in both an early ideation stage and a
Multiple Outputs, as it could help users understand the difer-
later evaluation stage to drive actionable design ideas.
ences amongst those outputs (especially for use cases where those
9.2.2 Relevance and clarity improvements. Workshop participants diferences might be subtle). But it was also included in Design
also provided feedback on the relevance and clarity of the de- for Imperfection, as it could help users more easily identify prob-
sign principles. This feedback primarily resulted in minor wording lematic outputs. Another example from Iteration 1 was the use
changes to the process-oriented strategies that we were unable to of a Sandbox / Playground Environment, which supported Design
evaluate in Iteration 3; no major organizational changes were made for Imperfection by not tainting an artifact-under-creation with
as a result of this feedback. After incorporating this feedback, we potentially-problematic generated content (e.g. source code with
produced our fnal set of design principles and strategies (Table 1). bugs or text containing factual errors). But it was also included in
Design for Exploration, as a sandbox provides a separate space
9.3 Final clarity evaluation for users to explore new candidates without interfering with their
To determine whether the wording changes we made after Itera- main working environment.
tions 3 & 4 impacted the clarity of the fnal design principles and As we worked through subsequent iterations, we wrestled with
strategies, we ran a follow-up survey with the evaluators from Iter- whether we should continue to allow strategies to overlap between
ation 3. Fourteen of 18 evaluators responded to our survey (77.7% principles or aim for a clean separation. During Iteration 2, with
response rate). the delineation between Design for Exploration and Design for
The clarity ratings collected in Iteration 3 were at the higher Optimization, even more redundancy was introduced as many
end of the 5-point scale (M (SD) = 4.29 (0.90) of 5). The changes we strategies support both kinds of uses. At this point, we even consid-
made in Iteration 4 made a small but signifcant improvement to ered completely decoupling the strategies from the principles and
clarity (M (SD) = 4.58 (0.58) of 5), � (1, 184) = 5.88, � = .02, ��2 = .03 providing an indication on each strategy (such as a tag) for which
(small). principle(s) it supported.
We ultimately decided to maintain the nesting of strategies
10 DISCUSSION within principles and aim for establishing clean boundaries. We
made this decision because the amount of overlap diminished when
We identifed a set of six principles important to the design of gen-
we refned the principles during Iteration 3 and separated out the
erative AI applications, along with a companion set of 24 strategies
user goals of optimization and exploration. Our evaluators also
for implementing those principles within a user experience. These
experienced frustration when they were unable to diferentiate be-
principles were developed iteratively using a combination of criti-
tween strategies, indicating a need to eliminate overlaps. However,
cal conceptual analyses (to ensure scientifc validity) and empirical
we note that a single UX feature may be used to implement more
work (to ensure real-world utility).
than one principle or strategy (see Appendix A for examples); there-
We collected formal feedback on the principles from a 18 design
fore, we only sought to reduce conceptual overlaps between the
practitioners who collectively evaluated them against 9 commercial
principles and strategies themselves, as opposed to overlaps when
applications. We then collected feedback from 12 design practition-
a specifc UX capability addresses multiple principles or strategies.
ers on two design teams who applied them to both the formative
and evaluative stages of product design. We found that the prin-
ciples helped design practitioners generate useful and actionable 10.2 User goals versus design principles
design improvements and were applicable to a range of generative In reviewing the Library of Mixed Initiative Creative Interfaces [161],
AI applications, including those that generate diferent types of we realized that generative capabilities are sometimes an end in
media (e.g. text, images, music). themselves, but other times are a means to achieving another goal.
We discuss two issues that kept surfacing throughout our de- We identifed these two diferent purposes of use as optimization
velopment process that required us to think deeply about where and exploration, respectively:
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

• In optimization use cases, the process of generating artifacts We conclude that the principles and strategies form a toolbox
is an end: users use the generative capability to produce one that design practitioners can use holistically or selectively as they
or more artifacts that satisfy their needs, such as a source craft user experiences for generative AI applications. Design prac-
code function that implements a desired operation, a mol- titioners know their users and their needs best – exemplifed by
ecule that possesses specifc properties, or an image that Use a human-centered approach – and it is our hope that we have
depicts a desired scene or character. We labeled this class of provided useful vocabulary for them to understand and design for
usage as “optimization” in recognition that the generative the new and diferent kinds of uses that generative AI systems ofer.
AI model may not produce a fawless or “perfect” output,
and some amount of refnement (either by the user or the 10.3 Adoption within our organization
AI) may be required before it is satisfactory. As discussed in Section 2.1, the HCI community has produced a
• In exploratory use cases, the process of generating artifacts prodigious number of design guidelines throughout its history. But,
is a means to an end: the purpose is not to generate the as noted by both Soni et al. [160] and Stark et al. [164], our commu-
artifact, but to use the generated artifacts in order to learn nity struggles with bridging the gap between the development of
about a domain (e.g. programming [100] or medicine [89]) scientifcally-grounded guidelines and real-world design practice.
or be inspired by seeing new or diferent possibilities (e.g. We developed our design principles specifcally to provide practi-
brainstorming [145] or pre-writing [175]). Few other types of cal and actionable support to design practitioners. Therefore, we
AI technology support this kind of usage, where the emphasis undertook a number of eforts to promote their adoption within
is on assisting people in conducting a thought process. our organization.
(1) Actionable activities. To bridge the gap between theory
During Iteration 3, we ultimately decided to remove Design
and practice, we developed activities for designers to apply
for Exploration and Design for Optimization as core design
the principles and strategies to their own work. Chief among
principles because of the strong degree of overlap between their
them is a heuristic evaluation that uses the principles as
strategies and the strategies of other principles. In fact, we could
heuristics for designers to evaluate the user experience of
not even clearly delineate diferent generative AI applications as
generative AI applications. We created and disseminated a
supporting exploratory versus optimization usage, because many
self-contained Mural template that guides designers through
applications supported both, and users might even alternate be-
this evaluation to identify new ideas and opportunities for
tween the two kinds of usage when using the application. For
design improvement. We also developed workshop activities
example, work by Weisz et al. [184] shows how software engineers
for identifying applications of generative AI that drive user
used generative technologies not only to produce source code trans-
value and evaluating a user’s mental model of an AI system.
lations (optimization), but also to improve their own knowledge of
(2) Progressive detail. When we initially developed the prin-
programming (exploration), within the same overall task context.
ciples and strategies, we wrote about them extensively in a
Hence, we drew a line between user goals and design principles
comprehensive guide that provided foundational knowledge
(depicted in Figure 1).
and case studies on generative AI, which we shared with
We assert that each principle broadly supports both user goals,
our internal design community. We received feedback that
but the extent of their support does difer by goal. Design for Im-
the level of detail was informative but too lengthy for busy
perfection is strongly aligned with optimization use cases, as the
designers. In response, we developed two condensed presen-
reason why optimization is even necessary is because of the imper-
tations: 1) paragraph-length descriptions for each principle
fect outputs produced by generative models. Concurrently, Design
and strategy (shown in Appendix A) which were included in
for Generative Variability is strongly aligned with exploration
the generative AI heuristic evaluation template, and 2) one-
use cases, as generative variability is a key enabler of exploration.
sentence descriptions of each principle and strategy (shown
However, we note that Design for Imperfection can also sup-
in Table 1) which were published on an internal website for
port exploratory use by embracing unexpected “imperfections” that
the design of AI applications.
arise from discrepancies between a user’s intent and the model’s
(3) Hands-on outreach. We conducted outreach activities to
output. In addition, Design for Generative Variability can also
raise awareness of the principles within our organization.
support optimization use cases by helping users narrow down on
Some of these eforts targeted a general design audience,
an option that fts their needs from a wide feld.
such as creating a discussion group22 for generative AI de-
Another principle that has a high degree of afnity to optimiza-
sign and presenting the principles at internal seminars. Other
tion is Design for Appropriate Trust & Reliance, especially
outreach targeted designers on key product teams. As one
when generated outputs are used within high-stakes domains (e.g.
example, we held a workshop attended by 62 people at an
code, customer service). Trust and reliance may be of a lesser con-
internal design event to teach designers how to conduct a
cern for exploration use cases, although users should still be wary
heuristic evaluation of generative AI applications. Instead
of bias, underrepresentation, and other harms that may occur.
of using a sample application, we evaluated a product re-
Finally, Design for Co-Creation can be applied to both explo-
cently released by our organization and invited the product’s
ration and optimization tasks, but the strategies of Help the user
design team to participate. In an hour-long session, partici-
craft efective outcome specifcations and Support co-editing of gener-
pants identifed 10 usability issues and 6 new feature ideas,
ated outputs may be more important when users need to optimize
generated outputs to ft certain criteria. 22 This group grew to over 1,200 members over the course of 9 months.
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

which we discussed in detail with the product team in follow- generative model’s inference latency. In addition, organizational
up meetings. They reported our fndings to be useful and policies may be created that govern the uses of generative AI. We
included several recommendations in their roadmap. believe there is room for expansion to identify how designers can
(4) Executive sponsorship. In addition to bottom-up dissemi- participate in these kinds of technical and policy decisions that
nation, we also worked with key executives in our design or- have an impact on the user experience.
ganization to encourage relevant product teams to adopt the
principles (as recommended by Madaio et al. [110]). Through
11 CONCLUSION
this efort, we introduced the principles to 10 product teams
who were in the process of learning about generative AI We developed a set of six principles for the design of applications
and identifying opportunities for incorporating it into their that incorporate generative AI technologies. Three principles – De-
product. sign Responsibly, Design for Mental Models, and Design for
Appropriate Trust & Reliance – ofer new interpretations of
Akin to Yildirim et al.’s observations on how their guidebook known issues with the design of AI systems when viewed through
improved AI literacy within their organization and helped designers the lens of generative AI. Three principles – Design for Gen-
establish credibility and advocate for user needs [192], we found erative Variability, Design for Co-Creation, and Design for
our materials had a similar impact. The executive sponsorship of Imperfection – identify issues that are unique to generative AI
our work and the adoption of the principles by numerous product applications. Each principle is coupled with a set of strategies for
teams speak not just to their practical utility, but also for the great how to implement it within a user experience, either through the
need to equip design practitioners and enable them to “have a seat inclusion of specifc types of UX features or by following a specifc
at the table” in the creation of generative AI applications. design process. We developed the principles and strategies using
an iterative process that involved reviewing relevant literature
10.4 Limitations and future work in human-AI collaboration and co-creation, collecting feedback
The feld of generative AI is undergoing rapid innovation, both in from design practitioners, and validating the principles against
the pace of technological development and in how those technolo- real-world generative AI applications. We also demonstrated the
gies are being brought to the market. We view our principles as value and applicability of the principles by applying them in the
beginning a discussion on how to design efective and safe gen- design process of two generative AI applications. As generative
erative AI applications. As the pace of innovation continues and AI technologies are rapidly being incorporated into existing appli-
new generative AI applications are developed, we anticipate new cations, and entirely new products are being created with these
challenges to be uncovered, necessitating new sets of guidelines, technologies, we see signifcant value in principles that aid design
tools, best practices, design patterns, and evaluative methods. practitioners in harnessing these technologies for the beneft of
One challenge we encountered with the modifed heuristic eval- their users in safe and efective ways.
uation was in its use to evaluate the design principles themselves
through the process of evaluating a generative AI application. Not ACKNOWLEDGMENTS
all of our evaluators understood this distinction, and as a result, we We thank everyone at IBM who provided valuable feedback and
sometimes received feedback about shortcomings of the applica- guidance in the development of these design principles.
tions that was less relevant to our goal of improving the principles.
We recommend providing stronger introductory examples that
focus on how they help evaluate the principles rather than the REFERENCES
products. [1] Adobe. 2023. Adobe Releases New Firefy Generative AI Models and
Web App; Integrates Firefy Into Creative Cloud and Adobe Express.
Another limitation of the modifed heuristic evaluation was our [Link]
focus on evaluating commercially-available generative AI applica- Firefy-Generative-AI-Models-and-Web-App-Integrates-Firefy-Into-
tions. There are also many experimental applications in this space, Creative-Cloud-and-Adobe-Express/[Link].
[2] Rukshan Alexander, David Murray, and Nik Thompson. 2017. Cross-cultural web
but we did not examine them. Our restriction to commercial ap- design guidelines. In Proceedings of the 14th International Web for All Conference.
plications excluded other ways of interacting with generative AI 1–4.
[3] Radwan Ali, Mike Gallivan, and Seema Sangari. 2019. A Study of mobile apps in
applications, such as through narrative [24], lyric and other poetic the banking Industry. International Journal of Digital Society (IJDS) 10, 3 (2019),
forms [147], and movement [174]. 1524–1533.
Finally, our design guidelines are entirely focused on helping [4] Reem Alnanih and Olga Ormandjieva. 2016. Mapping HCI principles to design
quality of mobile user interfaces in healthcare applications. Procedia Computer
design practitioners to develop the user experience for a generative Science 94 (2016), 75–82.
AI application. But, UX design is only one portion of the AI develop- [5] A Alsswey, IN Umar, and H Al-Samarraie. 2018. Towards mobile design
ment lifecycle, which includes other phases such as model selection, guidelines-based cultural values for elderly Arabic users. Journal of Funda-
mental and Applied Sciences 10, 2S (2018), 964–977.
model tuning, prompt engineering, deployment & monitoring, and [6] Oscar Alvarado and Annika Waern. 2018. Towards algorithmic experience:
more. As decisions made during those phases will ultimately impact Initial eforts for social media contexts. In Proceedings of the 2018 chi conference
on human factors in computing systems. 1–12.
the user experience, we believe design practitioners ought to have [7] Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira
their inputs considered. However, the design principles do not cur- Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen,
rently help them understand, for example, how to determine which et al. 2019. Guidelines for human-AI interaction. In Proceedings of the 2019 chi
conference on human factors in computing systems. 1–13.
generative model should be used to implement a Q&A use case and [8] Apple Computer, Inc. 1987. Apple human interface guidelines: The Apple desktop
when that model’s performance is “good enough,” or how to hide a interface. Addison Wesley Publishing Company.
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

[9] Apple, Inc. 2022. Human Interface Guidelines for Machine Learn- [30] Kulsiri Chirayus and Aziz Nanthaamornphong. 2020. Cognitive mobile design
ing. Retrieved 14-Aug-2023 from [Link] guidelines for the elderly: a preliminary study. In 2020 17th International Con-
interface-guidelines/technologies/machine-learning/introduction ference on Electrical Engineering/Electronics, Computer, Telecommunications and
[10] Maryam Ashoori and Justin D Weisz. 2019. In AI we trust? Factors that infu- Information Technology (ECTI-CON). IEEE, 673–678.
ence trustworthiness of AI-infused decision-making processes. arXiv preprint [31] Yaliang Chuang, Lin-Lin Chen, and Yoga Liu. 2018. Design vocabulary for
arXiv:1912.02675 (2019). human–IoT systems communication. In Proceedings of the 2018 CHI Conference
[11] Nagadivya Balasubramaniam, Marjo Kauppinen, Sari Kujala, and Kari Hiekka- on Human Factors in Computing Systems. 1–11.
nen. 2020. Ethical guidelines for solving ethical issues and developing AI [32] John Joon Young Chung, Minsuk Chang, and Eytan Adar. 2021. Gestural Inputs
systems. In Product-Focused Software Process Improvement: 21st International as Control Interaction for Generative Human-AI Co-Creation. In Workshops at
Conference, PROFES 2020, Turin, Italy, November 25–27, 2020, Proceedings 21. the International Conference on Intelligent User Interfaces (IUI).
Springer, 331–346. [33] Fabio Cuzzolin, Alice Morelli, Bogdan Cirstea, and Barbara J Sahakian. 2020.
[12] Shaowen Bardzell. 2010. Feminist HCI: taking stock and outlining an agenda for Knowing me, knowing you: theory of mind in AI. Psychological medicine 50, 7
design. In Proceedings of the SIGCHI conference on human factors in computing (2020), 1057–1061.
systems. 1301–1310. [34] Nicholas Davis, Chih-PIn Hsiao, Kunwar Yashraj Singh, Lisa Li, and Brian
[13] Kurtis Beavers, Shawndell Elfring, Hannah Reed, and Rachel Shepard. 2023. Magerko. 2016. Empirically studying participatory sense-making in abstract
UX: Designing for Copilot | DIS214H. Microsoft Developer. [Link] drawing with a co-creative cognitive agent. In Proceedings of the 21st Interna-
WiCVEMH4HTI tional Conference on Intelligent User Interfaces. 196–207.
[14] Shirley Ann Becker. 2004. A study of web usability for older adults seeking [35] Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu
online health resources. ACM Transactions on Computer-Human Interaction Wang, Tianwei Zhang, and Yang Liu. 2023. Jailbreaker: Automated Jailbreak
(TOCHI) 11, 4 (2004), 387–406. Across Multiple Large Language Model Chatbots. arXiv preprint arXiv:2307.08715
[15] Kirsty A Beilharz, Joanne Jakovich, and Sam Ferguson. 2006. Hyper- (2023).
shaku (Border-crossing): Towards the Multi-modal Gesture-controlled Hyper- [36] Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh,
Instrument.. In NIME. 352–357. Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 2022. Exploring how
[16] Nigel Bevan. 2005. Guidelines and standards for web usability. In Proceedings of machine learning practitioners (try to) use fairness toolkits. In Proceedings of the
HCI International, Vol. 2005. Citeseer, 10. 2022 ACM Conference on Fairness, Accountability, and Transparency. 473–484.
[17] Charlotte Bird, Eddie Ungless, and Atoosa Kasirzadeh. 2023. Typology of Risks [37] Sebastian Deterding, Jonathan Hook, Rebecca Fiebrink, Marco Gillies, Jeremy
of Generative Text-to-Image Models. In Proceedings of the 2023 AAAI/ACM Gow, Memo Akten, Gillian Smith, Antonios Liapis, and Kate Compton. 2017.
Conference on AI, Ethics, and Society. 396–410. Mixed-initiative creative interfaces. In Proceedings of the 2017 CHI Conference
[18] IBM AI Ethics Board. 2023. Foundation models: Opportunities, risks and mit- Extended Abstracts on Human Factors in Computing Systems. 628–635.
igations. Retrieved 05-Sep-2023 from [Link] [38] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert:
E5KE5KRZ Pre-training of deep bidirectional transformers for language understanding.
[19] Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, arXiv preprint arXiv:1810.04805 (2018).
Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma [39] DW Dodson and NL Shields Jr. 1978. Development of user guidelines for
Brunskill, et al. 2021. On the opportunities and risks of foundation models. ECAS display design, volume 1. (1978). [Link]
arXiv preprint arXiv:2108.07258 (2021). 19790007504/downloads/[Link]
[20] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, [40] Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu
Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey for in-context learning. arXiv
Askell, et al. 2020. Language models are few-shot learners. Advances in neural preprint arXiv:2301.00234 (2022).
information processing systems 33 (2020), 1877–1901. [41] Graham Dove, Kim Halskov, Jodi Forlizzi, and John Zimmerman. 2017. UX
[21] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Ka- design innovation: Challenges for working with machine learning as a design
plan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, material. In Proceedings of the 2017 chi conference on human factors in computing
Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, systems. 278–288.
Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jefrey Wu, [42] Elizabeth Edenberg and Alexandra Wood. 2023. Disambiguating Algorithmic
Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Bias: From Neutrality to Justice. In Proceedings of the 2023 AAAI/ACM Conference
Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec on AI, Ethics, and Society. 691–704.
Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are [43] Upol Ehsan, Q Vera Liao, Michael Muller, Mark O Riedl, and Justin D Weisz.
Few-Shot Learners. In Advances in Neural Information Processing Systems, 2021. Expanding explainability: Towards social transparency in ai systems. In
H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems.
Curran Associates, Inc., 1877–1901. [Link] 1–19.
fle/[Link] [44] Upol Ehsan, Philipp Wintersberger, Q Vera Liao, Elizabeth Anne Watkins, Carina
[22] Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to Manger, Hal Daumé III, Andreas Riener, and Mark O Riedl. 2022. Human-
think: cognitive forcing functions can reduce overreliance on AI in AI-assisted Centered Explainable AI (HCXAI): beyond opening the black-box of AI. In CHI
decision-making. Proceedings of the ACM on Human-Computer Interaction 5, conference on human factors in computing systems extended abstracts. 1–7.
CSCW1 (2021), 1–21. [45] Virginia K Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May.
[23] Paul Cairns, Christopher Power, Mark Barlet, and Greg Haynes. 2019. Future 2023. WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+
design of accessibility in games: A design vocabulary. International Journal of Bias in Large Language Models. arXiv preprint arXiv:2306.15087 (2023).
Human-Computer Studies 131 (2019), 64–71. [46] KJ Kevin Feng, Maxwell James Coppock, and David W McDonald. 2023. How Do
[24] Alex Calderwood, Vivian Qiu, Katy Ilonka Gero, and Lydia B Chilton. 2020. UX Practitioners Communicate AI as a Design Material? Artifacts, Conceptions,
How Novelists Use Generative Language Models: An Exploratory User Study.. and Propositions. In Proceedings of the 2023 ACM Designing Interactive Systems
In HAI-GEN+ user2agent@ IUI. Conference. 2263–2280.
[25] James J Cappel and Zhenyu Huang. 2007. A usability analysis of company [47] Giorgio Franceschelli and Mirco Musolesi. 2022. Copyright in generative deep
websites. Journal of Computer Information Systems 48, 1 (2007), 117–123. learning. Data & Policy 4 (2022), e17.
[26] PV Charan, Hrushikesh Chunduri, P Mohan Anand, and Sandeep K Shukla. 2023. [48] Batya Friedman. 1996. Value-sensitive design. interactions 3, 6 (1996), 16–23.
From Text to MITRE Techniques: Exploring the Malicious Use of Large Language [49] Helena Anna Frijns and Christina Schmidbauer. 2021. Design Guidelines for
Models for Generating Cyber Attack Payloads. arXiv preprint arXiv:2305.15336 Collaborative Industrial Robot User Interfaces. In Human-Computer Interaction–
(2023). INTERACT 2021: 18th IFIP TC 13 International Conference, Bari, Italy, August
[27] Jiaao Chen, Aston Zhang, Xingjian Shi, Mu Li, Alex Smola, and Diyi Yang. 2023. 30–September 3, 2021, Proceedings, Part III 18. Springer, 407–427.
Parameter-Efcient Fine-Tuning Design Spaces. arXiv preprint arXiv:2301.01821 [50] Sakiko Fukuda-Parr and Elizabeth Gibbons. 2021. Emerging consensus on
(2023). ‘ethical AI’: Human rights critique of stakeholder guidelines. Global Policy 12
[28] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde (2021), 32–44.
de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, [51] Noa Garcia, Yusuke Hirota, Yankun Wu, and Yuta Nakashima. 2023. Uncurated
Greg Brockman, et al. 2021. Evaluating large language models trained on code. image-text datasets: Shedding light on demographic bias. In Proceedings of the
arXiv preprint arXiv:2107.03374 (2021). IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6957–6966.
[29] Vijil Chenthamarakshan, Payel Das, Samuel Hofman, Hendrik Strobelt, Inkit [52] Werner Geyer, Lydia B Chilton, Justin D Weisz, and Mary Lou Maher. 2021. HAI-
Padhi, Kar Wai Lim, Benjamin Hoover, Matteo Manica, Jannis Born, Teodoro GEN 2021: 2nd Workshop on Human-AI Co-Creation with Generative Models.
Laino, et al. 2020. CogMol: Target-specifc and selective drug design for COVID- In 26th International Conference on Intelligent User Interfaces-Companion. 15–17.
19 using deep generative models. Advances in Neural Information Processing [53] Vlad Petre Glăveanu. 2014. Distributed creativity: What is it? Springer.
Systems 33 (2020), 4320–4332.
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

[54] Frederic Gmeiner, Humphrey Yang, Lining Yao, Kenneth Holstein, and Nikolas [78] Amy K Karlson, Shamsi T Iqbal, Brian Meyers, Gonzalo Ramos, Kathy Lee, and
Martelaro. 2023. Exploring Challenges and Opportunities to Support Designers John C Tang. 2010. Mobile taskfow in context: a screenshot study of smartphone
in Learning to Co-create with AI-based Manufacturing Design Tools. In Pro- usage. In Proceedings of the SIGCHI Conference on Human Factors in Computing
ceedings of the 2023 CHI Conference on Human Factors in Computing Systems. Systems. 2009–2018.
1–20. [79] Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator ar-
[55] Chao Gong, Yue Qiu, and Bin Zhao. 2018. Establishment of Design Strategies chitecture for generative adversarial networks. In Proceedings of the IEEE/CVF
and Design Models of Human Computer Interaction Interface Based on User conference on computer vision and pattern recognition. 4401–4410.
Experience. In Design, User Experience, and Usability: Theory and Practice: 7th [80] Markelle Kelly, Aakriti Kumar, Padhraic Smyth, and Mark Steyvers. 2023. Cap-
International Conference, DUXU 2018, Held as Part of HCI International 2018, Las turing Humans’ Mental Models of AI: An Item Response Theory Approach. In
Vegas, NV, USA, July 15-20, 2018, Proceedings, Part I 7. Springer, 60–76. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Trans-
[56] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, parency. 1723–1734.
Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial [81] Kamran Khowaja and Siti Salwah Salim. 2020. A framework to design
nets. Advances in neural information processing systems 27 (2014). vocabulary-based serious games for children with autism spectrum disorder
[57] Imke Grabe, Miguel González-Duque, Sebastian Risi, and Jichen Zhu. 2022. (ASD). Universal Access in the Information Society 19, 4 (2020), 739–781.
Towards a Framework for Human-AI Interaction Patterns in Co-Creative GAN [82] Huhn Kim. 2010. Efective organization of design guidelines refecting designer’s
Applications. Joint Proceedings of the ACM IUI Workshops 2022, March 2022, design strategies. International Journal of Industrial Ergonomics 40, 6 (2010),
Helsinki, Finland (2022). 669–688.
[58] Jonathan Grudin. 1990. The computer reaches out: The historical continuity of [83] Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and
interface design. In Proceedings of the SIGCHI conference on Human factors in Seong Joon Oh. 2023. Propile: Probing privacy leakage in large language models.
computing systems. 261–268. arXiv preprint arXiv:2307.01881 (2023).
[59] Thilo Hagendorf. 2020. The ethics of AI ethics: An evaluation of guidelines. [84] Taenyun Kim, Maria D Molina, Minjin Rheu, Emily S Zhan, and Wei Peng. 2023.
Minds and machines 30, 1 (2020), 99–120. One AI Does Not Fit All: A Cluster Analysis of the Laypeople’s Perception of AI
[60] Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, Roles. In Proceedings of the 2023 CHI Conference on Human Factors in Computing
and Ece Kamar. 2022. Toxigen: A large-scale machine-generated dataset for Systems. 1–20.
adversarial and implicit hate speech detection. arXiv preprint arXiv:2203.09509 [85] Tae Soo Kim, Arghya Sarkar, Yoonjoo Lee, Minsuk Chang, and Juho Kim. 2023.
(2022). LMCanvas: Object-Oriented Interaction to Personalize Large Language Model-
[61] Hartmut Hoehle, Ruba Aljafari, and Viswanath Venkatesh. 2016. Leveraging Powered Writing Environments. arXiv preprint arXiv:2303.15125 (2023).
Microsoft’s mobile usability guidelines: Conceptualizing and developing scales [86] Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes.
for mobile application usability. International Journal of Human-Computer arXiv preprint arXiv:1312.6114 (2013).
Studies 89 (2016), 35–53. [87] Marion Koelle, Swamy Ananthanarayan, and Susanne Boll. 2020. Social ac-
[62] Lars Erik Holmquist. 2017. Intelligence on tap: artifcial intelligence as a new ceptability in HCI: A survey of methods, measures, and design strategies. In
design material. interactions 24, 4 (2017), 28–33. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems.
[63] Stephanie Houde, Vera Liao, Jacquelyn Martino, Michael Muller, David Pi- 1–19.
orkowski, John Richards, Justin Weisz, and Yunfeng Zhang. 2020. Business [88] Max Kreminski, Isaac Karth, Michael Mateas, and Noah Wardrip-Fruin. 2022.
(mis) use cases of generative ai. arXiv preprint arXiv:2003.07679 (2020). Evaluating Mixed-Initiative Creative Interfaces via Expressive Range Coverage
[64] Johannes Huebner, Remo Manuel Frey, Christian Ammendola, Elgar Fleisch, Analysis.. In IUI Workshops. 34–45.
and Alexander Ilic. 2018. What people like in mobile fnance apps: An analysis [89] Tifany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lo-
of user reviews. In Proceedings of the 17th international conference on mobile and rie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-
ubiquitous multimedia. 293–304. Candido, James Maningo, et al. 2023. Performance of ChatGPT on USMLE:
[65] IBM. 2022. Design for AI. Retrieved 14-Aug-2023 from [Link] Potential for AI-assisted medical education using large language models. PLoS
design/ai/ digital health 2, 2 (2023), e0000198.
[66] IBM. 2022. What is AI ethics? Retrieved 14-Aug-2023 from [Link] [90] Sri Kurniawan and Panayiotis Zaphiris. 2005. derived web design guidelines for
com/topics/ai-ethics older people. In Proceedings of the 7th international ACM SIGACCESS conference
[67] Nanna Inie, Jeanette Falk, and Steve Tanimoto. 2023. Designing Participatory on Computers and accessibility. 129–135.
AI: Creative Professionals’ Worries and Expectations about Generative AI. In [91] Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat,
Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard:
Systems. 1–8. Scaling giant models with conditional computation and automatic sharding.
[68] Mikhail Jacob, Alexander Zook, and Brian Magerko. 2013. Viewpoints AI: arXiv preprint arXiv:2006.16668 (2020).
Procedurally Representing and Reasoning about Gestures.. In DiGRA conference. [92] Barbara Leporini and Fabio Paternò. 2008. Applying web usability criteria for
[69] Maurice Jakesch, Zana Buçinca, Saleema Amershi, and Alexandra Olteanu. 2022. vision-impaired users: does it really improve task performance? Intl. Journal of
How diferent groups prioritize ethical values for responsible AI. In Proceedings Human–Computer Interaction 24, 1 (2008), 17–47.
of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 310– [93] Clayton Lewis and Cathleen Wharton. 1997. Cognitive walkthroughs. In
323. Handbook of human-computer interaction. Elsevier, 717–732.
[70] Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang [94] Tianshi Li, Kayla Reiman, Yuvraj Agarwal, Lorrie Faith Cranor, and Jason I
Sun, Yizhou Wang, and Yaodong Yang. 2023. BeaverTails: Towards Improved Hong. 2022. Understanding challenges for developers to create accurate privacy
Safety Alignment of LLM via a Human-Preference Dataset. arXiv preprint nutrition labels. In Proceedings of the 2022 CHI Conference on Human Factors in
arXiv:2307.04657 (2023). Computing Systems. 1–24.
[71] Jiyou Jia and Bilan Zhang. 2018. Design guidelines for mobile MOOC learn- [95] Tianyi Li, Mihaela Vorvoreanu, Derek DeBellis, and Saleema Amershi. 2022.
ing—an empirical study. In Blended Learning. Enhancing Learning Success: 11th Assessing Human-AI Interaction Early through Factorial Surveys: A Study on the
International Conference, ICBL 2018, Osaka, Japan, July 31-August 2, 2018, Pro- Guidelines for Human-AI Interaction. ACM Transactions on Computer-Human
ceedings 11. Springer, 347–356. Interaction (2022).
[72] Anna Jobin, Marcello Ienca, and Efy Vayena. 2019. The global landscape of AI [96] Yucheng Li, Deyuan Chen, Tianshi Li, Yuvraj Agarwal, Lorrie Faith Cranor, and
ethics guidelines. Nature machine intelligence 1, 9 (2019), 389–399. Jason I Hong. 2022. Understanding iOS privacy nutrition labels: An exploratory
[73] Jef Johnson. 2020. Designing with the mind in mind: simple guide to understanding large-scale analysis of app store data. In CHI Conference on Human Factors in
user interface design guidelines. Morgan Kaufmann. Computing Systems Extended Abstracts. 1–7.
[74] Tristan E Johnson, Youngmin Lee, Miyoung Lee, Debra L O’Connor, Mo- [97] Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi-
hammed K Khalil, and Xiaoxia Huang. 2007. Measuring sharedness of team- hiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al.
related knowledge: Design and validation of a shared mental model instrument. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110
Human Resource Development International 10, 4 (2007), 437–454. (2022).
[75] Nick Jones and Matthew Moftt. 2016. Ethical guidelines for mobile app devel- [98] Q Vera Liao, Daniel Gruen, and Sarah Miller. 2020. Questioning the AI: informing
opment within health and mental health felds. Professional Psychology: Research design practices for explainable AI user experiences. In Proceedings of the 2020
and Practice 47, 2 (2016), 155. CHI Conference on Human Factors in Computing Systems. 1–15.
[76] Benjamin Kaiser, Akos Csiszar, and Alexander Verl. 2018. Generative models [99] Joseph Carl Robnett Licklider and Welden E Clark. 1962. On-line man-computer
for direct generation of cnc toolpaths. In 2018 25th International Conference on communication. In Proceedings of the May 1-3, 1962, spring joint computer con-
Mechatronics and Machine Vision in Practice (M2VIP). IEEE, 1–6. ference. 113–128.
[77] Anna Kantosalo and Tapio Takala. 2020. Five C’s for Human-Computer Co- [100] Mark Lifton, Brad Sheese, Jaromir Savelka, and Paul Denny. 2023. CodeHelp:
Creativity-An Update on Classical Creativity Perspectives.. In ICCC. 17–24. Using Large Language Models with Guardrails for Scalable Support in Program-
ming Classes. arXiv preprint arXiv:2308.06921 (2023).
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

[101] Weng Marc Lim, Asanka Gunasekara, Jessica Leigh Pallant, Jason Ian Pallant, [125] Jakob Nielsen. 1994. Enhancing the explanatory power of usability heuristics.
and Ekaterina Pechenkina. 2023. Generative AI and the future of education: Rag- In Proceedings of the SIGCHI conference on Human Factors in Computing Systems.
narök or reformation? A paradoxical perspective from management educators. 152–158.
The International Journal of Management Education 21, 2 (2023), 100790. [126] Jakob Nielsen. 1995. Scenarios in discount usability engineering. In Scenario-
[102] Vivian Liu. 2023. Beyond Text-to-Image: Multimodal Prompts to Explore Gen- Based Design: Envisioning work and technology in system development. 59–83.
erative AI. In Extended Abstracts of the 2023 CHI Conference on Human Factors [127] Jakob Nielsen. 2023. AI: First New UI Paradigm in 60 Years. Nielsen Norman
in Computing Systems. 1–6. Group (18 06 2023). [Link]
[103] Vivian Liu and Lydia B Chilton. 2021. Neurosymbolic Generation of 3D Animal [128] Jakob Nielsen and Rolf Molich. 1990. Heuristic evaluation of user interfaces. In
Shapes through Semantic Controls.. In IUI Workshops. Proceedings of the SIGCHI conference on Human factors in computing systems.
[104] Vivian Liu and Lydia B Chilton. 2022. Design guidelines for prompt engineering 249–256.
text-to-image generative models. In Proceedings of the 2022 CHI Conference on [129] Hugo Gonçalo Oliveira, Raquel Hervás, Alberto Díaz, and Pablo Gervás. 2014.
Human Factors in Computing Systems. 1–23. Adapting a Generic Platform for Poetry Generation to Produce Spanish Poems..
[105] Ryan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry, and Carrie J Cai. In ICCC. 63–71.
2020. Novice-AI music co-creation via AI-steering tools for deep generative [130] OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774 [[Link]]
models. In Proceedings of the 2020 CHI Conference on Human Factors in Computing [131] Jonas Oppenlaender. 2022. A taxonomy of prompt modifers for text-to-image
Systems. 1–13. generation. arXiv preprint arXiv:2204.13988 2 (2022).
[106] Todd Lubart. 2005. How can computers be partners in the creative process: [132] Long Ouyang, Jefrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela
classifcation and commentary on the special issue. International Journal of Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022.
Human-Computer Studies 63, 4-5 (2005), 365–369. Training language models to follow instructions with human feedback. Advances
[107] Nicola Lucchi. 2023. ChatGPT: A Case Study on Copyright Challenges for in Neural Information Processing Systems 35 (2022), 27730–27744.
Generative Artifcial Intelligence Systems. European Journal of Risk Regulation [133] Google PAIR. 2021. People+AI Guidebook. Google. [Link]
(2023), 1–23. guidebook/
[108] Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine [134] Kyudong Park, Taedong Goh, and Hyo-Jeong So. 2014. Toward accessible mobile
Jernite. 2023. Stable bias: Analyzing societal representations in difusion models. application design: developing mobile application accessibility guidelines for
arXiv preprint arXiv:2303.11408 (2023). people with visual impairment. HCI Korea (2014), 31–38.
[109] I Lupanda and JT Janse van Rensburg. 2021. Design Guidelines for Mobile Ap- [135] Alejandro Perez, Iaroslav Elistratov, Fynn Schmitt-Ulms, Ege Demir, Sadhana
plications. In International Conference Interfaces & Human Computer Interaction. Lolla, Elaheh Ahmadi, and Alexander Amini. 2023. Risk-Aware Image Genera-
92–99. tion by Estimating and Propagating Uncertainty. (2023).
[110] Michael A Madaio, Luke Stark, Jennifer Wortman Vaughan, and Hanna Wallach. [136] Aryo Pinandito, Hanifah Muslimah Az-zahra, Lutf Fanani, and Anggi Valeria
2020. Co-designing checklists to understand organizational challenges and Putri. 2017. Analysis of web content delivery efectiveness and efciency in
opportunities around fairness in AI. In Proceedings of the 2020 CHI conference on responsive web design using material design guidelines and User Centered
human factors in computing systems. 1–14. Design. In 2017 International Conference on Sustainable Information Engineering
[111] Mary Lou Maher. 2012. Computational and collective creativity: Who’s being and Technology (SIET). IEEE, 435–441.
creative?. In ICCC. Citeseer, 67–71. [137] Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018.
[112] Mary Lou Maher, Justin D Weisz, Lydia B Chilton, Werner Geyer, and Hen- Improving language understanding by generative pre-training. (2018).
drik Strobelt. 2023. HAI-GEN 2023: 4th Workshop on Human-AI Co-Creation [138] Alec Radford, Jefrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya
with Generative Models. In Companion Proceedings of the 28th International Sutskever, et al. 2019. Language models are unsupervised multitask learners.
Conference on Intelligent User Interfaces. 190–192. OpenAI blog 1, 8 (2019), 9.
[113] Céline Mariage, Jean Vanderdonckt, and Costin Pribeanu. 2011. State of the [139] Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen.
Art of Web Usability Guidelines. In Handbook of human factors in Web design, 2022. Hierarchical text-conditional image generation with clip latents. arXiv
Kim-Phuong L Vu and Robert W Proctor (Eds.). Crc Press. preprint arXiv:2204.06125 (2022).
[114] Christopher McComb, Peter Boatwright, and Jonathan Cagan. 2023. FOCUS [140] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn
AND MODALITY: DEFINING A ROADMAP TO FUTURE AI-HUMAN TEAM- Ommer. 2022. High-resolution image synthesis with latent difusion models. In
ING IN DESIGN. Proceedings of the Design Society 3 (2023), 1905–1914. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni-
[115] Donald McMillan, Alistair Morrison, and Matthew Chalmers. 2013. Categorised tion. 10684–10695.
ethical guidelines for large scale mobile HCI. In Proceedings of the SIGCHI [141] Dagfnn Rømen and Dag Svanæs. 2012. Validating WCAG versions 1.0 and 2.0
Conference on Human Factors in Computing Systems. 1853–1862. through usability testing with disabled users. Universal Access in the Information
[116] Fadel M Megahed, Ying-Ju Chen, Joshua A Ferris, Sven Knoth, and L Allison Society 11 (2012), 375–385.
Jones-Farmer. 2023. How generative ai models such as chatgpt can be (mis) [142] John Rooksby, Parvin Asadzadeh, Alistair Morrison, Claire McCallum, Cindy
used in spc practice, education, and research? an exploratory study. Quality Gray, and Matthew Chalmers. 2016. Implementing ethics for a mobile app
Engineering (2023), 1–29. deployment. In Proceedings of the 28th Australian Conference on Computer-
[117] Cade Metz. 2022. Meet GPT-3. It Has Learned to Code (and Blog and Ar- Human Interaction. 406–415.
gue). (Published 2020). [Link] [143] Steven I Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D
[Link] Weisz. 2023. The Programmer’s Assistant: Conversational Interaction with a
[118] Lalit Mohan, Neeraj Mathur, and Y Raghu Reddy. 2015. Mobile App Usability Large Language Model for Software Development. In 28th International Confer-
Index (MAUI) for improving mobile banking adoption. In 2015 International ence on Intelligent User Interfaces.
Conference on Evaluation of Novel Approaches to Software Engineering (ENASE). [144] Mattias Rost and Sebastian Andreasson. 2023. Stable Walk: An interactive
IEEE, 313–320. environment for exploring Stable Difusion outputs. 3124 (2023). [Link]
[119] Sina Mohseni, Niloofar Zarei, and Eric D Ragan. 2021. A multidisciplinary [Link]/Vol-3359/[Link]
survey and framework for design and evaluation of explainable AI systems. [145] Vildan Salikutluk, Dorothea Koert, and Frank Jäkel. 2023. Interacting with Large
ACM Transactions on Interactive Intelligent Systems (TiiS) 11, 3-4 (2021), 1–45. Language Models: A Case Study on AI-Aided Brainstorming for Guesstimation
[120] Michael Muller, Plamen Agelov, Hal Daume, Q Vera Liao, Nuria Oliver, David Problems. In HHAI 2023: Augmenting Human Intellect. IOS Press, 153–167.
Piorkowski, et al. 2022. HCAI@NeurIPS 2022, Human Centered AI. In Annual [146] Jose Ma Santiago III, Richard Lance Parayno, Jordan Aiko Deja, and Briane
Conference on Neural Information Processing Systems. Paul V Samson. 2023. Rolling the Dice: Imagining Generative AI as a Dungeons
[121] Michael Muller, Heloisa Candello, and Justin Weisz. 2023. Interactional Co- & Dragons Storytelling Companion. arXiv preprint arXiv:2304.01860 (2023).
Creativity of Human and AI in Analogy-Based Design. In International Confer- [147] Regina Schober. 2022. Passing the Turing test? AI generated poetry and posthu-
ence on Computational Creativity. man creativity. Artifcial Intelligence and Human Enhancement: Afrmative and
[122] Michael Muller, Lydia B Chilton, Anna Kantosalo, Charles Patrick Martin, and Critical Approaches in the Humanities 21 (2022), 151.
Greg Walsh. 2022. GenAICHI: Generative AI and HCI. In CHI Conference on [148] Isabella Seeber, Eva Bittner, Robert O Briggs, Triparna De Vreede, Gert-Jan
Human Factors in Computing Systems Extended Abstracts. 1–7. De Vreede, Aaron Elkins, Ronald Maier, Alexander B Merz, Sarah Oeste-Reiß,
[123] Michael Muller, Justin D. Weisz, and Werner Geyer. 2020. Mixed ini- Nils Randrup, et al. 2020. Machines as teammates: A research agenda on AI in
tiative generative AI interfaces: An analytic framework for generative team collaboration. Information & management 57, 2 (2020), 103174.
AI applications. ICCC 2020 Workshop, The Future of Co-Creative Sys- [149] Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang.
tems. [Link] 2022. On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in
Future_of_co-creative_systems_185.pdf Zero-Shot Reasoning. arXiv preprint arXiv:2212.08061 (2022).
[124] Stacey F Nagata. 2003. Multitasking and interruptions during mobile web tasks. [150] Yashraj Shashikant Sharma. 2023. Generating Wildfre Risk Maps for Critical
In Proceedings of the Human Factors and Ergonomics Society Annual Meeting, Infrastructure Systems Using Integrated Generative AI and Simulation Techniques
Vol. 47. SAGE Publications Sage CA: Los Angeles, CA, 1341–1345. Under Information Uncertainty. Ph. D. Dissertation. State University of New
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

York at Bufalo. [173] Mathias Peter Verheijden and Mathias Funk. 2023. Collaborative Difusion:
[151] Maria Shitkova, Justus Holler, Tobias Heide, Nico Clever, and Jörg Becker. 2015. Boosting Designerly Co-Creation with Generative AI. In Extended Abstracts of
Towards usability guidelines for mobile websites and applications. Wirtschaftsin- the 2023 CHI Conference on Human Factors in Computing Systems. 1–8.
formatik Proceedings (2015). [Link] [174] Benedikte Wallace, Charles P Martin, Jim Tørresen, and Kristian Nymoen. 2021.
1106&context=wi2015 Learning embodied sound-motion mappings: Evaluating AI-generated dance
[152] Ben Shneiderman. 2020. Bridging the gap between ethics and practice: guide- improvisation. In Creativity and Cognition. 1–9.
lines for reliable, safe, and trustworthy human-centered AI systems. ACM [175] Qian Wan, Siying Hu, Yu Zhang, Piaohong Wang, Bo Wen, and Zhicong Lu. 2023.
Transactions on Interactive Intelligent Systems (TiiS) 10, 4 (2020), 1–31. "It Felt Like Having a Second Mind": Investigating Human-AI Co-creativity in
[153] Ben Shneiderman. 2020. Human-centered artifcial intelligence: Reliable, safe & Prewriting with Large Language Models. arXiv preprint arXiv:2307.10811 (2023).
trustworthy. International Journal of Human–Computer Interaction 36, 6 (2020), [176] Qiaosi Wang, Michael Madaio, Shaun Kane, Shivani Kapania, Michael Terry,
495–504. and Lauren Wilcox. 2023. Designing Responsible AI: Adaptations of UX Practice
[154] Ben Shneiderman and Michael Muller. 2023. On AI Anthropomorphism. Human- to Meet Responsible AI Challenges. In Proceedings of the 2023 CHI Conference
Centered AI (Medium) (10 April 2023). [Link] on Human Factors in Computing Systems. 1–16.
ai/on-ai-anthropomorphism-abf4cecc5ae [177] Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel. 2021.
[155] Ben Shneiderman, Catherine Plaisant, Maxine S Cohen, Steven Jacobs, Niklas Towards mutual theory of mind in human-ai interaction: How language refects
Elmqvist, and Nicholas Diakopoulos. 2016. Designing the user interface: strategies what students perceive about a virtual teaching assistant. In Proceedings of the
for efective human-computer interaction. Pearson. 2021 CHI conference on human factors in computing systems. 1–14.
[156] RC Sidorsky, RN Parrish, JL Gates, SJ Munger, and SYNECTICS CORP FAIRFAX [178] Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Feris, Huan Sun, and
VA. 1984. Design guidelines for user transactions with battlefeld automated Yoon Kim. 2023. Multitask prompt tuning enables parameter-efcient transfer
systems: Prototype for a handbook. ARI Researoh Product (1984), 84–08. learning. arXiv preprint arXiv:2303.02861 (2023).
[157] Raymond C Sidorsky and Robert N Parrish. 1980. Guidelines and criteria for [179] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi,
human-computer interface design of battlefeld automated systems. In Proceed- Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits rea-
ings of the Human Factors Society Annual Meeting, Vol. 24. SAGE Publications soning in large language models. Advances in Neural Information Processing
Sage CA: Los Angeles, CA, 98–102. Systems 35 (2022), 24824–24837.
[158] Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan, Vaibhav [180] Laura Weidinger, John Mellor, Maribeth Rauh, Conor Grifn, Jonathan Uesato,
Aggarwal, Aaron Adcock, Armand Joulin, Piotr Dollár, Christoph Feichtenhofer, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al.
Ross Girshick, et al. 2023. The efectiveness of MAE pre-pretraining for billion- 2021. Ethical and social risks of harm from language models. arXiv preprint
scale pretraining. arXiv preprint arXiv:2303.13496 (2023). arXiv:2112.04359 (2021).
[159] Sidney L Smith and Jane N Mosier. 1986. Guidelines for designing user interface [181] Justin D Weisz, Mary Lou Maher, Hendrik Strobelt, Lydia B Chilton, David
software. Citeseer. Bau, and Werner Geyer. 2022. HAI-GEN 2022: 3rd Workshop on Human-AI Co-
[160] Nikita Soni, Aishat Aloba, Kristen S Morga, Pamela J Wisniewski, and Lisa An- Creation with Generative Models. In 27th International Conference on Intelligent
thony. 2019. A framework of touchscreen interaction design recommendations User Interfaces. 4–6.
for children (tidrc) characterizing the gap between research evidence and design [182] Justin D Weisz, Michael Muller, Jessica He, and Stephanie Houde. 2023. Toward
practice. In Proceedings of the 18th ACM international conference on interaction general design principles for generative AI applications. In Joint Proceedings
design and children. 419–431. of the IUI 2023 Workshops: HAI-GEN, ITAH, MILC, SHAI, SketchRec, SOCIALIZE
[161] Angie Spoto and Natalia Oleynik. 2017. Library of Mixed-Initiative Creative co-located with the ACM International Conference on Intelligent User Interfaces,
Interfaces. Retrieved 14-Aug-2023 from [Link] Vol. 3124. CEUR. [Link]
[162] Ayushi Srivastava, Shivani Kapania, Anupriya Tuli, and Pushpendra Singh. [183] Justin D Weisz, Michael Muller, Stephanie Houde, John Richards, Steven I Ross,
2021. Actionable UI Design Guidelines for Smartphone Applications Inclusive Fernando Martinez, Mayank Agarwal, and Kartik Talamadupula. 2021. Per-
of Low-Literate Users. Proceedings of the ACM on Human-Computer Interaction fection Not Required? Human-AI Partnerships in Code Translation. In 26th
5, CSCW1 (2021), 1–30. International Conference on Intelligent User Interfaces. 402–412.
[163] Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, [184] Justin D Weisz, Michael Muller, Steven I Ross, Fernando Martinez, Stephanie
Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Houde, Mayank Agarwal, Kartik Talamadupula, and John T Richards. 2022. Bet-
Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and ex- ter together? an evaluation of ai-supported code translation. In 27th International
trapolating the capabilities of language models. arXiv preprint arXiv:2206.04615 Conference on Intelligent User Interfaces. 369–391.
(2022). [185] Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry
[164] Luke Stark, Jen King, Xinru Page, Airi Lampinen, Jessica Vitak, Pamela Wis- Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023.
niewski, Tara Whalen, and Nathaniel Good. 2016. Bridging the gap between A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv
privacy by design and privacy in practice. In Proceedings of the 2016 CHI Confer- preprint arXiv:2302.11382 (2023).
ence Extended Abstracts on Human Factors in Computing Systems. 3415–3422. [186] Chathurika S Wickramasinghe, Daniel L Marino, Javier Grandio, and Milos
[165] Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Manic. 2020. Trustworthy AI development guidelines for human system inter-
Hanspeter Pfster, and Alexander M Rush. 2022. Interactive and visual prompt action. In 2020 13th International Conference on Human System Interaction (HSI).
engineering for ad-hoc task adaptation with large language models. IEEE IEEE, 130–136.
transactions on visualization and computer graphics 29, 1 (2022), 1146–1156. [187] Alan FT Winfeld. 2018. Experiments in artifcial theory of mind: From safety
[166] Jiao Sun, Q Vera Liao, Michael Muller, Mayank Agarwal, Stephanie Houde, to story-telling. Frontiers in Robotics and AI 5 (2018), 75.
Kartik Talamadupula, and Justin D Weisz. 2022. Investigating Explainability of [188] Austin P Wright, Zijie J Wang, Haekyu Park, Grace Guo, Fabian Sperrle, Men-
Generative AI for Code through Scenario-based Design. In 27th International natallah El-Assady, Alex Endert, Daniel Keim, and Duen Horng Chau. 2020.
Conference on Intelligent User Interfaces. 212–228. A comparative analysis of industry human-AI interaction guidelines. arXiv
[167] Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, preprint arXiv:2010.11761 (2020).
Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019. [189] Sang Michael Xie and Sewon Min. 2022. How does in-context learning work?
Mitigating gender bias in natural language processing: Literature review. arXiv A framework for understanding the diferences from traditional supervised
preprint arXiv:1906.08976 (2019). learning. [Link]
[168] Shari Trewin, Sara Basson, Michael Muller, Stacy Branham, Jutta Treviranus, [190] Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou,
Daniel Gruen, Daniel Hebert, Natalia Lyckowski, and Erich Manser. 2019. Con- and Xinyun Chen. 2023. Large Language Models as Optimizers. arXiv preprint
siderations for AI fairness for people with disabilities. AI Matters 5, 3 (2019), arXiv:2309.03409 (2023).
40–63. [191] Nur Yildirim, Alex Kass, Teresa Tung, Connor Upton, Donnacha Costello, Robert
[169] Josh Urban Davis, Fraser Anderson, Merten Stroetzel, Tovi Grossman, and Giusti, Sinem Lacin, Sara Lovic, James M O’Neill, Rudi O’Reilly Meehan, et al.
George Fitzmaurice. 2021. Designing co-creative ai for virtual environments. In 2022. How Experienced Designers of Enterprise Applications Engage AI as a
Creativity and Cognition. 1–11. Design Material. In Proceedings of the 2022 CHI Conference on Human Factors in
[170] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Computing Systems. 1–13.
Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you [192] Nur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg, and Fer-
need. Advances in neural information processing systems 30 (2017). nanda Viégas. 2023. Investigating How Practitioners Use Human-AI Guidelines:
[171] Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Shomir A Case Study on the People+ AI Guidebook. In Proceedings of the 2023 CHI
Wilson, et al. 2023. Nationality Bias in Text Generation. arXiv preprint Conference on Human Factors in Computing Systems. 1–13.
arXiv:2302.02463 (2023). [193] JD Zamfrescu-Pereira, Richmond Y Wong, Bjoern Hartmann, and Qian Yang.
[172] Oleksandra Vereschak, Gilles Bailly, and Baptiste Caramiaux. 2021. How to eval- 2023. Why Johnny can’t prompt: how non-AI experts try (and fail) to design
uate trust in AI-assisted decision making? A survey of empirical methodologies. LLM prompts. In Proceedings of the 2023 CHI Conference on Human Factors in
Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–39. Computing Systems. 1–21.
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

[194] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, cases that apply these technologies for their own sake may fail to
Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. deliver real user value.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv preprint
arXiv:2306.05685 (2023). Example: Human-centered approaches such as design thinking
[195] Shuoyang Zheng. 2023. StyleGAN-Canvas: Augmenting StyleGAN3 for Real- and participatory design allow you to observe users’ workfows and
Time Human-AI Co-Creation. (2023).
[196] Daniel M Ziegler, Nisan Stiennon, Jefrey Wu, Tom B Brown, Alec Radford,
pain points to ensure proposed uses of generative AI are aligned
Dario Amodei, Paul Christiano, and Geofrey Irving. 2019. Fine-tuning language with users’ actual needs. For example, involving stakeholders in
models from human preferences. arXiv preprint arXiv:1909.08593 (2019). the co-design of prototypes can serve as a probe for discussions
around user value and technological feasibility.
A EXTENDED DESCRIPTIONS AND
A.1.2 Identify and resolve value tensions*. Often, there are multiple
EXAMPLES stakeholders involved in the creation of a generative AI application,
We provide extended descriptions for each design principle and including the end users, those who design and build the application
strategy to ofer deeper insight into their meaning and application, (e.g. designers, developers, product managers), and those who make
written in second-person for an audience of design practitioners. purchasing or licensing decisions (e.g. CIOs, CEOs). When these
We also provide examples for each design strategy to illustrate stakeholders’ values are not aligned, it results in a value tension.
how they have been used in a realistic context. Each example was Addressing these tensions is important for creating feasible and
drawn either from a commercial generative AI application or an valuable products that meet end users’ needs.
experimental generative AI system. Most examples were identifed Example: Value Sensitive Design (VSD) [48] is a method that can
in the modifed heuristic evaluation of Iteration 3; new examples help designers identify who the important stakeholders are and
were found for the three new strategies added after that iteration. navigate the value tensions that exist across them.
For process-related strategies (denoted with an asterisk (*)), we
discuss how the process would theoretically be applied as we did A.1.3 Expose or limit emergent behaviors*. Generative AI can ex-
not have visibility into the actual design process for the generative hibit emergent behaviors — the ability to perform tasks beyond
AI applications we examined. ones they were trained for. These emergent behaviors can be a
We make reference to the following commercial systems in our delighter, such as when a conversational Q&A system answers an
examples: out-of-domain question, or they can be a risk, such as when that
output is toxic or aggressive. As a designer, you should carefully con-
• Adobe Firefy: [Link] sider the trade-ofs between optimizing your user experience for a
• Adobe Photoshop: [Link] well-defned set of capabilities versus providing a more open-ended
html experience that may surface potentially risky emergent behaviors.
• AIVA: [Link] Example: Conversational interfaces that enable open-ended inter-
• ChatGPT: [Link] actions will allow such emergent behaviors to surface. For example,
• DALL-E: [Link] a user may discover that ChatGPT can perform sentiment analysis,
• DreamStudio (powered by Stable Difusion): [Link] a task that it (likely) wasn’t explicitly trained to do. By contrast,
ai graphical user interfaces (GUIs), such as AIVA, can place limits on
• Github Copilot: [Link] the ways a user can interact with the underlying generative model
• Google Bard: [Link] by only exposing selected functionality.
• Midjourney: [Link]
We also note that a single type of UX feature or functionality A.1.4 Test & monitor for user harms*. Generative models may pro-
can be used to implement more than one design strategy. We have duce a variety of harmful outputs, such as language that is hateful,
included similar or duplicate examples below to illustrate this point. abusive, profane, or otherwise toxic. They may also harm users by
failing to produce outputs that provide a fair or accurate representa-
A.1 Design responsibly tion of diversity. It is imperative to work closely with technologists
to understand and evaluate the potential risks stemming from the
The most important principle to follow when designing genera- use of a generative model in an application. It is also imperative
tive AI systems is to design responsibly. The use of all AI systems, to assume that user harms will occur and develop reporting and
including those that incorporate generative capabilities, may un- escalation mechanisms for when they do.
fortunately lead to diverse forms of harms, especially for people in Example: One way to test for harms is by benchmarking models
vulnerable situations. As designers, it is imperative that we adopt on known data sets of hate speech [60] and bias [45, 149, 171]. After
a socio-technical perspective toward designing responsibly: when deploying an application, harms can be fagged through mecha-
technologists recommend new technical mechanisms to incorpo- nisms that allow users to report problematic model outputs.
rate into a generative AI system, we should question how those
mechanisms will improve the user’s experience, provide them with A.2 Design for mental models
new capabilities, or address their pain points.
A mental model is a simplifed representation of the world that
A.1.1 Use a human-centered approach*. Technosolutionism is the people use to process new information and make predictions [80].
idea that technology will solve all of our (human) problems, and It is their own understanding of how something works and how
it should be avoided at all costs. Human-centered approaches can their actions afect it. Generative AI poses new challenges to users,
determine whether the use of generative AI is appropriate; use and designers must carefully consider how to impart useful mental
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

models to their users to help them understand how a system works you talk about for hours?” In this way, users teach ChatGPT about
and how their actions afect that system. Also consider the user’s themselves in order to receive more personalized responses.
background and goals and how to help the AI form a “mental model”
of the user. A.3 Design for appropriate trust & reliance
A.2.1 Orient the user to generative variability. Help the user under- Trustworthy generative AI applications are those that produce
stand the AI system’s behavior, and that it may produce multiple, high-quality, useful, and (where applicable) factual outputs that
varied outputs that may not be reproducible, even when given the are faithful to a source of truth. Calibrating users’ trust is crucial
same input. This behavior will be unexpected for novice users be- for establishing appropriate reliance: teaching users to scrutinize
cause it is fundamentally diferent from traditional AI systems that a model’s outputs for quality issues, inaccuracies, biases, under-
always give the same outcome for the same input. representation, and other issues to determine whether they are
Example: Google Bard provides answers in the form of multiple acceptable (e.g. because they achieve a certain level of quality or
drafts, indicating that it came up with multiple, varied answers for veracity) or if they should be modifed or rejected.
the same question.
A.3.1 Calibrate trust using explanations. Be clear and upfront about
A.2.2 Teach efective use. Users need to understand how to work what the application can and cannot do by explaining its capabilities
efectively with a generative AI application to accomplish their and limitations. Teach users to be skeptical of potentially imperfect
goals. Mechanisms such as tutorials, examples, explanations, and model outputs and help them understand when they can trust the
social transparency [43] (i.e. showing other users’ inputs and out- system.
puts) can help users form mental models for how to efectively use Example: ChatGPT explains its capabilities (e.g. “answer ques-
an application. tions, help you learn, write code, brainstorm together”) and limi-
Example: DALL-E provides curated examples of generated out- tations (e.g. “ChatGPT may give you inaccurate information. It’s
puts and the prompts used to generate them. Adobe Photoshop not intended to give advice.”) directly on its introduction screen.
provides pop-ups and tooltips to introduce the user to its Generative Google Bard provides a notice below the prompt feld that states,
Fill feature. “Bard may display inaccurate or ofensive information that doesn’t
represent Google’s views.”
A.2.3 Understand the user’s mental model*. Conduct evaluations,
such as interviews, to determine whether a user has formed a useful A.3.2 Provide rationales for outputs. Show the user why a par-
mental model of a generative AI application. One prompt that can ticular output was generated by showing the model’s “chain of
be useful to ask is how they think the application provides a certain thought” [179] or identifying the source materials used to gen-
capability, which forces the user to articulate their theory of how erate it. For example, identifying source documents for answers
the system works. The goal is not for the user to possess an accurate expected to be factually correct or revealing the image sets that a
model, but rather, one that is useful for working efectively with text-to-image model was trained on can help users calibrate their
the system. Furthermore, understanding the user’s mental model trust.
can also help you leverage their existing knowledge of similar Example: Google Bard provides a list of sources it used to produce
applications to inform your design decisions. answers to questions. Adobe discloses that its Generative Fill feature
Example: In evaluating a Q&A application, you might ask the was trained on “stock imagery, openly licensed work, and public
user, “how did the system answer your question about who the domain content where the copyright has expired” [1].
current President is?” Answers such as, “it looked it up on the
web” might indicate a need to educate users about hallucination A.3.3 Use friction to avoid overreliance. Encourage the user to
issues. Users’ existing mental models of other applications can also review and think critically about the generative model’s outputs
be useful to understand. For example, Github Copilot builds on by introducing mechanisms that slow them down at key decision-
users’ mental models by following the same interaction pattern as making points. These mechanisms are known as cognitive forcing
its existing code completion features, which are familiar to many functions [22]. Examples include ofering multiple AI-generated
developers, hence easing their learning curve. options for the user to select from, highlighting uncertainty, or
requiring the user to create content or render a decision before
A.2.4 Teach the AI system about the user. LLMs are adept at tailor- showing the AI’s output.
ing their language to a target audience. Designers can induce these Example: Google Bard displays multiple drafts for the user to
models to produce personalized responses to users – in essence, review, which can encourage them to slow down and consider
teaching the model about the user – by including additional prompt which drafts may be of lower or higher quality.
text such as, “explain like I’m fve” or “please give me a detailed,
technical answer.” Capturing the user’s expectations, behaviors, and A.3.4 Signify the role of the AI. AI systems may act in diferent
preferences can improve the AI’s interactions with them. Users can roles, such as “tools,” “partners,” “analytics,” or “coaches” [114].
also provide information about their background in UI outside of These roles shape users’ expectations of how the AI system fts into
their prompt, which the application can then opaquely incorporate their workfow, such as the extent to which it takes initiative (e.g.
back into the prompt. by acting proactively vs. reactively) and agency (e.g. by directly
Example: ChatGPT provides a form for “Custom Instructions” manipulating an artifact vs. making suggestions or recommenda-
in which users provide answers to questions such as, “Where are tions). The role that is signifed shapes users’ perceptions of the
you based?”, “What do you do for work?”, and “What subjects can system: research by Kim et al. [84] shows that AI “tools” are viewed
Design Principles for Generative AI Applications CHI ’24, May 11–16, 2024, Honolulu, HI, USA

as less genuine and caring than AI “mediators” or AI “assistants.” and name multiple public or private collections to organize their
Anthropomorphic signifers such as giving a human name to an work.
AI, representing the AI with a human-like avatar, having the AI
refer to itself using frst-person pronouns, or showing an animated A.4.4 Draw atention to diferences or variations across outputs. A
typing bubble to hide inference latency may all give users the false generative model can sometimes produce a set of similar outputs
impression that they are interacting with a human. Although there that are difcult to tell apart. When outputs are similar, tools that
is no clear consensus on the extent to which AI-infused user expe- aid users in identifying the similarities and diferences between
riences should favor or discourage anthropomorphism [154], as a multiple outputs can be useful.
designer, it is crucial to clearly signify to the user when they are Example: DreamStudio, DALL-E, and Midjourney all display
interacting with an AI system and what content was AI-generated. multiple outputs in a grid-like fashion to allow the user to identify
Example: Github Copilot’s tagline is “Your AI pair programmer”, diferences, but fne-grained diferences between outputs are not
which elicits the role of a partner. Copilot fulflls this role by proac- explicitly highlighted. A prototype source code translation inter-
tively making suggestions as the user writes code. It also possesses face by Weisz et al. [183] visualizes the diferences across multiple
a limited form of agency by making autocompletion suggestions generated code translations through granular highlights, as well as
directly in the user’s code editor, although it requires the user to interactively through a list of alternate translations.
explicitly accept or reject those suggestions (e.g. by pressing tab or
escape). A.5 Design for co-creation
Generative AI ofers new co-creative capabilities. Help the user cre-
A.4 Design for generative variability ate outputs that meet their needs by providing controls that enable
One distinguishing characteristic of generative AI systems is that them to infuence the generative process and work collaboratively
they can produce multiple outputs that vary in character or quality, with the AI.
even when the user’s input does not change. This characteristic
A.5.1 Help the user craf efective outcome specifications. Genera-
raises important design considerations: to what extent should mul-
tive AI has introduced a new interaction paradigm, intent-based
tiple outputs be visible to users, and how might we help users
outcome specifcation, in which users specify what they want but
organize and select amongst varied outputs?
not how it should be produced. Orient the user to this new para-
A.4.1 Leverage multiple outputs. Take advantage of multiple out- digm and assist them in prompting efectively to produce outputs
puts to help users produce the one that fts their needs. Multiple that ft their needs.
outputs can be exposed to the user or remain under the hood and Example: Google Bard sends out a newsletter that includes a
instead allow the model to select the best option(s) to surface. section called “Prompt Engineering 101,” which features tips and
Example: DreamStudio, DALL-E, and Midjourney all generate examples to help users improve their prompt writing. The IBM
multiple distinct outputs for a given prompt; for example, Dream- [Link] Prompt Lab documentation includes a set of tips and
Studio produces four images by default and can be confgured to examples to help the user understand how to improve their prompts.
produce up to 10. ChatGPT allows the user to regenerate a response
to see more options. A.5.2 Provide generic input parameters. Let the user control generic
aspects of the generative process such as the number of outputs and
A.4.2 Visualize the user’s journey. Users may not be able to repro- the random seed used to produce those outputs. Generic controls
duce prior outputs. Capturing and displaying the user’s history of apply across most use cases, independent of which model is used.
outputs, along with the input parameters used (e.g. prompts, con- Example: DreamStudio provides a slider for users to indicate the
trols) can help them track their work. Also consider ways to guide number of images they want to produce for a given prompt, along
the user to new output possibilities that they have not explored. with an input feld for random seed.
Example: DreamStudio, DALL-E, and Midjourney all show a
history of the user’s inputs and resulting image outputs. Research A.5.3 Provide controls relevant to the use case and technology. Let
prototypes extend the idea of “visualizing the user’s journey” even the user control parameters specifc to their use case, domain, or
further using visualization techniques to show unexplored parts of model architecture. Some model architectures provide specialized
an output space [88] (using dots overlaid on 2D histograms) or of a means of control, such as semantic sliders for latent space mod-
parameter confguration space [144] (by showing confgurations a els [103] or decoding strategy and temperature for LLMs. Other
user has and has not yet tried in a grid, with untried confgurations models have controls specifc to the domain of the application (e.g.
rendered as placeholders). code, music, art).
Example: AIVA allows the user to customize domain-specifc
A.4.3 Enable curation & annotation. When a generative model pro- characteristics of the musical compositions it generates, such as
duces multiple outputs, users may need to curate or annotate them. the type of ensemble and emotion.
Curation may include collecting, fltering, or organizing outputs
(possibly from the generation history). Annotation may include A.5.4 Support co-editing of generated outputs. Allow both the user
the ability to tag artifacts (e.g. “I like this picture”) or make notes and the AI system to improve generated outputs. Although the user
within an artifact (e.g. “this line of code looks suspicious”). should retain agency and decision-making authority, the AI can
Example: DALL-E allows the user to mark images as favorites co-create with the user to improve outputs, such as by providing
and store them within groups called collections. Users may create suggestions, organizing outputs, or editing an artifact.
CHI ’24, May 11–16, 2024, Honolulu, HI, USA Weisz et al.

Example: Adobe Photoshop exposes generative AI capabilities A.6.2 Evaluate outputs using domain-specific metrics. In some cases,
within the same design surface as its other image editing tools, the quality of a generative model’s outputs can be assessed with
enabling both the user and the generative AI model to co-edit an measurable criteria. For example, answers to customer queries can
image. be evaluated for faithfulness to source documents, and design mock-
ups can be evaluated for adherence to a UI style guide.
A.6 Design for imperfection Example: Molecular candidates generated by CogMol [29], a
Users must understand that generative model outputs may be im- prototype generative application for drug design, are evaluated
perfect according to objective metrics (e.g. untruthful or misleading with a molecular simulator to compute domain-specifc attributes
answers, violations of prompt specifcations) or subjective metrics such as molecular weight, water solubility, and toxicity.
(e.g. the user doesn’t like the output). Provide transparency by iden- A.6.3 Ofer ways to improve outputs. Provide ways for the user to
tifying or highlighting possible imperfections, and help the user fx imperfections and improve output quality, such as editing tools,
understand and work with outputs that may not align with their an option to regenerate, or providing alternative outputs to select
expectations. from.
A.6.1 Make uncertainty visible. Caution the user that outputs may Example: DALL-E and DreamStudio allow users to refne outputs
not align with their expectations and identify detectable uncertain- by erasing and regenerating parts of an image (inpainting) or gen-
ties or faws. Provide disclaimers about the potential for imperfec- erating new parts of the image beyond its boundaries (outpainting).
tion, show the model’s confdence level when possible, and utilize Google Bard ofers options for the user to modify outputs to be
“I don’t know” responses when confdence is low. shorter, longer, simpler, more casual, or more professional.
Example: Google Bard’s interface states, “Bard may display inac- A.6.4 Provide feedback mechanisms. Allow users to provide feed-
curate info, including about people, so double-check its responses.” back on the quality of a model’s output. Feedback may be implicit
This disclaimer alerts the user to the possibility of uncertainties or (e.g. the user makes edits to a generated artifact indicating potential
imperfections in its outputs. A prototype source code translation in- issues) or it may be explicit (e.g. rating the quality with a thumbs
terface proposed by Weisz et al. [183] makes the generative model’s up / thumbs down).
uncertainty visible to the user by highlighting source code tokens Example: ChatGPT ofers an option for the user to provide a
based on the degree to which the underlying model is confdent thumbs up or thumbs down rating for its responses, along with
that they were correctly translated. These highlights help guide the open-ended textual feedback.
user’s attention toward places that require their review.

You might also like