0% found this document useful (0 votes)
5 views6 pages

User Studies in Visualization Techniques

The document discusses the importance of user studies in evaluating visualization techniques, emphasizing their role in measuring performance and understanding the effectiveness of different methods. It outlines best practices for designing experiments, including considerations for participant selection and task appropriateness, as well as the application of psychophysical theories to improve visualization design. Additionally, it highlights the need for controlled studies to explore how texture and color can enhance shape perception in visualizations.

Uploaded by

Patrick Ogao
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views6 pages

User Studies in Visualization Techniques

The document discusses the importance of user studies in evaluating visualization techniques, emphasizing their role in measuring performance and understanding the effectiveness of different methods. It outlines best practices for designing experiments, including considerations for participant selection and task appropriateness, as well as the application of psychophysical theories to improve visualization design. Additionally, it highlights the need for controlled studies to explore how texture and color can enhance shape perception in visualizations.

Uploaded by

Patrick Ogao
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Visualization Viewpoints

Editor: Theresa-Marie Rhyne

User Studies: Why, How, and When? __________________


Robert Kosara
VRVis Research
Center, Austria
I n crafting today’s visualizations, we often design and
evaluate methods by presenting results informally to
potential users. No matter how efficient a visualization
puter vision may not extend to a visualization environ-
ment. We can run user studies to test this hypothesis.
Results can show when the theories hold and how they
technique may be, or how well motivated from theory, if need to be modified to function correctly for real-world
Christopher G. it doesn’t convey information effectively, it’s of little use. data and tasks.
Healey A good starting point in any study is the scientific or
North Carolina Why conduct user studies? visual design question to be examined. This drives the
State University User studies offer a scientifically sound method to process of experimental design. A poorly designed
measure a visualization’s performance. The reasons experiment will yield results of only limited value.
Victoria abound to pursue user studies, particularly when eval- Although a comprehensive discussion of experimental
Interrante uating the strengths and weaknesses of different visu- design is beyond the scope of this article, we offer some
University of alization techniques. For example, in Figure 1 Laidlaw suggestions and lessons learned in the “Basics of User
Minnesota compared six methods for visualizing 2D vector fields.1 Study Design” sidebar. We also describe in the sidebar
His experiments measured user performance on three how we designed experiments to answer important
David H. flow-related tasks for each of the six methods. He used questions from our own research.
Laidlaw the results to identify what makes a 2D vector field visu-
Brown alization effective. Color sequences
University Studies can show that a new visualization technique One reason for conducting studies is to determine if
is useful in a practical sense, according to some objec- we can apply theoretical principles derived from other
Colin Ware tive criteria, for a specific task. Even more exciting are disciplines (such as psychophysics) to visualization
University of studies (like Laidlaw’s) that show that a new technique design. Researchers have studied human color vision
New Hampshire is more effective than an existing technique for an theory for more than a century. Results from this work
important task. User studies can objectively establish provide a solid foundation for using color in visualiza-
which method is most appropriate for a given situation. tion. However, choosing colors for a particular visual-
A more fundamental goal of conducting user studies ization problem is normally very different from the
is to seek insight into why a particular technique is effec- extremely simple displays used by experimental psy-
tive. This can guide future efforts to improve existing chologists. We need to experiment to bridge this gap
techniques. We want to understand what types of tasks between theory and practice.
and conditions yield high-quality results for a particu- Consider the problem of designing pseudocolor
lar method. This knowledge is critical because different sequences for scientific images. We have a continuous
analysis tasks require different visualization techniques. data field over a plane (for example, an energy or densi-
A final use for studies in visualization is to show that ty distribution), and we want to use color to illustrate fea-
an abstract theory applies under certain practical con- tures in the data. Human vision theory dictates that
ditions. For example, results from psychophysics or com- neural signals from the rods and cones in the retina are
1 Three of six
visualization
methods1 com-
pared with a
user study. Each
method shows
the same vector
field. User
performance on
different tasks
provides quanti-
tative compar-
isons of the
methods.

20 July/August 2003 Published by the IEEE Computer Society 0272-1716/03/$17.00 © 2003 IEEE
Basics of User Study Design

While a complete tutorial on user studies is appropriate. At one end of the spectrum is the
beyond the scope of a short article, we hope to rigorous application of signal detection methods.1
share some useful lessons we’ve learned. We can use these to assess the detectability of a
The approach we advocate is a form of applied target structure from a background of noise. A
perception research. Proper use of this technique more common experiment type is the evaluation
requires an understanding of how to build of a number of different visual features. For
experiments that include human participants. It’s example, a study might address the question of
challenging to design an experiment that will give how well motion parallax, stereoscopic depth, and
robust answers to the questions of interest. A surface texture contribute to the perception of
typical study might ask, Which prospective method surface shape. Such an experiment calls for a
is most promising? Do any of these methods factorial design with analysis of variance (ANOVA)
perform better than the best available alternative? to evaluate the results.
Unfortunately, many problems can compromise a Another concern is how many participants to
study’s validity or make it difficult to draw useful use. The answer depends on what’s being studied.
insights from the results. Is the task appropriate? Is For psychophysical experiments that measure low-
it possible that participants were using cues other level visual phenomena, it’s acceptable to use only
than the ones being examined to perform the task? a few participants. This is because there’s little
Is there a control condition to provide a baseline variation in viewer’s reactions. These experiments
for comparison between different methods? Do all contain numerous repeated measures (that is,
participants have a correct and equivalent multiple trials with the same experimental
understanding of the task? Are all participants conditions) to ensure a sufficient total number of
sufficiently willing and able to perform the task? Is trials. If cognitive (as opposed to purely perceptual)
there a learning effect, wherein the participant processes are involved, more participants are
performs the task better because he or she has normally required. Counterbalancing participants
already solved a similar task before? based on characteristics like gender, age, or
We can address these problems by experience may also be necessary. A detailed
description of both participants and methods is an
■ testing participants for adequate spatial acuity, essential component for any publication involving
stereo ability, and absence of color blindness; user studies.
■ randomizing the presentation order of the tri- Finally, researchers at US universities should be
als, by using written instructions; aware that they may be required to obtain prior
■ letting participants rest during the experiment approval (or exemption) from the Institutional
to avoid becoming fatigued; Review Board at their institution before
■ devising robust methods to identify when par- conducting any work involving human subjects. In
ticipants are giving garbage answers; and other countries, similar requirements may apply.
■ asking participants to successfully complete a In all cases, consulting with an expert on
training task before proceeding to the record- experiments can be invaluable. This will help with
ed trials. design and in applying appropriate statistical
analyses to study the experimental results.
Because of the significant costs associated with
running an experiment, it’s often valuable to
conduct a pilot study with one or two viewers.
This allows testing and refining the experimental Reference
design before starting a full-fledged study with 1. J.A. Swets and R.M. Pickett, Evaluation of Diagnostic Sys-
numerous participants. tems: Methods from Signal Detection Theory, Academic
A wide range of experimental methods may be Press, 1982.

transformed by neural connections in the visual cortex that simultaneous contrast (the phenomenon by which
into three opponent color channels: a luminance chan- perceived color is affected by surrounding colors) occurs
nel (black–white) and two chromatic channels in all three opponent channels. This can cause large errors
(red–green and yellow–blue). The luminance channel when viewers try to read values in the data based on color.
conveys the most information, letting us see form, shape, We can use these theories to draw some conclusions
and detailed patterns to a much greater extent than the regarding the design of color sequences:
chromatic channels. Perception in the chromatic chan-
nels tends to be categorical. That is, we tend to place col- ■ If we want our color sequence to reveal form (such as
ors into categories like red, green, yellow, and blue. local maxima, minima, and ridges), or if we need to
However, we see hues such as turquoise or lime green display detailed patterns, then we should use a
more ambiguously. Another relevant theoretical point is sequence with a substantial luminance component.

IEEE Computer Graphics and Applications 21


Visualization Viewpoints

(a) (b) (c)


2 Three color sequences: (a) a chromatic sequence, good for representing categories; (b) a luminance sequence,
good for representing form; and (c) a combined chromatic-luminance sequence, good for representing both cate-
gories and form.

Shape from texture


Numerous applications in scien-
tific visualization involve the com-
putation and display of arbitrarily
shaped, smoothly curving surfaces.
3 Four exam- A common case is level surfaces in
ples from a volume data. By default, the stan-
study testing dard practice is to render these sur-
different meth- faces with a smooth, Phong-shaded
ods to enhance finish. One important question that
shape percep- arises is, Can we better convey the
(a) (b)
tion: (a) Phong 3D shape by rendering the surface
shading, as if it were made from a subtly tex-
(b) one princi- tured material, rather than polished
pal direction, plastic? Ample evidence from psy-
(c) two princi- chophysics3 suggests that certain
pal directions, kinds of surface texture can facili-
and (d) a line tate shape perception (see Figure
integral 3).4 Unfortunately, the exact mech-
convolution. anisms by which surface texture
affects shape perception—and
hence the specific characteristics of
texture patterns that best show
shape—remain unknown. Compli-
(c) (d)
cating any naive attempt to use tex-
ture to enhance shape appearance is
the complementary evidence that
■ If we want to display categories of information—for under many conditions texture can camouflage surface
example, the classification of a terrain into regions of shape features.5
different geological type—then we should use a chro- Through carefully designed experiments, it’s possi-
matic sequence. ble to gain concrete insights into how we might use tex-
■ If we want to minimize errors from contrast effects, ture most effectively to support accurate shape
then we should arrange a sequence to cycle through perception. More specifically, we can start to answer the
many colors. question, If we want to design the ideal texture that best
conveys the shape of a smoothly curving surface, what
We can also construct a general solution that cycles should the characteristics be? Visualization researchers’
through many colors (to allow categorization) while user studies are essential in this endeavor for several
continuously increasing luminance. reasons.
Figure 2 illustrates three different color sequences First, traditional vision researchers are primarily con-
selected to emphasize a different aspect of the underly- cerned with elucidating the neural processes involved
ing data. Experimental studies have verified that these in the perception of shape from texture, and their inves-
theoretical predictions apply in the case of color tigations don’t fully encompass the scope of questions
sequences.2 This demonstrates the use of well- that we’d like to ask.
established theories to build design guidelines, togeth- Second, there’s a limit to the depth of understanding
er with experiments that validate the guidelines in an we can derive purely from introspection and informal
applied setting. empirical comparison. In the absence of a clear task, view-

22 July/August 2003
ers may adopt differing opinions about which textures 4 Perceptual texture elements
are most effective. Without concrete experimental evi- (pexels) used to visualize a typhoon
dence, it may be impossible to sort out these differences. striking the island of Taiwan: pexel
Furthermore, complex problems rarely yield simple height represents wind speed
answers. If texturing can help, it’s unlikely that any (taller for stronger winds), density
method we initially attempt will turn out to be the best represents pressure (denser for
in all cases. We expect to discover complicated interac- lower pressure), and color repre-
tions between surface texture and shading, between tex- sents precipitation (blue and green
ture orientation and surface geometry, and between for light rainfall to purple and red
aesthetics and convention. We may also find numerous for heavy rainfall; yellow indicates
task dependencies. This suggests that we’ll need to iter- an unknown rainfall amount).
ate to achieve progressively more effective methods for
different purposes. These goals are best achieved
through carefully controlled, quantitative user studies ■ target patch size (the number of pexels used as tar-
that objectively assess the impact of particular texture gets), and
pattern characteristics on the accuracy of performance ■ background texture pattern (whether nontarget tex-
on specific tasks. ture properties were held constant or varied ran-
domly).
Perceptual textures
One key issue we must address when we design an Each condition served a specific function. Target type
experiment is which conditions to study. As the number let us test three different texture dimensions.
of conditions (and the interaction between conditions) Target–background pairing searched for differences in
grows, so does the number of trials needed to test each performance based on the target dimension’s value.
condition properly. Therefore, we often restrict experi- Display duration measured the time needed to perform
ments to the most important conditions. a target detection task. Target patch size asked whether
Understanding how we see the basic properties of an smaller texture patches were harder to identify. Finally,
image lets us create representations that take advantage background texture pattern tested for visual interfer-
of the human visual system. An important discovery in ence when secondary texture dimensions varied ran-
psychophysics from the past 25 years is that human domly across the display. Even these basic conditions
vision doesn’t resemble the largely passive process of produced 108 different display types (three target types
modern photography. A much better metaphor is a by two target–background pairings by three display
dynamic and ongoing construction project, in which the durations by two patch sizes by three background pat-
products are short-lived models of the external world terns). Each viewer who participated during the exper-
specifically designed for the viewer’s current visual iment observed 576 trials from one target type (16
tasks. Harnessing human vision for visualization there- repetitions of a target’s 36 different display types). We
fore requires that we construct images that draw atten- randomly selected eight trials (from the 16 repetitions)
tion to their important parts. in each display type to contain a target patch; the
Previous work in computer vision and psychophysics remaining eight did not.
decomposed texture patterns into a number of basic tex- Results from the experiment showed a preference for
ture dimensions like size, contrast, regularity, and direc- target type (taller targets were easier to identify than
tionality. Based on this, we wondered whether we could shorter, denser, and sparser targets, which were them-
use individual texture dimensions to display multiple selves easier to identify than irregular or regular tar-
attribute values. Controlled experiments offer a way to gets). High accuracy was possible for many target types,
answer this question. even for display durations of 150 ms or less. Finally, vari-
We showed viewers regularly spaced 20 × 15 arrays of ations in regularity interfered with the identification of
perceptual texture elements (or pexels) that look like shorter, sparser, and denser targets (but not taller ones).
upright paper strips. The pexels allow multiple texture A complete description of the experiment’s results is
dimension variations including height, density, and reg- available in Healey and Enns.6 We applied these results
ularity of placement. Viewers saw the pexel grid for a as guidelines for using texture in multidimensional visu-
short duration. We then asked whether a group of pex- alizations. Figure 4 shows an example of using pexels to
els with a particular target value was present or absent. visualize typhoon activity in southeast Asia.
Our experiment tested five different conditions selected
from models of human vision and from texture segmen- Usability testing
tation and classification experiments in computer vision. We designed much of the work presented in this arti-
We varied cle to test basic perceptual features or visualization tech-
niques. We’ve found, however, that visualization
■ target type (target pexels were defined by height, den- applications have important aspects that we should
sity, or spatial regularity), study within the application’s context.
■ target–background pairing (different types of tar- The approach to this type of study is quite different
gets—for example, both medium and tall targets), from basic perception experiments. Participants must
■ display duration (the amount of time the viewer saw solve a relatively complex task, where there’s a greater
the pexel array), freedom of actions and a higher potential for mistakes.

IEEE Computer Graphics and Applications 23


Visualization Viewpoints

cluded there were two problems with the study. First,


the maps we used were visually too simple. Second, the
number of tasks was too small; more examples per user
5 An image might lead to significant results. We plan to consider
from the these ideas in future work on SDOF.
semantic depth
of field study. When do user studies help?
The image was While user studies are an important tool for visual-
displayed for ization design, they aren’t the proper choice in every sit-
200 ms, after uation. Experiments don’t always work as expected and
which partici- other techniques are available.
pants were
asked to point Other techniques
at the quadrant It’s important to consider other options before design-
with the sharp ing and running a user study. Studies are time consum-
object. ing to design, implement, run, and analyze. Typically,
we can only use them to answer small questions, and
any larger conclusions rely on generalizations that might
not be valid. Often, measures that are less precise, quan-
titative, and objective may provide sufficient insight
about a visualization question to let us move forward.
Studying a technique in an application setting (as In our investigation of virtual reality tools for archae-
opposed to an artificially simple environment) is criti- ological analysis,8 we labored long and hard to design
cal because we can’t assume that low-level results auto- a good user study to test the system we developed.
matically apply to more complex displays. However, the experimental design eluded us. In the end,
Comments from participants are often more impor- we videotaped a pair of archaeologists using the system
tant than the other data we collect because they provide to evaluate some of their scientific hypotheses. They also
valuable hints about what’s happening during the exper- generated several new ideas, some of which would have
iment. Close observation of the participants can also been difficult to generate with other analysis methods.
offer information about experiment details that possibly This approach was sufficient to demonstrate the visual-
weren’t part of the original hypotheses. ization application’s use.
An example of this type of study is the evaluation of In another context, we can also transcend the tradi-
semantic depth of field (SDOF),7 a technique for guid- tional user study. Artists and designers have been cre-
ing a viewer to specific information in an image. SDOF ating visualizations for centuries and have invented
is based on the depth-of-field effect from photography, effective methods. User studies come from science—in
where different parts of a picture are in or out of focus fact, they embody the scientific method of posing
based on their distance from the focal point of the lens. hypotheses, taking measurements, analyzing them, and
SDOF generalizes this concept. An object’s sharpness iterating to gain insight. For the scientific study of low-
depends not on its physical position, but on its relevance. level vision, the methodology works, but as we rise up
Viewers are immediately drawn to the sharp (that is, to the level of a scientific visualization application, it
highly relevant) parts of the image, but they can still might not be possible to use these techniques to answer
choose to look at other, out-of-focus objects (see Figure important questions.
5). We designed an experiment that contained both Can we replace some parts of user testing with expert
basic perception and application components. The per- visual designers? This is a conjecture we can likely test
ception studies produced significant results, which were (not surprisingly) with a user study, comparing results
close to what we expected to find. The application find- of a standard user study with expert visual designer
ings, however, were much less conclusive. input. Preliminary results suggest that visual designers
One application was a map viewer that presented can replicate some user study results more quickly and
users with a map containing nine layers of information with more insight about why differences occur.
(for example, roads, elevations, and cities). We asked However, we still have much to learn about the space
them to position a project (for example, a factory) based between perceptual psychology and visual design.
on three very important and three somewhat important
factors. Users could reorder the layers by selecting which When things go wrong
layer was on top. The layers were displayed in three dif- In some studies, experimental design may lead to
ferent ways: opaque, semitransparent, and SDOF (the results that aren’t statistically significant. For example,
top layer was sharp and underlying layers were increas- in a recent study we hypothesized that users would per-
ingly blurred). The hypothesis was that SDOF would form differently for a visual search task in virtual reali-
make it easier to stack the layers in order of importance, ty if the virtual environment were different. In fact, we
and thus to answer more quickly and correctly. found that statistically there was no significant differ-
While some useful results were identified during the ence. Perhaps our conjecture was wrong, but it’s also
application study, we didn’t find statistically significant possible that our choice of task or other parts of the
results in either response time or correctness. We con- experimental design misled us. The virtual environment

24 July/August 2003
may really matter in some cases. We continue to think methods to select the best tool for the problem at hand.
about how the virtual environment might make a dif- One reason visualization is such a fascinating part of
ference, particularly since visual context is important in computer science is because so many other fields (such
2D visual search tasks. Some studies aren’t published as psychology and the visual arts) overlap with our
because of null results, or because the results are incon- research. ■
clusive or uncompelling.
Null results are completely natural because they show Acknowledgments
that the original hypothesis wasn’t supported by the We’d like to thank Helwig Hauser, who played a key
data. This can be because the difference is too small for role in proposing the idea for this article.
the amount of data collected, but most often it’s because We completed this work in part as a component of the
the hypothesized difference is insignificant. This is why basic research on visualization ([Link]
the study was done in the first place and it should there- vis/) at the VRVis Research Center in Vienna, which is
fore not be considered a failure. In visualization, we can’t funded by the Austrian research program Kplus. The US
publish null results (at least not on their own) easily. National Science Foundation also supported our work
Nevertheless, the results can provide insight about (ACI-0083421, CCR-0086065, CCR-0093238).
which directions of research to pursue and which to
abandon.
Inconclusive results are a much more serious prob-
lem. They usually mean that there was a design error in References
the study and that it must be run again. Usually, how- 1. D.H. Laidlaw et al., “Quantitative Comparative Evaluation
ever, this affects only one part of a study, so the effort is of 2D Vector Field Visualization Methods,” Proc. IEEE Visu-
considerably smaller the second time. Also, we can test alization 2001, IEEE CS Press, 2001, pp. 143-150.
additional hypotheses emerging from the successful 2. C. Ware, “Color Sequences for Univariate Maps: Theory,
parts of the study. Experiments and Principles,” IEEE Computer Graphics and
Uncompelling results can result from choosing the Applications, vol. 8, no. 5, Sept./Oct. 1998, pp. 41-49.
wrong task or measuring the wrong performance quan- 3. J.T. Todd and F.D. Reichel, “Visual Perception of Smooth-
tity. For example, in “The Great Potato Search,”9 we ly Curved Surfaces from Double-Projected Contour Pat-
chose a 3D visual search task. Unfortunately, it was a terns,” J. Experimental Psychology: Human Perception and
task that involved looking inward at a relatively small Performance, vol. 16, no. 3, 1990, pp. 665-674.
model. We believe a task that involved searching more 4. S. Kim, H. Hagh-Shenas, and V. Interrante, “Showing
broadly around the user might have shown important Shape with Texture: Two Directions are Better than One,”
performance differences correlated with changes in the Human Vision and Electronic Imaging VIII, SPIE, no. 5007,
virtual context. While we can (and will) go on to test July 2003.
that new hypothesis, if we had chosen a different task 5. J.A. Ferwerda et al., “A Model of Visual Masking for Com-
in the first place, we would have been better off. There’s puter Graphics,” Proc. Siggraph, ACM Press, 1997, pp. 143-
always a tension between executing an experiment 152.
quickly and spending time on design. Practice can help 6. C.G. Healey and J. T. Enns, “Large Datasets at a Glance:
reduce or alleviate these types of mistakes. Combining Textures and Colors in Scientific Visualization,”
IEEE Trans. Visualization and Computer Graphics, vol. 5,
Conclusions no. 2, Apr.–June 1999, 145-167.
In this article we tried to advance the current state of 7. R. Kosara et al., “Useful Properties of Semantic Depth of
the art in two ways: Field for Better F + C Visualization,” Proc. Joint Euro-
graphics and IEEE Trans. Visualization and Computer Graph-
■ Promote evaluating visualization methods with user ics Symp. Visualization (VisSym 2002), IEEE CS Press,
studies. This is being done in certain cases, but it’s still 2002, pp. 205-210.
far from standard practice in our field. 8. E. Vote et al., “Discovering Petra: Archaeological Analysis
■ Ask where user studies might be useful and where in VR,” IEEE Computer Graphics and Applications, vol. 22,
other techniques might be more appropriate (such as no. 5, Sept./Oct. 2002, pp. 38-50.
ideas from the visual arts). 9. C.D. Jackson et al., “The Great Potato Search: The Effects
of Visual Context on Users’ Feature Search and Recogni-
User studies can improve the quality of our research. tion Abilities in an IVR Scene,” Poster Proc. IEEE Visualiza-
Although it’s difficult to design a good experiment and tion 2002, [Link]
the relevant skills require substantial study tempered pdf/Jackson:2002:[Link].
with experience, a well-conducted study is usually
worth the effort. The results can ultimately have a con-
siderable impact and potentially contribute to the dis- Readers may contact Robert Kosara by email at
cipline’s scientific foundations. Kosara@[Link].
Even though we advocate more user studies, we rec-
ognize that other methods may be more appropriate in Contact department editor Theresa-Marie Rhyne by
certain situations. Designers should be aware of these email at tmrhyne@[Link].

IEEE Computer Graphics and Applications 25

You might also like