0% found this document useful (0 votes)
3 views16 pages

Vuo S Koski 2016

The study investigates the interaction of visual and auditory cues in the perception of musical performance, highlighting the importance of visual kinematic information in shaping emotional reactions and evaluations of expressivity. Experiments reveal that visual cues can significantly influence observers' subjective emotional responses and perceptions of loudness, though their effect on tempo variability remains unclear. The findings suggest that both sight and sound play crucial roles in the overall experience of music, challenging the traditional emphasis on auditory elements alone.

Uploaded by

dfm7txjywr
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views16 pages

Vuo S Koski 2016

The study investigates the interaction of visual and auditory cues in the perception of musical performance, highlighting the importance of visual kinematic information in shaping emotional reactions and evaluations of expressivity. Experiments reveal that visual cues can significantly influence observers' subjective emotional responses and perceptions of loudness, though their effect on tempo variability remains unclear. The findings suggest that both sight and sound play crucial roles in the overall experience of music, challenging the traditional emphasis on auditory elements alone.

Uploaded by

dfm7txjywr
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Interaction of Sight and Sound in the Perception and Experience of Musical

Performance
Author(s): Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence and Eric F. Clarke
Source: Music Perception: An Interdisciplinary Journal , Vol. 33, No. 4 (APRIL 2016), pp.
457-471
Published by: University of California Press

Stable URL: [Link]

REFERENCES
Linked references are available on JSTOR for this article:
[Link]
reference#references_tab_contents
You may need to log in to JSTOR to access the linked references.

JSTOR is a not-for-profit service that helps scholars, researchers, and students discover, use, and build upon a wide
range of content in a trusted digital archive. We use information technology and tools to increase productivity and
facilitate new forms of scholarship. For more information about JSTOR, please contact support@[Link].

Your use of the JSTOR archive indicates your acceptance of the Terms & Conditions of Use, available at
[Link]

University of California Press is collaborating with JSTOR to digitize, preserve and extend
access to Music Perception: An Interdisciplinary Journal

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 457

I N T E R AC T I O N OF SIGHT AND SOUND IN THE PERCEPTION AND EXPERIENCE OF


MUSICAL PERFORMANCE

M
J O N NA K. V U O S KO S K I USIC IS AN INHERENTLY MULTISENSORY
University of Oxford, Oxford, United Kingdom phenomenon, comprising auditory, visual,
and somatosensory components. In musical
M A R C R. T H O M P S O N performance, a performer’s body movements and ges-
University of Jyväskylä, Jyväskylä, Finland tures can convey a range of meaningful information to
audiences and co-performers alike, including emotional
C HA R L E S S P E N C E , & E R I C F. C L A R K E expression (Castellano, Mortillaro, Camurri, Volpe, &
University of Oxford, Oxford, United Kingdom Scherer, 2008; Dahl & Friberg, 2007; Davidson, 1993,
1994) and phrasing (Juchniewicz, 2008; Vines, Krum-
RECENTLY, VUOSKOSKI, THOMPSON, CLARKE, AND hansl, Wanderley, & Levitin, 2006), as well as musical
Spence (2014) demonstrated that visual kinematic ideas and timing (Glowinski et al., 2013; Goebl &
performance cues may be more important than audi- Palmer, 2009; Williamon & Davidson, 2002). The
tory performance cues in terms of observers’ ratings of salience of visual kinematic information (i.e., visual
expressivity perceived in audiovisual excerpts of piano information about performers’ body movements and
playing, and that visual kinematic performance cues gestures) for an observer’s perception and experience
had crossmodal effects on the perception of auditory of a musical performance has been widely documented
expressivity. The present study was designed to extend (e.g., Chapados & Levitin, 2008; Davidson, 1993; Tsay,
these findings, and to provide additional information 2013; Vines, Krumhansl, Wanderley, Dalca, & Levitin,
about the roles of sight and sound in the perception 2011; Vines et al., 2006), and a recent meta-analysis
and experience of musical performance. Experiment 1 by Platz and Kopiez (2012) revealed that, compared
investigated the relative contributions of auditory and to audio-only presentations, audiovisual information
visual kinematic performance features to participants’ has a moderate effect on participants’ evaluations of
subjective emotional reactions evoked by piano perfor- a musical performance.
mances, while Experiment 2 was designed to explore Although it has been established that visual informa-
the effect of visual kinematic cues on the perception of tion about the performer’s movements consistently
loudness and tempo variability. Experiment 1 revealed enhances the appreciation of a musical performance
that visual performance cues seem to be just as impor- (Platz & Kopiez, 2012), previous studies have not reli-
tant as auditory performance cues in terms of the ably estimated the relative contributions of visual and
subjective emotional reaction of the observer, thus auditory performance cues to observers’ experience.
highlighting the importance of non-auditory cues for Although previous investigations have shown that the
music-induced emotions. The results of Experiment 2 effect size of the visual component on observers’ evalua-
revealed that visual kinematic cues only affected rat- tions could on average be characterized as ‘‘medium’’
ings of loudness variability, but not ratings of tempo (Platz & Kopiez, 2012), it is not known how that relates
variability. to the effect size of auditory performance cues – espe-
cially across different levels of expressivity. Variations in
Received: November 26, 2013, accepted February 20, performance features – often collectively referred to as
2015. ‘‘expressivity’’ – are what differentiate performances of
the same notated work, and serve to articulate musical
Key words: Multisensory integration, piano perfor-
structure (Clarke, 1988), communicate emotional
mance, expressivity, music-induced emotion, audio-
meaning (see Juslin, 2001, for a review), and convey
visual perception
a sense of biological motion (Juslin, 2003). In order to

Music Perception, VOLUM E 33, ISSU E 4, PP. 457–471, IS S N 0730-7829, EL ECTR ONI C ISSN 1533-8312. © 2016 B Y THE R E GE N TS OF THE UN IV E RS I T Y O F CA LI FOR NIA A LL
R IG HTS RES ERV ED . PLEASE DIR ECT ALL REQ UEST S F OR PER MISSION T O PHOT O COPY OR R EPRO DUC E A RTI CLE CONT ENT T HRO UGH T HE UNI VE R S IT Y OF CALI FO RNIA P R E SS ’ S
R EPRIN TS AND P ERMISSI ONS WEB PAG E , HT T P :// W W W. UC PRESS . EDU / JO URNA LS . PHP ? P ¼R E P RI NTS . DOI: 10.1525/ M P.2016.33.4.457

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
458 Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence, & Eric F. Clarke

investigate the relative contributions of auditory and explored potential crossmodal effects in the perception
visual performance cues to observers’ evaluations, an of auditory and visual expressivity, addressing the ques-
experimental method is needed in which the expressiv- tion of whether simultaneously presented visual kine-
ity conveyed by the visual and auditory components of matic information might alter the way in which auditory
a performance can be manipulated independently, so as expressivity is perceived, or vice versa. They found that
to result in matched and mismatched audiovisual pair- relative to auditory cues, visual kinematic cues actually
ings. Such experimental designs have been successfully contributed slightly more to a participant’s overall eval-
used to investigate the interaction of auditory and visual uation of expressivity, and that there appeared to be
components in the perception of note duration (Schutz crossmodal interactions at play in the evaluation of both
& Lipscomb, 2007), loudness (Rosenblum & Fowler, auditory and visual expressivity.
1991), timbre (Saldaña & Rosenblum, 1993), pitch Although Vuoskoski et al.’s (2014) study provides
(Thompson, Graham, & Russo, 2005), and interval preliminary evidence for the existence of crossmodal
affect (Thompson, Russo, & Quinto, 2008), demonstrat- effects in the evaluation of expressivity – as well as shed-
ing that visual information can significantly alter the ding light on the relative salience of auditory and visual
perception of various auditory features. However, the kinematic performance cues – there are some limitations
difficulty with applying such a design to a complex and questions that require further investigation. First,
action such as musical performance is that musicians when considering the relative importance of auditory and
find it very difficult to control expressivity in one visual cues from the observer’s point of view, the evalu-
modality independent of the other (e.g., Thompson & ation of perceived expressivity may not capture the most
Luck, 2012), and the temporal properties of a musical salient or essential aspects of an observer’s experience of
performance also vary greatly from one performance to a musical performance. Instead of the objective appraisal
the next. of the expressive components of a musical performance,
Previous studies have attempted to tackle this issue by it is arguably an observer’s subjective emotional experi-
combining a constant auditory stimulus with visual ence of the performance that ultimately determines their
information of actors portraying different expressive evaluation (cf. Hargreaves & North, 2010). Although
intentions (e.g., Juchniewicz, 2008; Morrison, Price, there is some evidence to indicate that visual information
Geiger, & Cornacchio, 2009), or have settled for com- might enhance emotional reactivity to musical perfor-
bining structurally incongruent auditory and visual mances (Chapados & Levitin, 2008), it is not yet known
components, thus resulting in functionally incongruent how the effect of visual performance cues relates to that
and temporally asynchronous stimuli (e.g., Krahé, of auditory performance cues with regard to the subjec-
Hahn, & Whitney, 2013; Petrini, McAleer, & Pollick, tive emotional reaction of the observer. Furthermore, the
2010). The former approach is problematic because of explicit instructions used by Vuoskoski et al. to take both
the limited validity of ‘‘faked’’ expressive movements, auditory and visual aspects of the performance into
and the latter because the movements and gestures in account in the evaluations of overall expressivity might
musical performance have been found to arise from have affected which aspects of the material the partici-
a representation of the musical structure, and thus con- pants attended to (for details, see Vuoskoski et al., 2014;
vey meaning in association with specific musical pas- Experiment 1). In other words, it may be that as a result
sages (e.g., MacRitchie, Buck, & Bailey, 2013). of the instructions, the participants paid more attention
To address these limitations, a recent study by to visual kinematic performance features than they oth-
Vuoskoski et al. (2014) presented a novel method for erwise would.
creating matched and mismatched audiovisual combi- Second, although the study by Vuoskoski et al. (2014,
nations of different expressive intentions. By utilizing Experiment 2) demonstrated that visual kinematic cues
motion-capture animations of piano performances can have an impact on evaluations of auditory expres-
and time-warping algorithms, Vuoskoski et al. were sivity, the exact nature of these crossmodal effects
able to investigate the relative contributions of auditory remains unclear. It is not yet established whether there
and visual kinematic performance cues to the perception are crossmodal effects at play in the perception of lower-
of expressivity in a systematic and balanced way. In con- level auditory features such as, for example, loudness.
trast to previous studies, the mismatched stimuli utilized Furthermore, it is possible that the outcome reflects
by Vuoskoski et al. were temporally synchronized and response bias, the participants’ ratings of auditory
structurally congruent (i.e., the visual kinematic infor- expressivity being affected by the expressive qualities
mation always represented the same composed structure of the simultaneously presented visual kinematic infor-
as the auditory information). Vuoskoski et al. also mation without their perception of the auditory features

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 459

actually having been affected (see, e.g., Schutz & By comparison, it is less obvious how temporally
Kubovy, 2009). aligned visual kinematic information could affect the
The aim of the present study was therefore to extend auditory perception of tempo variability. Previous
the findings of Vuoskoski et al. (2014), and to provide research has shown the temporal resolution of the audi-
new information regarding the roles of visual kinematic tory modality to be superior to that of the visual modal-
and auditory cues in the subjective emotional reactions ity (e.g., Burr, Banks, & Morrone, 2009; Freides, 1974;
evoked by musical performance, as well as investigating Repp & Penel, 2002), resulting in superior auditory
the possible effect of visual kinematic cues on the eval- rhythm and beat perception (e.g., Grahn, 2012). How-
uation of specific auditory performance features. Exper- ever, previous research has also shown that visual kine-
iment 1 was designed to investigate the relative matic information can influence the perceived duration
contributions of auditory and visual kinematic perfor- of notes played on a marimba (Schutz & Kubovy, 2009;
mance features on participants’ subjective emotional Schutz & Lipscomb, 2008), and that the sensitivity to
reactions, and thus to provide a more ecologically rhythmic deviations can be modulated by point-light
relevant account of the roles of sight and sound in an animations of a bouncing person (Su, 2014). Neverthe-
observer’s experience of a musical performance. The less, since the movements of the pianists were tempo-
difference between the previous Experiment 1 reported rally synchronized with the music in all of our stimuli,
by Vuoskoski et al. (2014) and the current Experiment 1 we hypothesize that the visual kinematic information
mirrors the well-established distinction between per- will have an effect on the perception of loudness vari-
ceived and felt emotion (see, e.g., Sloboda & Juslin, ability but not on the perception of tempo variability.
2010). The former experiment investigated evaluations
of a perceived characteristic of the performances (i.e.,
Experiment 1
perceived expressivity), while the current experiment
investigates the subjective emotional reactions experi-
METHOD
enced by participants (i.e., felt emotion). Previous
Participants. Nineteen participants (7 male, 12 female)
research has suggested that visual information may
aged 18-31 years (M ¼ 23.1, SD ¼ 4.1) were recruited
have a significant impact on an observer’s emotional
from the University of Oxford community. Fourteen
reactions to a musical performance (Chapados & Levi-
participants (73.7%) reported having received at least
tin, 2008; Krahé et al., 2013; Vines et al., 2006), but
some music training on an instrument (ranging from 1
the effect size of visual kinematic performance cues
to 18 years; M ¼ 10.6, SD ¼ 5.0). The participants
relative to that of auditory performance cues has yet
received a monetary incentive (£5) for taking part in
to be investigated.
the study. All of the experimental procedures followed
The aim of Experiment 2 was to explore the effect of
the University of Oxford Policy on the Ethical Conduct
visual kinematic cues on the evaluations of auditory
of Research Involving Human Participants and received
expressivity in more detail. The two main auditory
approval from the Research Ethics Committee.
characteristics contributing to expressivity in piano
performance are variations in timing and dynamics Stimuli. The stimuli were obtained from a recent study
(i.e., tempo and loudness variation; e.g., Gabrielsson, by Vuoskoski et al. (2014), where the stimulus genera-
1999; Palmer, 1997), with the amount of variation tion process is reported in some detail. However, the
being positively associated with perceived expressivity process is briefly outlined here, as the method of stim-
(e.g., Bhatara, Tirovolas, Marie Duan, Levy, & Levitin, ulus generation is crucial for the questions addressed in
2011). Perceived expressivity is also positively associ- the study. Two pianists – one male and one female –
ated with how much a performer moves (e.g., David- performed Chopin’s Prelude in E minor (Op. 28, No. 4)
son, 1994; Thompson & Luck, 2012). Since the size of with three different levels of expression: Deadpan
a performer’s movements reflects the physical energy (reduced level of expressive intensity); Normal (normal
used to play the notes, the kinematic visual informa- level of expressive intensity); and Exaggerated (maxi-
tion specifying a performer’s movements might be mum level of expressive intensity); while their move-
expected to affect the perception of loudness, which ments were captured at 120 frames per second using an
is directly related to physical energy. Visual kinematic 8-camera optical motion capture system (Qualisys Pro-
cues have previously been shown to affect loudness Reflex). In addition, the MIDI output of the digital
perception in the context of simple clapping sounds, piano keyboard used in the performances was recorded,
with the size of clapping motions positively associated providing a complete record of the performances. To
with perceived loudness (Rosenblum & Fowler, 1991). create the audio stimuli, the MIDI data were imported

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
460 Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence, & Eric F. Clarke

TABLE 1. Mean Tempo, Tempo Variation, Mean Root-mean-square Energy, and Total Amount of Movement in the Six Performance Excerpts

Performance type Mean tempo (bpm) Tempo variability (%) Mean RMS (SD)* Amount of movement (m)
Pianist 1 Deadpan 65.48 5.84 2.05 (0.79) 15.69
Normal 62.22 15.08 2.45 (1.22) 36.70
Exaggerated 58.87 17.31 3.23 (1.50) 44.03
Pianist 2 Deadpan 59.29 8.24 1.72 (0.73) 18.27
Normal 58.72 16.18 2.23 (1.30) 29.61
Exaggerated 57.39 24.26 2.32 (1.54) 33.36
Tempo variability reflects the standard deviation of the divergence from mean tempo, calculated for each eighth note. Root-mean-square energy reflects the mean loudness
(and loudness variability) of the audio excerpts. RMS values were calculated for 500 millisecond segments. Amount of movement indicates the total distance travelled by the
motion capture markers. *RMS values and standard deviations have been multiplied by 1000.

into GarageBand ’11 (version 6.0.5), running on Mac


OS X. The ‘‘Grand Piano’’ software instrument with
50% reverb was used to generate high-quality renditions
of the performances. The segment from the beginning
of measure 13 to the end of measure 20 was used to
create the experimental stimuli, as this section includes
the expressive climax of the piece (Sloboda & Lehmann,
2001), and should therefore allow for the greatest
amount of variation in terms of expressive intensity. The
duration of the resulting six performance excerpts (2
performers x 3 expressive intentions) ranged from 29
to 33 s (M ¼ 31.3, SD ¼ 1.5). The descriptive details of
the performance excerpts (mean tempo, mean loudness,
tempo and loudness variability, and the total amount of
movement) are displayed in Table 1.
In order to generate audiovisual stimuli that would be
incongruent in terms of their expressive intention (e.g., FIGURE 1. A sample frame of the point-light animations used in
deadpan audio þ normal movement) yet temporally Experiments 1 and 2.

synchronized, the motion capture data from each per-


former were temporally aligned to each of the three unaltered. Finally, the resulting splines were sampled
audio tracks of that performer using a time-warping to create time-warped motion capture data that could
algorithm (Verron, 2005; see also Wanderley, Vines, be used to generate point-light animations. This method
Middleton, McKay, & Hatch, 2005). This procedure has previously been used for analysis purposes (see
involved the generation of timing profiles for each per- Wanderley et al., 2005), as it enables the comparison
formance by annotating the timing of each eighth-note of different performances independent of original
chord played by the left hand, producing an average tempo or timing variations. However, the present study
resolution of 2.04 time points per s. The motion capture (and the previous one by Vuoskoski et al., 2014) are – to
data was then functionalized using cubic splines. Using the best of our knowledge – the first to use the method
the annotated timing profiles for each performance, to generate time-warped point-light animations.
curve-stretching algorithms (see Verron, 2005, for Point-light animations of the original and time-warped
details) were used to stretch and compress the motion motion capture data were generated using MATLAB and
capture data of a given performance so that it matched the Motion Capture Toolbox (Burger & Toiviainen,
the timing profile of another performance. More specif- 2013). Light points – connected by lines to form a stick-
ically, the splines between each note onset were made to figure shape – represented each pianist’s hands, wrists,
match the time separation of the corresponding note elbows, shoulders, head (midpoint and four markers
onsets in the other performance. Two time-warped ver- around the head), torso (mid-shoulder and mid-torso),
sions of each performance were generated to match the and hips. The keyboard was represented by two markers
timing profiles of the other two performances by the connected by a line (see Figure 1 for a sample frame).
same performer. Note that only the movement data The 18 animations were combined with the appropri-
were time-warped while the audio data remained ate audio to create 6 matching (e.g., normal audio þ

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 461

normal video) and 12 mismatching (e.g., exaggerated time-warped animations that had been altered to match
audio þ deadpan video) audiovisual stimuli (example the different audio tracks. The order in which the two
stimuli can be downloaded from [Link] unimodal blocks were presented was counterbalanced
[Link]/u/311821/Video_examples.zip). Note that across participants. Again, the excerpts within the
the audio and video from different performers were blocks were presented in a different random order for
never combined. In addition, unimodal versions of the each participant. After the experiment, the participants
stimuli (6 audio-only and 18 video-only stimuli) were completed a short questionnaire about their music
also generated. training and music listening habits, and were fully
debriefed.
Procedure. The Max/MSP (version 5.1.9; Cycling ‘74)
graphical programming environment (running on Mac RESULTS
OS X) was used to present the stimuli and to collect the Emotional impact in unimodal rating conditions. In
data. The animations were presented with a resolution order to investigate whether the unimodal (audio-only
of 800 x 600 pixels and a frame rate of 30 fps. The audio and video-only) representations of different expressive
was presented in WAV format through high quality intentions resulted in differing evaluations of felt emo-
headphones (Sennheiser HD 219). The participants tional impact, repeated-measures ANOVAs were con-
were instructed to evaluate the intensity of their subjec- ducted to investigate the ratings obtained in the two
tive emotional reactions to the performances, and were unimodal conditions. There were two within-
informed that a given performance might leave them participant factors; Performance Condition (Deadpan,
cold, while another performance might move them in Normal, or Exaggerated) and Pianist (Pianist 1 or 2), and
a profound way. The evaluations were made using a hor- one between-participants factor; Block Order. The latter
izontal analog scale (width 278 pixels) ranging from factor was added in order to investigate whether the
‘‘did not move me at all’’ to ‘‘moved me very strongly.’’ presentation order of the unimodal blocks (audio-only
The participants were instructed to base their ratings on first or video-only first) had any effect on participants’
their own emotional reactions rather than any specific ratings. Note that the audiovisual block always preceded
aspect of the performances (such as the auditory or the two unimodal blocks. In the audio-only condition,
visual components of the stimuli), so as not to direct there was a significant main effect of Performance Con-
the participants towards perceived rather than felt emo- dition; F(2, 34) ¼ 7.07, p < .01, G2 (generalized eta
tion. The output of the scale, as a default property of the squared; Bakeman, 2005) ¼ .17, as well as a significant
Max/MSP-object, provided data in the range 0-127. The main effect of Pianist; F(1, 17) ¼ 5.84, p < .05, G2 ¼ .04.
participants were instructed to make their evaluations There was no main effect of Block Order, and no inter-
immediately after each excerpt had ended. action effects. Multiple comparisons of means (paired
The experiment started with two practice trials using t-tests, p < .05 significance level adjusted using the
audiovisual excerpts that were similar to – but not part Holm-Bonferroni method; Holm, 1979) revealed that
of – the actual stimulus set, to which the participants ratings of emotional impact for the Deadpan perfor-
were instructed to respond. These responses were not mances were significantly lower than those for the Nor-
included in the data. The practice trials were followed mal and Exaggerated performances, but that the latter
by the 18 audiovisual excerpts, which were presented in two did not differ significantly from each other. A com-
a different random order for each participant. The parison of means also revealed that the performances of
audiovisual block was followed by two unimodal blocks Pianist 2 were rated as having a stronger emotional
(audio-only, consisting of 6 audio excerpts; and video- impact on average than those of Pianist 1. The mean
only, consisting of 18 video excerpts), in which evalua- ratings for the three different types of performances by
tions of felt emotional impact were based only on what the two pianists are displayed in Figure 2.
was perceived in the presented modality. The audiovi- A similar repeated-measures ANOVA was conducted
sual block was always presented first, as the audiovisual to analyze the ratings of felt emotional impact obtained
condition was the main focus of interest in the current in the video-only rating condition, with the difference
study. Furthermore, the initial exposure to the audiovi- that two factors regarding performance condition were
sual excerpts provided participants with a relevant included: Type of Video, and Type of Time-warp. As
framework in which to view the silent point-light ani- the video component of the mismatched stimuli had
mations, which might have seemed strange or arbitrary been slightly altered to fit the accompanying audio
if presented first. The video-only condition included track, Type of Time-warp was included to determine
both the six original animations as well as the twelve whether there were any differences between the

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
462 Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence, & Eric F. Clarke

FIGURE 2. The mean ratings of emotional impact (+ standard error of the mean) obtained in the unimodal audio-only and video-only conditions of
Experiment 1. The ratings have been scaled to a range of 0-100.

different time-warped and original animations. Type


of Video and Type of Time-warp both had three levels:
Deadpan, Normal, and Exaggerated. Due to a technical
failure, one participant’s video-only ratings were not
recorded, and thus n ¼ 18 for this analysis. The anal-
ysis revealed significant main effects of Type of Video;
F(2, 32) ¼ 29.17, p < .001, G2 ¼ .26, Type of Time-
warp; F(2, 32) ¼ 4.32, p < .05, G2 ¼ .01, and Pianist;
F(1, 16) ¼ 18.04, p < .001, G2 ¼ .07. There were no
main or interaction effects related to Block Order.
Multiple comparisons of means revealed that the emo-
tional impact of the Deadpan video type was rated as
significantly weaker than the impact of the Normal or
Exaggerated video types, as expected; but that the dif-
ference between the latter two – although in the FIGURE 3. The mean ratings of felt emotional impact (+ standard error
expected direction – was not statistically significant. of the mean) obtained in the audiovisual condition of Experiment 1,
grouped by Type of Audio and Type of Video. The ratings have been
Multiple comparisons regarding the effect of Type of
scaled to a range of 0-100.
Time-warp did not reveal any significant differences
between the different time-warped and original anima-
tions after the Holm-Bonferroni correction had been and visual modalities with regard to the emotional
applied. A comparison of means confirmed that the impact induced by the audiovisual performance
emotional impact of the performances by Pianist 1 excerpts, a repeated-measures ANOVA was conducted.
were evaluated as significantly stronger than for those There were three within-participant factors in the
by Pianist 2. There was also a significant interaction ANOVA: Type of Audio, Type of Video, and Pianist.
between Type of Video and Pianist; F(2, 32) ¼ 10.02, The analysis yielded significant main effects of Type
p < .001, G2 ¼ .04. Multiple comparisons of means of Audio: F(2, 36) ¼ 11.22, p < .001, G2 ¼ .10; and Type
revealed that the emotional impact of the perfor- of Video: F(2, 36) ¼ 9.12, p < .001, G2 ¼ .09. The mean
mances by Pianist 1 was rated as significantly stronger ratings (grouped by Type of Audio and Type of Video)
(than those of Pianist 2) only in the Normal and Exag- are displayed in Figure 3. Multiple comparisons of
gerated video conditions. The mean ratings given for means (paired t-tests, p < .05 significance level adjusted
the three different types of performance by the two using the Holm-Bonferroni method) revealed that all
pianists are shown in Figure 2. three types of audio were significantly different from each
other, with the Deadpan condition receiving the lowest
Ratings of emotional impact in the audiovisual condi- and the Exaggerated condition receiving the highest rat-
tion. In order to investigate the salience of the auditory ings of felt emotional impact. Regarding the different

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 463

video types, multiple comparisons revealed that the Vuoskoski et al., 2014), it is nevertheless striking that
Deadpan videos received significantly lower ratings of point-light animations of pianists performing were none-
emotional impact than the Normal and Exaggerated theless able to evoke significantly differentiated emo-
videos, but that the latter two did not differ significantly tional responses in participants. However, it may also
from one other. There was no main effect of Pianist, and be that participants’ evaluations were affected by demand
no interaction effects. characteristics (e.g., Orne, 1962). When asked to evaluate
To further investigate the relative contribution of the emotional impact of stimuli that clearly represent an
auditory and visual cues to the emotional impact evoked emotional expression of some kind, it might be that even
by the performance excerpts, a linear regression analysis in the absence of genuine emotional reactions the parti-
was conducted. The dependent variable was the mean cipants nonetheless move the slider on the basis of per-
ratings of felt emotional impact for audiovisual stimuli, ceived expressivity rather than felt emotion (cf. Konečni,
while the mean ratings of emotional impact for audio- 2008). This possibility is supported by the fact that three
only and video-only stimuli were the independent vari- of the participants reported extremely low (or nonexis-
ables. The two predictor variables were not significantly tent) levels of emotional impact in response to the video-
intercorrelated, r(16) ¼ –.06, ns, but both were signifi- only stimuli (but not in response to the audiovisual or
cantly correlated with the dependent variable: r(16) ¼ audio-only stimuli), perhaps reflecting a more rigorous
.67, p < .01, for audio-only ratings, and r(16) ¼ .62, p < rating strategy on their part than for the other partici-
.01, for video-only ratings. Audio-only and video-only pants. Furthermore, having already responded to an
ratings of emotional impact both significantly predicted audiovisual block (which was always presented first) it
felt emotional impact in the audiovisual condition, ¼ is possible that the participants’ unimodal ratings were
.71, t(17) ¼ 7.94, p < .001, and ¼ .66, t(17) ¼ 7.42, p < influenced by previous audio-visual associations. Since
.001, respectively. Together they explained a significant the participants were exposed to both matched and mis-
proportion of the variance in the emotional impact felt matched combinations in the audiovisual block, it is
in the audiovisual condition; R2 ¼ .88, F(2, 17) ¼ 55.93, unlikely that they would have associated a specific
p < .001. audio-only stimulus with a specific video component or
vice versa; but it may be that a more generic association
DISCUSSION between the two modalities may nonetheless have been
The results of Experiment 1 demonstrate that each induced.
audio type – Deadpan, Normal, and Exaggerated – was The results of the audiovisual rating condition
rated as eliciting a different level of emotional impact in revealed that both Type of Audio and Type of Video had
the audio-only condition. The effect size of audio type a significant effect (G2 ¼ .10 and .09, respectively) on
(G2 ¼ .17) was notably smaller than that in a previous the emotional impact of the piano performances. The
experiment measuring perceived expressivity (using the effect sizes of audio type and video type were compara-
same stimuli; G2 ¼ .59; Vuoskoski et al., 2014). This ble, in contrast to the differences observed in the unim-
difference in effect size may be attributable to the more odal rating conditions. This pattern of results is
subjective and internal character of participants’ own somewhat different from that found for the perception
emotional reactions as compared to the more manifest of expressivity (Vuoskoski et al., 2014), where Type of
and external character of the expressive intentions on Video (G2 ¼ .29) revealed a stronger effect compared to
which participants were asked to focus in the previous Type of Audio (G2 ¼ .23). Again, the overall difference
study. Indeed, previous research on music-induced in effect size may be related to the subjective and elusive
emotions has found that there tends to be more inter- character of emotional reactions as compared with per-
individual variability in evaluations of felt emotion ceived expressive intentions, but the difference in the
compared to evaluations of perceived emotion (e.g., relative contribution of auditory and visual modalities
Juslin, 2009). suggests that while visual kinematic cues may be more
Interestingly, the effect of video type in the video-only salient than auditory cues in communicating expressive
rating condition (G2 ¼ .26) was somewhat larger than intentions, their contribution to the emotional impact of
the effect of audio type in the audio-only condition, performances is comparable to that of auditory perfor-
though there was no statistically significant difference mance cues. The results of the linear regression analysis
between the Normal and Exaggerated video types in support this conclusion, by showing that audio-only
terms of their emotional impact. Although this effect size ratings and video-only ratings explain comparable pro-
is smaller than that observed in a previous experiment portions of the variance in the audiovisual ratings of
investigating the perception of expressivity (G2 ¼ .61; emotional impact. As in the case of the unimodal rating

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
464 Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence, & Eric F. Clarke

blocks, it is possible that some of the participants based expressivity are affected by visual kinematic cues. Thus,
their ratings of emotional impact on perceived expres- the aim of Experiment 2 was to investigate whether
sivity rather than their actual emotional reactions. Note, visual kinematic cues might affect the perception of the
though, that this is an issue that affects all studies aim- key auditory features contributing to perceived expres-
ing to investigate music-induced emotions using self- sivity, namely loudness and tempo variability. Since the
report measures, and can be minimized by giving clear aim was to obtain as reliable and consistent an evalua-
instructions to participants (see e.g., Konečni, 2008). We tion of loudness and tempo variation as possible, only
gave our participants explicit instructions to focus on those participants with musical instrument training
the ‘‘emotional effect that the performance has on you,’’ were recruited to take part in this experiment.
and used unambiguous labels (‘‘did not move me at all’’
and ‘‘moved me very strongly’’) to signify the extremes METHOD
of the rating scale. Participants. Seventeen participants (7 male, 10 female)
Finally, the contribution of either modality to the aged 18-61 years (M ¼ 26.3, SD ¼ 11.7) were recruited
emotional impact of a performance may depend on the from the University of Oxford community. All of the
performer and her or his efficacy in conveying expres- participants had received a minimum of two years of
sive intentions via body movements and auditory cues. music training on an instrument (ranging from 2 to 17
In the present study, the audio-only excerpts of Pianist 2 years; M ¼ 10.2, SD ¼ 4.9). The participants received
were evaluated as having a stronger emotional impact a monetary incentive (£5) for taking part in the study.
than those of Pianist 1, while the video-only ratings
Stimuli. The stimuli were the same as those in Experi-
revealed the opposite pattern. These results are in line
ment 1.
with the objective measures of auditory and kinematic
features (see Table 1), with Pianist 2 displaying more Procedure. The procedure was almost identical to that of
tempo variability, and Pianist 1 displaying more move- Experiment 1, with the difference that instead of emo-
ment overall. However, there was no effect of Pianist in tional impact, the participants were asked to evaluate
the ratings obtained in the audiovisual condition (and the amount of loudness (dynamic) variation, and the
no interaction effects), thus suggesting that the relative amount of tempo variation, in the performances. They
contribution of the auditory and visual modalities to the were instructed that ‘‘A performance with no variation
emotional impact of audiovisual performances may not in dynamics or timing would sound flat and mechani-
be significantly affected by differences in expressive effi- cal, while a performance with an extreme amount of
cacy. Furthermore, it should be noted that the facial variation would have continuous changes in tempo and
expressions of performers – which would sometimes loudness.’’ Both evaluations were made using horizontal
be visible to the audience in real-life performance situa- visual analog scales (width 278 pixels) ranging from
tions and which are eliminated in this study by the use ‘‘No variation at all’’ to ‘‘An extreme amount of varia-
of stick figures – may add significantly to the overall tion.’’ The order in which the scales were presented was
emotional impact of a musical performance. balanced across participants. The same rating scales were
also used in two unimodal rating conditions. In the
Experiment 2 video-only condition, the participants were instructed
to ‘‘try to imagine how the music produced by the pia-
The results of Experiment 1 revealed that auditory and nists’ movements would sound, and evaluate the amount
visual kinematic performance cues seem to account for of variation in the timing and dynamics of the imagined
comparable proportions of participants’ subjective emo- performances.’’ The audiovisual block was always pre-
tional reactions to piano performance excerpts. How- sented first, followed by the audio-only and video-only
ever, the potential crossmodal effects involved in the blocks. As the presentation order of the unimodal blocks
process remain unclear. A previous study by Vuoskoski had no significant effect on participants’ ratings in Exper-
et al. (2014) revealed that visual kinematic cues can iment 1, all participants in this experiment completed the
affect the ratings of perceived auditory expressivity, but unimodal blocks in the same order.
it is not yet known whether this effect reflects actual
crossmodal interactions between the auditory and visual RESULTS
modalities, or whether instead it could be attributed, for Unimodal perception of loudness and tempo variability.
example, to some kind of response bias. Furthermore, if In order to determine whether the perceived amount of
the observed effects were due to crossmodal interac- loudness and tempo variation differed significantly
tions, it is unclear which aspects of perceived auditory between the different performance conditions, the ratings

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 465

FIGURE 4. The mean ratings of loudness and tempo variability (+ standard error of the mean) obtained in the unimodal audio-only and video-only
conditions of Experiment 2. The ratings have been scaled to a range of 0-100.

obtained in the unimodal audio-only rating condition were able to consistently estimate the amount of loudness
were analysed using repeated-measures ANOVAs. The and tempo variation based on the pianists’ movements
mean ratings are displayed in Figure 4. There were two alone. Type of Time-warp was included as a factor in
within-participant factors: Type of Audio (Deadpan, order to see whether there were any differences between
Normal, or Exaggerated), and Pianist (1 or 2). One par- the time-warped and the original animations, since time-
ticipant’s audio-only ratings were not recorded due to warping changes the timing of the movements. In the
a technical failure, and thus n ¼ 16 for this analysis. In ratings of loudness variation, there were significant main
the ratings of the perceived amount of loudness variation, effects of Type of Video; F(2, 32) ¼ 49.35, p < .001, G2 ¼
there was a significant main effect of Type of Audio; F(2, .38, and Pianist; F(1, 16) ¼ 26.29, p < .001, G2 ¼ .09, but
30) ¼ 41.01, p < .001, G2 ¼ .40, but no effect of Pianist no main effect of Type of Time-warp. Multiple compar-
nor any interaction. Multiple comparisons of means isons of means revealed that all three video types differed
(paired t-tests, level of statistical significance adjusted significantly from one another in terms of loudness var-
using the Holm-Bonferroni method) revealed that all iability, with the Deadpan video type receiving the lowest
three audio types differed significantly from each other and the Exaggerated video type the highest ratings. Fur-
in terms of the perceived amount of loudness variation, thermore, a comparison of means revealed that Pianist 1
with the Deadpan audio type receiving the lowest and the was rated as exhibiting more loudness variation. There
Exaggerated audio type the highest ratings. A similar were also interaction effects between Type of Video and
analysis was conducted on the ratings of the amount of Pianist; F(2, 32) ¼ 20.77, p < .001, G2 ¼ .07, and between
tempo variation. This analysis yielded a significant main Type of Time-warp and Pianist; F(2, 32) ¼ 7.30, p < .01,
effect of Type of Audio; F(2, 30) ¼ 46.61, p < .001, G2 ¼ G2 ¼ .02. Multiple comparisons of means revealed that
.48, but once again no effect of Pianist and no interaction Pianist 1 was rated as exhibiting more loudness variation
effect was observed. Multiple comparisons of means than Pianist 2 only in the Normal and Exaggerated video
revealed that all three audio types differed significantly types. Multiple comparisons investigating the interaction
from each other in terms of the perceived amount of effect between Type of Time-warp and Pianist failed to
tempo variation, with the Deadpan audio type receiving reach statistical significance after the Holm-Bonferroni
the lowest and the Exaggerated audio type the highest correction had been applied.
ratings. A similar analysis was conducted on the ratings of
The next step was to investigate the ratings of loud- tempo variation obtained in the video-only condition,
ness and tempo variation obtained in the video-only yielding significant main effects of Type of Video; F(2,
condition, where the participants were instructed to base 32) ¼ 38.73, p < .001, G2 ¼ .26, Type of Time-warp; F(2,
their ratings on how they imagined the music produced 32) ¼ 4.94, p < .05, G2 ¼ .02, and Pianist; F(1, 16) ¼ 6.18,
by the observed movements would sound. Repeated- p < .05, G2 ¼ .02. Multiple comparisons of means
measures ANOVAs with three within-participants revealed that the Deadpan video type was rated as sig-
factors – Type of Video, Type of Time-warp, and Pianist nificantly lower in tempo variation than the Normal
– were conducted to investigate whether the participants and Exaggerated video types, but that there was no

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
466 Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence, & Eric F. Clarke

FIGURE 5. The mean ratings of loudness variability (+ standard error FIGURE 6. The mean ratings of tempo variability (+ standard error of
of the mean) obtained in the audiovisual rating condition of Experiment the mean) obtained in the audiovisual rating condition in Experiment 2,
2, grouped by Type of Audio and Type of Video. The ratings have been grouped by Type of Audio and Type of Video. The ratings have been
scaled to a range of 0-100. scaled to a range of 0-100.

statistically significant difference between the latter type receiving the lowest and the Exaggerated audio type
two. Multiple comparisons for the main effect of Type the highest ratings. For the effect of Type of Video, mul-
of Time-warp failed to reach statistical significance tiple comparisons of means revealed that there was a sta-
after the Holm-Bonferroni correction had been tistically significant difference only between the Deadpan
applied. A comparison of means also revealed that Pia- and Normal video types, with the Deadpan video type
nist 1 was rated as exhibiting more tempo variation receiving significantly lower ratings of loudness variation.
than Pianist 2, with interaction effects between Type A comparison of the means also revealed that Pianist 2
of Video and Pianist; F(2, 32) ¼ 3.36, p < .05, G2 ¼ was rated as exhibiting more loudness variation than
.02, and between Type of Time-warp and Pianist; F(2, Pianist 1.
32) ¼ 5.78, p < .01, G2 ¼ .01. Multiple comparisons Finally, the potential effect of visual cues on the per-
revealed that Pianist 1 was rated as exhibiting more ception of tempo variation was investigated by conduct-
tempo variation than Pianist 2 only in the case of the ing a similar repeated-measures ANOVA on the ratings
Exaggerated video type. Furthermore, multiple compar- of tempo variation (see Figure 6 for mean ratings). Once
isons revealed that Type of Time-warp only had a sig- again, there were three within-participants factors: Type
nificant effect on the ratings of tempo variation in the of Audio, Type of Video, and Pianist. The analysis yielded
case of Pianist 2, with the videos warped to Exaggerated significant main effects of Type of Audio; F(2, 32) ¼
audio receiving higher ratings than those warped to 61.47, p < .001, G2 ¼ .45, and Pianist; F(1, 16) ¼ 7.70,
Normal or Deadpan audio. p < .05, G2 ¼ .03, but no effect of Type of Video, nor any
interaction effects. Multiple comparisons of means
Bimodal perception of loudness and tempo variability. In
revealed that all three audio types were rated as signifi-
order to investigate the potential effect of visual cues on
cantly different in terms of the amount of tempo varia-
the perception of loudness variation, the ratings of loud-
tion, with the Deadpan audio type receiving the lowest
ness variation – obtained in the audiovisual rating con-
and the Exaggerated audio type the highest ratings. A
dition – were analysed using a repeated-measures
comparison of means also revealed that Pianist 2 was
ANOVA. The mean values are displayed in Figure 5.
rated as exhibiting more tempo variation than Pianist 1.
There were three within-participants factors: Type of
Audio, Type of Video, and Pianist. The analysis yielded
significant main effects of Type of Audio; F(2, 32) ¼ DISCUSSION
72.69, p < .001, G2 ¼ .38, Type of Video; F(2, 32) ¼ 3.71, The ratings of loudness and tempo variability obtained
p < .05, G2 ¼ .01, and Pianist; F(1, 16) ¼ 6.70, p < .05, in the audio-only condition demonstrated – in line with
G2 ¼ .03. There were no interaction effects. Multiple the objective measures of loudness and tempo variability
comparisons of means revealed that all three audio types (see Table 1) – that the performances produced under all
were rated as significantly different in terms of the three expressive conditions were evaluated as signifi-
amount of loudness variation, with the Deadpan audio cantly different in terms of the perceived loudness and

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 467

timing variation. Furthermore, there were no signifi- These results are consistent with the hypothesis that
cant differences between the two pianists in terms of visual kinematic information exerts a crossmodal influ-
perceived loudness and tempo variability. In the silent ence on the perception of loudness variability, but not
video-only rating condition, where the participants on the perception of tempo variability. However, the
were instructed to imagine how the music produced pattern of crossmodal effects observed in the two
by the pianists’ movements would sound, the partici- experiments reported here was not entirely straightfor-
pants rated all three video types as significantly differ- ward. The theory of optimal sensory integration (e.g.,
ent in terms of their loudness variability. Since the total Alais & Burr, 2004; Ernst & Banks, 2002), which pro-
amount of movement increased significantly from poses that more weight is given to the modality that
Deadpan to Exaggerated performances (see Table 1), provides the more reliable sensory information, does
this suggests that participants used the size of move- not fully explain why the loudness variability of the
ments as the cue in their evaluations. This conclusion Normal video type was evaluated as significantly higher
is further supported by the finding that Pianist 1 – who than that of the Deadpan video type while the Exagger-
showed more movement variation across the different ated video type was not. An alternative account is
performance types (see Table 1, right hand column) – offered by those studies that have demonstrated that
was evaluated as exhibiting more loudness variation when sounds and sights are perceived as originating
than Pianist 2 in the video-only condition. In the from a common event (i.e., the unity assumption), the
video-only ratings of tempo variability, the notably process of sensory integration is altered in a way that
larger effect size for Type of Video relative to Type of differs from the traditional understanding of optimal
Time-warp (which represented the timing model to integration (e.g., Schutz & Kubovy, 2009). However,
which the animation was time-warped and matched) studies that have investigated the unity assumption
suggests that participants used the simple amount of using musical instrument stimuli have reported con-
movement – rather than the pattern of timing of those flicting findings, either succeeding (Schutz & Kubovy,
movements – as a cue. This finding may be explained 2009) or failing (Vatakis & Spence, 2008) to find an
by the limited temporal resolution of the visual modal- effect of the unity assumption. Mitterer and Jesse
ity (e.g., Freides, 1974; Welch, DuttonHurt, & Warren, (2010) propose that multisensory integration may actu-
1986), as well as the strong real-world association ally be driven by learned co-occurrences of visual and
between the size of performers’ movements and the auditory stimuli rather than their perceived common
amount of tempo and loudness variation. causation: using piano stimuli showing either a key
Although Pianist 1 was evaluated as exhibiting more stroke or the actual sound-producing hammer stroke,
loudness and tempo variation than Pianist 2 in the video- they demonstrated that multisensory integration was
only condition, this pattern of results was reversed in the stronger in the case of key strokes. As there is a strong
audiovisual rating condition. The audiovisual ratings real-world correlation between auditory and visual cues
revealed that Pianist 2 was evaluated as exhibiting more of musical expressivity – with performers finding it dif-
loudness and tempo variation than Pianist 1 – a result ficult to retain their normal level of expression while
that is in line with the objective measures of audio fea- restricting their movements (Thompson & Luck,
tures (see Table 1). Interestingly, however, there was no 2012) – this account may also reflect the process under-
effect of Pianist in the audio-only condition. As in the lying the effects observed in the present study.
audio-only condition, all three audio types were evalu- In line with this proposal, it may be that the degree of
ated as significantly different in terms of their loudness crossmodal effect observed in the perception of loud-
and tempo variability in the audiovisual rating condi- ness variability varied depending on the ecological plau-
tion. The effect of Type of Audio on the evaluations of sibility of the audio-video combinations, suggesting that
loudness variability was comparable to that observed in only those cues that could be meaningfully paired with
the audio-only condition, but Type of Video also had cues in the other modality resulted in crossmodal effects
a statistically significant effect. More specifically, when (cf. Warren, Welch, & McCarthy, 1981). This interpre-
the different audio types were presented in combination tation is in line with the findings of Vuoskoski et al.
with the Deadpan video type, they received lower rat- (2014), who observed that the more contrasting
ings of loudness variability than when presented audio-visual combinations resulted in weaker crossmo-
together with the Normal video type; while for the rat- dal effects.
ings of tempo variability, the effect of Type of Audio was Finally, there is a need to consider the potential effect
comparable to the audio-only ratings, and showed no of response bias on the observed effects. It may be that
effect of Type of Video. only participants’ evaluations of loudness variability

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
468 Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence, & Eric F. Clarke

were affected by visual cues, while their perceptions of observers’ subjective emotional reactions to a musical
loudness variability remained unaltered. We did not performance, the visual modality appears to be just as
explicitly instruct the participants to base their evalua- important as the auditory modality.
tions only on the auditory modality, as we expected The significant contribution of visual cues to our par-
musically trained participants to have an established ticipants’ emotional experiences is striking, since the
understanding of loudness and tempo variability as effects of performance features on the perception and
musical features; and asking participants to base their induction of emotion have often been considered only
ratings on one modality while still attending to the from an auditory perspective (see e.g., Juslin & Timmers,
other, risks drawing participants’ attention to the phe- 2010) – despite more widespread recognition of the role
nomenon under investigation, thus increasing the likeli- of visual factors in judgements of performance expres-
hood of demand characteristics. The fact that the ratings sivity (e.g., Davidson, 1993,1994; Tsay, 2013). There is
of tempo variability were not affected by the simulta- some evidence to suggest that the type of emotional
neously presented visual kinematic information, and that expression communicated via visual kinematic cues can
visual information affected ratings of loudness variability have an effect on the type (and intensity) of emotions
only in the case of certain audio-visual pairings (across perceived and experienced by the observer of a musical
both pianists), suggests that the observed crossmodal performance (Chapados & Levitin, 2008; Krahé et al.,
effects cannot be explained solely in terms of response 2013; Timmers et al., 2006; Vines et al., 2011), but more
bias. However, further investigation is undoubtedly controlled and systematic investigations (e.g., within-
required to clarify whether visual information about participants rather than between-participants designs,
a piano performance could affect the perception of loud- and more systematically generated stimuli) are needed
ness at a sensory level. to explore this issue further. Moreover, recent findings
suggest that the emotions felt by a performer also alter
General Discussion the way in which he or she moves, since observers seem
to perceive visually and audiovisually presented (but not
This study provides further evidence for the significance solely auditorily presented) violin performances as sad-
of visual kinematic cues in the perception and experi- der when the performer was actually feeling sad, com-
ence of musical performance. Although previous studies pared to when they were only expressing sadness (Van
have shown that visual information can influence the Zijl & Luck, 2013). These findings – as well as those of the
emotions induced by a musical performance (e.g., Cha- present study – support the view that observers of a musi-
pados & Levitin, 2008; Krahé et al., 2013; Timmers, cal performance are able to detect very subtle yet infor-
Marolt, Camurri, & Volpe, 2006, Vines et al., 2011), they mative cues from visual kinematic information – without
haven’t been able to reliably estimate the effects size of necessarily attending to them in a conscious manner
visual performance cues relative to that of auditory per- (Tsay, 2013).
formance cues. The present study revealed that – in The results of the present study also provide evidence
terms of the emotional impact of musical performances to support the view that visual kinematic information
– the contribution of visual kinematic performance cues can have an effect on the judgment of certain auditory
appears to be comparable to that of auditory perfor- performance cues. The results of Experiment 2 revealed
mance cues. This is not to say that the effect of visual that visual kinematic information had an impact on rat-
cues would be equal to that of musical cues as a whole, ings of loudness variability – but not on ratings of tempo
since there is the significant impact of the music’s com- variability – suggesting that the crossmodal effects in the
posed structure to consider in addition to auditory per- perception of auditory expressivity observed in a previous
formance features. The emotions conveyed and induced study (Vuoskoski et al., 2014) may be attributed to the
by music emerge from the combination of structural effect of visual cues on perceived loudness (rather than
and performance features, and are also affected by indi- tempo) variability. In order to tease out the relative con-
vidual and situational factors (e.g., Scherer & Zentner, tributions of timing and loudness variability – as well as
2001). In relation to this complex range of factors, the the effects of visual kinematic cues – on perceived audi-
present study was only designed to investigate the rela- tory expressivity in more detail, future studies could
tive contributions of auditory and visual kinematic per- apply time-warping algorithms to MIDI data as well.
formance cues by comparing different performances In the case of both experiments, there seemed to be
(and combinations of different performances) of the a clearer difference between the Deadpan and Normal
same musical piece. Thus, the results of this study sug- performance types than between the Normal and Exag-
gest that in terms of the effect of performance cues on gerated performance types. This is in line with the

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 469

findings of Vuoskoski et al. (2014), as well as those of of musicians better than do professional concert
Davidson (1993) suggesting that performers may find it pianists.
easier to ‘‘withhold expression from the piece than exag- In conclusion, the results of the two experiments
gerate the expressivity of a piece beyond its normal reported here demonstrate that visual information about
level’’ (Davidson, 1993, p. 109). It should also be noted a performer’s movements not only has an impact on the
that performers can differ greatly in terms of how much intensity of emotional reactions evoked by the perfor-
they move while performing, as well as how much loud- mance, but can also change how that performance
ness and timing variation they use when communicat- sounds to an observer. The study has shown that visual
ing their expressive intentions. Indeed, this was the case performance cues may be just as important as auditory
in the present study, where Pianist 1 displayed more performance cues in terms of the subjective emotional
movement variability, whereas Pianist 2 exhibited more experience of the observer, suggesting that non-auditory
tempo variability. Although the effects observed in this cues may contribute more to music-induced emotions
study were consistent across pianists (as evidenced by than has previously been established. These results con-
the lack of interaction effects related to Pianist), it may firm the significant role of visual kinematic cues for audi-
be that the relative contributions of auditory and visual ence members, and encourage further investigations into
kinematic performance cues may vary across different the ways in which visual information may interact with
pianists, especially in the case of more extreme perfor- auditory information in our perception and experience of
mance styles. Indeed, differences in expressive efficacy – a musical performance.
between different performers and between different
instruments – may explain the contrasting findings Author Note
observed in the present study and a previous study by
Vines et al. (2011), where different expressive intentions We are grateful to three anonymous reviewers and the
led to differing emotional reactions only in the audio- Action Editor for their helpful comments on an earlier
visual and video-only conditions, but not in the audio- version of the paper. This research was supported by the
only condition. However, it might also be argued that Andrew W. Mellon Foundation.
the pianists included in this study – music students Correspondence concerning this article should be
rather than professional concert pianists – utilize more addressed to Jonna K. Vuoskoski, Faculty of Music, Uni-
conventional (i.e., less idiosyncratic) expressive devices versity of Oxford, St Aldate’s, OX1 1DB, Oxford, United
in their performances, and thus represent the majority Kingdom. E-mail: [Link]@[Link]

References

A LAIS , D., & B URR , D. (2004). The ventriloquist effect results C ASTELLANO, G., M ORTILLARO, M., C AMURRI , A., V OLPE , G., &
from near-optimal bimodal integration. Current Biology, 14, S CHERER , K. (2008). Automated analysis of body movement in
257-262. emotionally expressive piano performances. Music Perception,
B AKEMAN , R. (2005). Recommended effect size statistics for 26, 103-119.
repeated measures designs. Behavior Research Methods, 37, C HAPADOS , C., & L EVITIN , D. J. (2008). Cross-modal interac-
379-384. tions in the experience of musical performances: Physiological
B HATARA , A., T IROVOLAS , A. K., M ARIE D UAN , L., L EVY, B., & correlates. Cognition, 108, 639-651.
L EVITIN , D. J. (2011). Perception of emotional expression in C LARKE , E. F. (1988). Generative principles in music perfor-
musical performance. Journal of Experimental Psychology: mance. In J. A. Sloboda (Ed.), Generative processes in music:
Human Perception and Performance, 37, 921-934. The psychology of performance, improvisation, and composition
B URGER , B., & T OIVIAINEN , P. (2013). MoCap Toolbox – (pp. 1-26). Oxford: Oxford University Press.
A Matlab toolbox for computational analysis of movement DAHL , S., & F RIBERG , A. (2007). Visual perception of expres-
data. In R. Bresin (Ed.), Proceedings of the 10th Sound and siveness in musicians’ body movements. Music Perception, 24,
Music Computing Conference. Stockholm, Sweden: KTH Royal 433-454.
Institute of Technology. DAVIDSON , J. W. (1993). Visual perception of performance
B URR , D., B ANKS , M. S., & M ORRONE , M. C. (2009). Auditory manner in the movements of solo musicians. Psychology of
dominance over vision in the perception of interval duration. Music, 21, 103-113.
Experimental Brain Research, 198, 49-57.

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
470 Jonna K. Vuoskoski, Marc R. Thompson, Charles Spence, & Eric F. Clarke

DAVIDSON , J. W. (1994). What type of information is conveyed in M AC R ITCHIE , J., B UCK , B., & B AILEY, N. J. (2013). Inferring
the body movements of solo musician performers? Journal of musical structure through bodily gestures. Musicae Scientiae,
Human Movement Science, 6, 279-301. 17, 86-108.
E RNST, M. O., & B ANKS , M. S. (2002). Humans integrate visual M ITTERER , H., & J ESSE , A. (2010). Correlation versus causation
and haptic information in a statistically optimal fashion. in multisensory perception. Psychonomic Bulletin and Review,
Nature, 415(6870), 429-433. 17, 329-334.
F REIDES , D. (1974). Human information processing and sensory M ORRISON , S. J., P RICE , H. E., G EIGER , C. G., & C ORNACCHIO,
modality: Cross-modal functions, information complexity, R. A. (2009). The effect of conductor expressivity on ensemble
memory, and deficit. Psychological Bulletin, 81, 284-310. performance evaluation. Journal of Research in Music
G ABRIELSSON , A. (1999). The performance of music. In D. Education, 57, 37-49.
Deutsch (Ed.), The psychology of music (2nd ed., pp. 501-602). O RNE , M. T. (1962). On the social psychology of the psycho-
San Diego, CA: Academic Press. logical experiment with particular reference to demand char-
G LOWINSKI , D., M ANCINI , M., C OWIE , R., C AMURRI , A., acteristics and their implications. American Psychologist, 17,
C HIORRI , C., & D OHERTY, C. (2013). The movements made by 776-783.
performers in a skilled quartet: A distinctive pattern, and the PALMER , C. (1997). Music performance. Annual Review of
function that it serves. Frontiers in Psychology, 4, 841. Psychology, 48(1), 115-138.
G OEBL , W., & PALMER , C. (2009). Synchronization of timing and P ETRINI , K., M C A LEER , P., & P OLLICK , F. (2010). Audiovisual
motion among performing musicians. Music Perception, 26, integration of emotional signals from music improvisation
427-438. does not depend on temporal correspondence. Brain Research,
G RAHN , J. A. (2012). See what I hear? Beat perception in auditory 1323, 139-148.
and visual rhythms. Experimental Brain Research, 220, 51-61. P LATZ , F., & KOPIEZ , R. (2012). When the eye listens: A meta-
H ARGREAVES , D. J., & N ORTH , A. C. (2010). Experimental aes- analysis of how audio-visual presentation enhances the
thetics and liking for music. In P. N. Juslin & J. A. Sloboda appreciation of music performance. Music Perception, 30,
(Eds.), Handbook of music and emotion: Theory, research, 71-83.
applications (pp. 515-546). Oxford: Oxford University Press. R EPP, B. H., & P ENEL , A. (2002). Auditory dominance in tem-
H OLM , S. (1979). A simple sequentially rejective multiple test poral processing: New evidence from synchronization with
procedure. Scandinavian Journal of Statistics, 6, 65-70. simultaneous visual and auditory sequences. Journal of
J UCHNIEWICZ , J. (2008). The influence of physical movement on Experimental Psychology: Human Perception and Performance,
the perception of musical performance. Psychology of Music, 28, 1085-1099.
36, 417-427. R OSENBLUM , L. D., & F OWLER , C. A. (1991). Audiovisual
J USLIN , P. N. (2001). Communicating emotion in music perfor- investigation of the loudness-effort effect for speech and
mance: A review and a theoretical framework. In P. N. Juslin & nonspeech events. Journal of Experimental Psychology: Human
J. A. Sloboda (Eds.), Music and emotion: Theory and research Perception and Performance, 17, 976-985.
(pp. 309-337). Oxford: Oxford University Press. S ALDAÑA , H. M., & R OSENBLUM , L. D. (1993). Visual influences
J USLIN , P. N. (2003). Five facets of musical expression: A psy- on auditory pluck and bow judgments. Perception and
chologist’s perspective on music performance. Psychology of Psychophysics, 54, 406-416.
Music, 31, 273-302. S CHUTZ , M., & K UBOVY, M. (2009). Causality and cross-modal
J USLIN , P. N. (2009). Emotional responses to music. In S. Hallam, integration. Journal of Experimental Psychology: Human
I. Cross, & M. Thaut (Eds.), The Oxford handbook of music Perception and Performance, 35, 1791-1810.
psychology (pp. 131-140). Oxford: Oxford University Press. S CHUTZ , M., & L IPSCOMB , S. (2007). Hearing gestures, seeing
J USLIN , P. N., & T IMMERS , R. (2010). Expression and commu- music: Vision influences perceived tone duration. Perception,
nication of emotion in music performance. In P. N. Juslin & J. 36, 888-897.
A. Sloboda (Eds.), Handbook of music and emotion: Theory, S LOBODA , J. A., & J USLIN , P. N. (2010). At the interface between
research, applications (pp. 453-489). Oxford: Oxford the inner and outer world: Psychological perspectives. In P. N.
University Press. Juslin & J. A. Sloboda (Eds.), Handbook of music and emotion:
KONENI , V. J. (2008). Does music induce emotion? A theoretical Theory, research, applications (pp. 73- 98). Oxford: Oxford
and methodological analysis. Psychology of Aesthetics, University Press.
Creativity, and the Arts, 2, 115-129. S LOBODA , J. A., & L EHMANN , A. C. (2001). Tracking perfor-
K RAHÉ , C., HAHN , U., & W HITNEY, K. (2013). Is seeing (musical) mance correlates of changes in perceived intensity of emotion
believing? The eye versus the ear in emotional responses to during different interpretations of a Chopin piano prelude.
music. Psychology of Music [online before print]. DOI: Music Perception, 19, 87-120.
0305735613498920.

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]
Interaction of Sight and Sound 471

S U, Y. H. (2014). Audiovisual beat induction in complex auditory V ERRON , C. (2005). Traitement et visualisation de données ges-
rhythms: Point-light figure movement as an effective visual turalles captées par Optotrak [Processing and visualizing ges-
beat. Acta Psychologica, 151, 40-50. ture data captured by Optotrak]. Unpublished Report. Input
T HOMPSON , M. R., & LUCK , G. (2012). Exploring relationships Devices and Music Interaction Laboratory, McGill University.
between pianists’ body movements, their expressive intentions, Retrieved from [Link]
and structural elements of the music. Musicae Scientiae, 16, V INES , B. W., K RUMHANSL , C. L., WANDERLEY, M. M., DALCA , I.
19-40. M., & L EVITIN , D. J. (2011). Music to my eyes: Cross-modal
T HOMPSON , W. F., G RAHAM , P., & RUSSO, F. A. (2005). Seeing interactions in the perception of emotions in musical perfor-
music performance: Visual influences on perception and mance. Cognition, 118, 157-170.
experience. Semiotica, 156, 203-227. V INES , B. W., K RUMHANSL , C. L., WANDERLEY, M. M., &
T HOMPSON , W. F., R USSO, F. A., & Q UINTO, L. (2008). Audio– L EVITIN , D. J. (2006). Cross-modal interactions in the per-
visual integration of emotional cues in song. Cognition and ception of musical performance. Cognition, 101, 80-113.
Emotion, 22, 1457-1470. V UOSKOSKI , J. K., T HOMPSON , M. R., C LARKE , E. F., & S PENCE ,
T IMMERS , R., M AROLT, M., C AMURRI , A., & V OLPE , G. (2006). C. (2014). Crossmodal interactions in the perception of
Listeners’ emotional engagement with performances of expressivity in musical performance. Attention, Perception, and
a Scriabin étude: An explorative case study. Psychology of Psychophysics, 76, 591-604.
Music, 34, 481-510. WANDERLEY, M., V INES , B. W., M IDDLETON , N., M C K AY, C., &
T SAY, C. J. (2013). Sight over sound in the judgment of music H ATCH , W. (2005). The musical significance of clarinetists’
performance. Proceedings of the National Academy of Sciences ancillary gestures: An exploration of the field. Journal of New
of the USA, 110, 14580-14585. Music Research, 34, 97-113.
VAN Z IJL , A. G., & LUCK , G. (2013). The sound of sadness: The WARREN , D. H., W ELCH , R. B., & M C C ARTHY, T. J. (1981). The
effect of performers’ emotions on audience ratings. In G. Luck, role of visual-auditory ‘‘compellingness’’ in the ventriloquism
& O. Brabant (Eds.), Proceedings of the 3rd International effect: Implications for transitivity among the spatial senses.
Conference on Music & Emotion (ICME3). Jyväskylä, Finland: Perception and Psychophysics, 30, 557-564.
ICME3. W ELCH , R. B., D UTTON H URT, L. D., & WARREN , D. H. (1986).
VATAKIS , A., & S PENCE , C. (2008). Evaluating the influence of the Contributions of audition and vision to temporal rate per-
‘‘unity assumption’’ on the temporal perception of realistic ception. Perception and Psychophysics, 39, 294-300.
audiovisual stimuli. Acta Psychologica, 127, 12-23. W ILLIAMON , A., & DAVIDSON , J. W. (2002). Exploring
co-performer communication. Musicae Scientiae, 6, 53-72.

This content downloaded from


[Link] on Sat, 05 Aug 2023 14:59:49 +00:00
All use subject to [Link]

You might also like