0% found this document useful (0 votes)
9 views1 page

Portable Singing Voice Synthesizer Project

Most singers are untrained in composition and most composers are untrained in singing. Rather than attempting to train singers, it would be much easier to remove them from the equation using automation.

Uploaded by

Aaron Thapa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views1 page

Portable Singing Voice Synthesizer Project

Most singers are untrained in composition and most composers are untrained in singing. Rather than attempting to train singers, it would be much easier to remove them from the equation using automation.

Uploaded by

Aaron Thapa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Senior Research Project Proposal

Aaron Thapa
August 2022

Most singers are untrained in composition and most composers are untrained in singing. Rather
than attempting to train singers, it would be much easier to remove them from the equation using
automation.
One of the first text-to-speech technologies to support pitch control was DECtalk. Released in
1984, it allowed PC users to create speech using commands sent over a serial interface. Various
parameters allowed users to specify spesific phonemes, pitches, and timings to create a singing-
like voice. The technology suffered from issues such as low audio fidelity and difficulty with
integration into music.
In 2004, Yamaha released Vocaloid, a voice synthesis technology now entierly focused on
music production. As software product, it was trivial to integrate into music production software
and provide much more detailed control over the produced voice. Their methods along with the
convenience of the Japanese language allowed them to create a much more realistic-sounding
virtual voice. Yamaha continued to iterate on the technology, with the latest version, Vocaloid 5,
releasing in 2019.
In 2012, Yamaha experimented with implementing a version of this technology in one of their
synthesizer ICs, the YMW820 (NSX-1). It has an integrated RISC CPU which it uses to commu-
nicate with audio devices, MIDI devices, and a wavetable synthesis core.
Using this IC (more commonly available as the eVY1 module or eVY1 shield) would allow for
the creation of a portable singing voice synthesizer. While many have used Vocaloid in concert,
such a device would allow for live ”playing” of the synthesized voice. It could also allow a mute
person to speak or sing with a voice more directly than through typing.
The project should meet the following objectives: allow for easy selection of phoneme, prefer-
ably through a tactile method; use a method familiar to musicians for pitch control; produce sound
through integrated speakers and allow for connection of external output; integrate with other MIDI
devices through USB and/or the DIN connector; expose vocal parameters to the user through
knobs/sliders.
There is little documentation on the internet for the IC as a lot of it has been deleted after the
original publication. Most of it has been archived on the sandsoftwaresound blog. ([Link]
[Link]/yamaha-nsx-1-resources/) During the last school year, I also did some preliminary
research on the chip. ([Link] In terms of support from other
humans, the extreme degree to which this topic is niche, especially among English speakers, means
that short of contacting Yamaha there is basically no option.

You might also like