KABURI TTS improves two person conversations with overlapping speech
KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction
SoundComputation and Language
Summary
It is hard for machines to create realistic conversations where two people talk at the same time or interrupt each other, like in normal chats. To solve this, the authors developed KABURI-TTS, a system that separately creates speech for two speakers and includes moments when their talk overlaps naturally. This system uses detailed sound information (phonemes) and voice activity to control who is speaking and when. Tests show KABURI-TTS sounds more natural and captures typical talk patterns such as when speakers take turns or speak over each other.
text-to-speechphonemevoice activity detectionconversational speechfull-duplex dialogueoverlapping speechbackchannelturn-takingutterancespeech synthesis
Authors
Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji Ogawa
Abstract
Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps that occur while the interlocutor is speaking. In this work, aiming at conversational speech synthesis that reproduces human-like overlap, we propose KABURI-TTS. KABURI-TTS takes a per-speaker phoneme raster as input and renders the speech of the two speakers on separate channels, conditioned on the per-frame phonemes and the voice activity derived from them. Because the phoneme raster is supplied by a separate module, the proposed method enables controllable generation of one-speaker-per-channel, two-party spoken dialogue. A user evaluation shows that, compared with strong baselines, the proposed method attains higher naturalness at both the utterance and the interaction level. Furthermore, an analysis of voice activity confirms that the proposed method produces more overlap and more frequent turn-taking.