English
Related papers

Related papers: Transhuman Ansambl - Voice Beyond Language

200 papers

This paper describes the design of NNSVS, an open-source software for neural network-based singing voice synthesis research. NNSVS is inspired by Sinsy, an open-source pioneer in singing voice synthesis research, and provides many…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-02 Ryuichi Yamamoto , Reo Yoneyama , Tomoki Toda

This article presents an interactive system for stage acoustics experimentation including considerations for hearing one's own and others' instruments. The quality of real-time auralization systems for psychophysical experiments on music…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-30 Ernesto Accolti , Lukas Aspöck , Manuj Yadav , Michael Vorländer

The advent of increasingly-growing virtual realities poses unprecedented opportunities and challenges to different societies. Artistic collectives are not an exception, and we here aim to put special attention into musicians. Compositions,…

Visually impaired people face numerous challenges when it comes to transportation. Not only must they circumvent obstacles while navigating, but they also need access to essential information related to available public transport,…

Human-Computer Interaction · Computer Science 2017-03-08 Gourav G. Shenoy , Mangirish A. Wagle , Kay Connelly

To foster effective human-agent interactions, designers must understand how vocal cues influence the perception of agent personality and the role of user-agent alignment in shaping these perceptions. In this work, we examine whether users…

Robotic ultrasound systems can enhance medical diagnostics, but patient acceptance is a challenge. We propose a system combining an AI-powered conversational virtual agent with three mixed reality visualizations to improve trust and…

Human-Computer Interaction · Computer Science 2025-07-18 Tianyu Song , Felix Pabst , Ulrich Eck , Nassir Navab

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Aggelina Chatziagapi , Louis-Philippe Morency , Hongyu Gong , Michael Zollhoefer , Dimitris Samaras , Alexander Richard

Hearing-impaired individuals often face significant barriers in daily communication due to the inherent challenges of producing clear speech. To address this, we introduce the Omni-Model paradigm into assistive technology and present…

Computation and Language · Computer Science 2025-11-17 Zhiming Ma , Shiyu Gan , Junhao Zhao , Xianming Li , Qingyun Pan , Peidong Wang , Mingjun Pan , Yuhao Mo , Jiajie Cheng , Chengxin Chen , Zhonglun Cao , Chonghan Liu , Shi Cheng

Physically assistive robots present an opportunity to significantly increase the well-being and independence of individuals with motor impairments or other forms of disability who are unable to complete activities of daily living. Speech…

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Muhammad Saad Saeed , Shah Nawaz , Pietro Morerio , Arif Mahmood , Ignazio Gallo , Muhammad Haroon Yousaf , Alessio Del Bue

The task of isolating a target singing voice in music videos has useful applications. In this work, we explore the single-channel singing voice separation problem from a multimodal perspective, by jointly learning from audio and visual…

Sound · Computer Science 2021-10-20 Juan F. Montesinos , Venkatesh S. Kadandale , Gloria Haro

Breath is a significant component in singing performance, which is still underresearched in most singing-related music interfaces. In this paper, we present a multimodal system that detects the learner's singing pitch and breathing states…

Human-Computer Interaction · Computer Science 2022-07-06 Ziyue Piao , Gus Xia

In-person human interaction relies on our spatial perception of each other and our surroundings. Current remote communication tools partially address each of these aspects. Video calls convey real user representations but without spatial…

Human-Computer Interaction · Computer Science 2023-09-06 Rishi Vanukuru , Suibi Che-Chuan Weng , Krithik Ranjan , Torin Hopkins , Amy Banic , Mark D. Gross , Ellen Yi-Luen Do

Machine Listening, as usually formalized, attempts to perform a task that is, from our perspective, fundamentally human-performable, and performed by humans. Current automated models of Machine Listening vary from purely data-driven…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-27 Laurie M. Heller , Benjamin Elizalde , Bhiksha Raj , Soham Deshmukh

We investigate the feasibility of a singing voice synthesis (SVS) system by using a decomposed framework to improve flexibility in generating singing voices. Due to data-driven approaches, SVS performs a music score-to-waveform mapping;…

Sound · Computer Science 2024-07-15 Lester Phillip Violeta , Taketo Akama

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the naturalness of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-23 Kazuhiro Nakamura , Shinji Takaki , Kei Hashimoto , Keiichiro Oura , Yoshihiko Nankaku , Keiichi Tokuda

The paper presents results from a project aiming to create horizontally distributed surround sound sources and virtual sound images as auditory BCI (aBCI) stimuli. The purpose is to create evoked brain wave response patterns depending on…

Human-Computer Interaction · Computer Science 2012-10-11 Nozomu Nishikawa , Yoshihiro Matsumoto , Shoji Makino , Tomasz M. Rutkowski

We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unified framework, and…

Sound · Computer Science 2019-12-24 Liqiang Zhang , Chengzhu Yu , Heng Lu , Chao Weng , Yusong Wu , Xiang Xie , Zijin Li , Dong Yu

Singing Voice Synthesis (SVS) has witnessed significant advancements with the advent of deep learning techniques. However, a significant challenge in SVS is the scarcity of labeled singing voice data, which limits the effectiveness of…

Sound · Computer Science 2024-12-17 Yifeng Yu , Jiatong Shi , Yuning Wu , Yuxun Tang , Shinji Watanabe

This paper introduces a novel task in generative speech processing, Acoustic Scene Transfer (AST), which aims to transfer acoustic scenes of speech signals to diverse environments. AST promises an immersive experience in speech perception…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-19 Miseul Kim , Soo-Whan Chung , Youna Ji , Hong-Goo Kang , Min-Seok Choi