English
Related papers

Related papers: Transhuman Ansambl - Voice Beyond Language

200 papers

There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models,…

Computation and Language · Computer Science 2024-11-01 Chenyang Le , Yao Qian , Dongmei Wang , Long Zhou , Shujie Liu , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Sheng Zhao , Michael Zeng

Normally, a system that translates speech into text consists of separate modules for speech recognition and text-to-text translation. Combining those tasks into a SpeechLLM promises to exploit paralinguistic information in the speech and to…

Computation and Language · Computer Science 2026-05-15 Titouan Parcollet , Shucong Zhang , Xianrui Zheng , Rogier C. van Dalen

In this paper I examine the process of getting affected by and the process of making sense of non-language sounds and propose the idea of the contextual cognitive apparatus or exo-brain. We are affected by a singing voice even when we do…

Neurons and Cognition · Quantitative Biology 2024-03-29 Ankur Betageri

Text-to-speech (TTS) and singing voice synthesis (SVS) aim at generating high-quality speaking and singing voice according to textual input and music scores, respectively. Unifying TTS and SVS into a single system is crucial to the…

Sound · Computer Science 2022-12-07 Yi Lei , Shan Yang , Xinsheng Wang , Qicong Xie , Jixun Yao , Lei Xie , Dan Su

Imagine AI assistants that enhance conversations without interrupting them: quietly providing relevant information during a medical consultation, seamlessly preparing materials as teachers discuss lesson plans, or unobtrusively scheduling…

Computation and Language · Computer Science 2025-09-23 Andrew Zhu , Chris Callison-Burch

The relationships between the listener, physical world and virtual environment (VE) should not only inspire the design of natural multimodal interfaces but should be discovered to make sense of the mediating action of VR technologies. This…

Human-Computer Interaction · Computer Science 2022-04-22 Michele Geronazzo , Stefania Serafin

We present the Voice Conversion Challenge 2018, designed as a follow up to the 2016 edition with the aim of providing a common framework for evaluating and comparing different state-of-the-art voice conversion (VC) systems. The objective of…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-13 Jaime Lorenzo-Trueba , Junichi Yamagishi , Tomoki Toda , Daisuke Saito , Fernando Villavicencio , Tomi Kinnunen , Zhenhua Ling

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating…

Human-Computer Interaction · Computer Science 2025-09-19 Taesoo Kim , Yongsik Jo , Hyunmin Song , Taehwan Kim

The COVID-19 pandemic shifted many events in our daily lives into the virtual domain. While virtual conference systems provide an alternative to physical meetings, larger events require a muted audience to avoid an accumulation of…

Multimedia · Computer Science 2023-10-30 Tamay Aykut , Markus Hofbauer , Christopher Kuhn , Eckehard Steinbach , Bernd Girod

As AI systems increasingly become embedded in interactive and im-mersive artistic environments, artists and technologists are discovering new opportunities to engage with their interpretive and autonomous capacities as creative…

Human-Computer Interaction · Computer Science 2026-02-09 Pavlos Panagiotidis , Jocelyn Spence , Nils Jaeger , Paul Tennent

With advancements in multimodal communication technologies, remote learning environments such as, distance universities are increasing. Remote learning typically happens asynchronously. As a consequence, unlike face-to-face in-person…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-14 Sargam Vyas , Bogdan Vlasenko , André Mayoraz , Egon Werlen , Per Bergamin , Mathew Magimai. -Doss

This work presents our advancements in controlling an articulatory speech synthesis engine, \textit{viz.}, Pink Trombone, with hand gestures. Our interface translates continuous finger movements and wrist flexion into continuous speech…

Sound · Computer Science 2021-02-03 Pramit Saha , Debasish Ray Mohapatra , Sidney Fels

Singing voice conversion is a task to convert a song sang by a source singer to the voice of a target singer. In this paper, we propose using a parallel data free, many-to-one voice conversion technique on singing voices. A phonetic…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-12 Xin Chen , Wei Chu , Jinxi Guo , Ning Xu

The way we engage with digital spaces and the digital world has undergone rapid changes in recent years, largely due to the emergence of the Metaverse. As technology continues to advance, the demand for sophisticated and immersive…

Human-Computer Interaction · Computer Science 2024-09-04 Senthil Kumar Jagatheesaperumal , Praveen Sathikumar , Harikrishnan Rajan

In this work, we investigate the phenomenon of transverse resonance and transverse standing waves that occur within the cochlea of living organisms. It is demonstrated that the predisposing factor for their occurrence is the cochlear shape,…

Neurons and Cognition · Quantitative Biology 2023-10-10 M. V. Semotiuk , A. V. Palagin

Recent progress in large language model (LLM) technology has significantly enhanced the interaction experience between humans and voice assistants (VAs). This project aims to explore a user's continuous interaction with LLM-based VA…

Human-Computer Interaction · Computer Science 2024-09-04 Szeyi Chan , Shihan Fu , Jiachen Li , Bingsheng Yao , Smit Desai , Mirjana Prpa , Dakuo Wang

We present VOICE, a novel approach to science communication that connects large language models' (LLM) conversational capabilities with interactive exploratory visualization. VOICE introduces several innovative technical contributions that…

Human-Computer Interaction · Computer Science 2026-05-19 Donggang Jia , Alexandra Irger , Lonni Besancon , Ondrej Strnad , Deng Luo , Johanna Bjorklund , Anders Ynnerman , Ivan Viola

Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer from a disconnect between perception and expression, resulting…

Sound · Computer Science 2026-03-02 Yueran Hou , Peilei Jia , Zihan Sun , Qihang Lu , Wenbing Yang , Yingming Gao , Ya Li , Jun Gao

The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present…

Sound · Computer Science 2025-05-15 Yicheng Gu , Chaoren Wang , Junan Zhang , Xueyao Zhang , Zihao Fang , Haorui He , Zhizheng Wu

The voice conversion challenge is a bi-annual scientific event held to compare and understand different voice conversion (VC) systems built on a common dataset. In 2020, we organized the third edition of the challenge and constructed and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-31 Yi Zhao , Wen-Chin Huang , Xiaohai Tian , Junichi Yamagishi , Rohan Kumar Das , Tomi Kinnunen , Zhenhua Ling , Tomoki Toda
‹ Prev 1 4 5 6 7 8 10 Next ›