English
Related papers

Related papers: STAR: Speech-to-Audio Generation via Representatio…

200 papers

One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. While this cascaded…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-06 Do June Min , Karel Mundnich , Andy Lapastora , Erfan Soltanmohammadi , Srikanth Ronanki , Kyu Han

This paper proposes a novel 3D speech-to-animation (STA) generation framework designed to address the shortcomings of existing models in producing diverse and emotionally resonant animations. Current STA models often generate animations…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Xulong Zhang , Xiaoyang Qu , Haoxiang Shi , Chunguang Xiao , Jianzong Wang

Speech to speech translation (S2ST) is a transformative technology that bridges global communication gaps, enabling real time multilingual interactions in diplomacy, tourism, and international trade. Our review examines the evolution of…

Computation and Language · Computer Science 2025-03-10 Mohammad Sarim , Saim Shakeel , Laeeba Javed , Jamaluddin , Mohammad Nadeem

Diffusion models have significantly advanced the field of talking head generation (THG). However, slow inference speeds and prevalent non-autoregressive paradigms severely constrain the application of diffusion-based THG models. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Haotian Wang , Yuzhe Weng , Jun Du , Haoran Xu , Xiaoyan Wu , Shan He , Bing Yin , Cong Liu , Qingfeng Liu

A cascaded speech translation model relies on discrete and non-differentiable transcription, which provides a supervision signal from the source side and helps the transformation between source speech and target text. Such modeling suffers…

Computation and Language · Computer Science 2020-11-25 Parnia Bahar , Tobias Bieschke , Ralf Schlüter , Hermann Ney

Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a…

Sound · Computer Science 2026-02-20 Bjorn Johnson , Jared Levy

Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to…

Sound · Computer Science 2020-07-30 Viet-Trung Dang , Tianyu Zhao , Sei Ueno , Hirofumi Inaguma , Tatsuya Kawahara

Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such as high costs in voice data collection, weakness in dynamic…

Computation and Language · Computer Science 2025-02-19 Ailin Huang , Boyong Wu , Bruce Wang , Chao Yan , Chen Hu , Chengli Feng , Fei Tian , Feiyu Shen , Jingbei Li , Mingrui Chen , Peng Liu , Ruihang Miao , Wang You , Xi Chen , Xuerui Yang , Yechang Huang , Yuxiang Zhang , Zheng Gong , Zixin Zhang , Hongyu Zhou , Jianjian Sun , Brian Li , Chengting Feng , Changyi Wan , Hanpeng Hu , Jianchang Wu , Jiangjie Zhen , Ranchen Ming , Song Yuan , Xuelin Zhang , Yu Zhou , Bingxin Li , Buyun Ma , Hongyuan Wang , Kang An , Wei Ji , Wen Li , Xuan Wen , Xiangwen Kong , Yuankai Ma , Yuanwei Liang , Yun Mou , Bahtiyar Ahmidi , Bin Wang , Bo Li , Changxin Miao , Chen Xu , Chenrun Wang , Dapeng Shi , Deshan Sun , Dingyuan Hu , Dula Sai , Enle Liu , Guanzhe Huang , Gulin Yan , Heng Wang , Haonan Jia , Haoyang Zhang , Jiahao Gong , Junjing Guo , Jiashuai Liu , Jiahong Liu , Jie Feng , Jie Wu , Jiaoren Wu , Jie Yang , Jinguo Wang , Jingyang Zhang , Junzhe Lin , Kaixiang Li , Lei Xia , Li Zhou , Liang Zhao , Longlong Gu , Mei Chen , Menglin Wu , Ming Li , Mingxiao Li , Mingliang Li , Mingyao Liang , Na Wang , Nie Hao , Qiling Wu , Qinyuan Tan , Ran Sun , Shuai Shuai , Shaoliang Pang , Shiliang Yang , Shuli Gao , Shanshan Yuan , Siqi Liu , Shihong Deng , Shilei Jiang , Sitong Liu , Tiancheng Cao , Tianyu Wang , Wenjin Deng , Wuxun Xie , Weipeng Ming , Wenqing He , Wen Sun , Xin Han , Xin Huang , Xiaomin Deng , Xiaojia Liu , Xin Wu , Xu Zhao , Yanan Wei , Yanbo Yu , Yang Cao , Yangguang Li , Yangzhen Ma , Yanming Xu , Yaoyu Wang , Yaqiang Shi , Yilei Wang , Yizhuang Zhou , Yinmin Zhong , Yang Zhang , Yaoben Wei , Yu Luo , Yuanwei Lu , Yuhe Yin , Yuchu Luo , Yuanhao Ding , Yuting Yan , Yaqi Dai , Yuxiang Yang , Zhe Xie , Zheng Ge , Zheng Sun , Zhewei Huang , Zhichao Chang , Zhisheng Guan , Zidong Yang , Zili Zhang , Binxing Jiao , Daxin Jiang , Heung-Yeung Shum , Jiansheng Chen , Jing Li , Shuchang Zhou , Xiangyu Zhang , Xinhao Zhang , Yibo Zhu

End-to-end speech-in speech-out dialogue systems are emerging as a powerful alternative to traditional ASR-LLM-TTS pipelines, generating more natural, expressive responses with significantly lower latency. However, these systems remain…

This work seeks the possibility of generating the human face from voice solely based on the audio-visual data without any human-labeled annotations. To this end, we propose a multi-modal learning framework that links the inference stage and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-14 Hyeong-Seok Choi , Changdae Park , Kyogu Lee

Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained perceptual reasoning.…

End-to-end spoken language understanding (SLU) predicts intent directly from audio using a single model. It promises to improve the performance of assistant systems by leveraging acoustic information lost in the intermediate textual…

Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits expressivity. Recent end-to-end methods reduce latency and…

While various end-to-end models for spoken language understanding tasks have been explored recently, this paper is probably the first known attempt to challenge the very difficult task of end-to-end spoken question answering (SQA). Learning…

Computation and Language · Computer Science 2020-08-12 Yung-Sung Chuang , Chi-Liang Liu , Hung-Yi Lee , Lin-shan Lee

This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs competitively at this task. TacoSpawn is a recurrent…

We aim to solve the highly challenging task of generating continuous sign language videos solely from speech segments for the first time. Recent efforts in this space have focused on generating such videos from human-annotated text…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Parul Kapoor , Rudrabha Mukhopadhyay , Sindhu B Hegde , Vinay Namboodiri , C V Jawahar

A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these components often requires extensive domain expertise and may contain…

Simultaneous translation models play a crucial role in facilitating communication. However, existing research primarily focuses on text-to-text or speech-to-text models, necessitating additional cascade components to achieve…

Computation and Language · Computer Science 2024-10-22 Zhengrui Ma , Qingkai Fang , Shaolei Zhang , Shoutao Guo , Yang Feng , Min Zhang

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Jinting Wang , Jun Wang , Hei Victor Cheng , Li Liu