English
Related papers

Related papers: VITA-Audio: Fast Interleaved Cross-Modal Token Gen…

200 papers

Current audio language models are predominantly text-first, either extending pre-trained text LLM backbones or relying on semantic-only audio tokens, limiting general audio modeling. This paper presents a systematic empirical study of…

Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements in AIGC technologies for text and image generation, the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-18 Ruibo Fu , Shuchen Shi , Hongming Guo , Tao Wang , Chunyu Qiang , Zhengqi Wen , Jianhua Tao , Xin Qi , Yi Lu , Xiaopeng Wang , Zhiyong Wang , Yukun Liu , Xuefei Liu , Shuai Zhang , Guanjun Li

We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting module, which processes…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-26 Joun Yeop Lee , Myeonghun Jeong , Minchan Kim , Ji-Hyun Lee , Hoon-Young Cho , Nam Soo Kim

Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humming a song, which we formalize as role-playing and singing. We present VITA-QinYu, the…

Computation and Language · Computer Science 2026-05-11 Jiacheng Xu , Heting Gao , Liufei Xie , Zhenchuan Yang , Lijiang Li , Yiting Chen , Bin Zhang , Meng Chen , Chaoyu Fu , Weifeng Zhao , Wenjiang Zhou

The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Xiulong Liu , Kun Su , Eli Shlizerman

This paper presents VoiceLDM, a model designed to produce audio that accurately follows two distinct natural language text prompts: the description prompt and the content prompt. The former provides information about the overall…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-26 Yeonghyeon Lee , Inmo Yeon , Juhan Nam , Joon Son Chung

Text-to-audio (TTA), which generates audio signals from textual descriptions, has received huge attention in recent years. However, recent works focused on text to monaural audio only. As we know, spatial audio provides more immersive…

Sound · Computer Science 2025-06-09 Lei Zhao , Sizhou Chen , Linfeng Feng , Jichao Zhang , Xiao-Lei Zhang , Chi Zhang , Xuelong Li

Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusion models still needs improvement. And the effectiveness of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Wenhao Guan , Kaidi Wang , Wangjin Zhou , Yang Wang , Feng Deng , Hui Wang , Lin Li , Qingyang Hong , Yong Qin

In this paper we propose a novel virtual simulation-pilot engine for speeding up air traffic controller (ATCo) training by integrating different state-of-the-art artificial intelligence (AI) based tools. The virtual simulation-pilot engine…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-18 Juan Zuluaga-Gomez , Amrutha Prasad , Iuliia Nigmatulina , Petr Motlicek , Matthias Kleinert

Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models experience high latency due to their unordered temporal nature.…

Sound · Computer Science 2026-03-17 Zhengyan Sheng , Zhihao Du , Shiliang Zhang , Zhijie Yan , Liping Chen

We are interested in a novel task, namely low-resource text-to-talking avatar. Given only a few-minute-long talking person video with the audio track as the training data and arbitrary texts as the driving input, we aim to synthesize…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Zhenhui Ye , Ziyue Jiang , Yi Ren , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

We present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours. The model predicts pitch contours during inference. By altering these predictions, the generated speech can be…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-17 Adrian Łańcucki

Diffusion-based models have gained wide adoption in the virtual human generation due to their outstanding expressiveness. However, their substantial computational requirements have constrained their deployment in real-time interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Haojie Yu , Zhaonian Wang , Yihan Pan , Meng Cheng , Hao Yang , Chao Wang , Tao Xie , Xiaoming Xu , Xiaoming Wei , Xunliang Cai

Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mainly use sequence-to-sequence models to solve this problem,…

Sound · Computer Science 2022-07-14 Yongqi Wang , Zhou Zhao

With the availability of massive general-domain dialogue data, pre-trained dialogue generation appears to be super appealing to transfer knowledge from the general domain to downstream applications. In most existing work, such transferable…

Computation and Language · Computer Science 2022-10-25 Xueliang Zhao , Lemao Liu , Tingchen Fu , Shuming Shi , Dongyan Zhao , Rui Yan

The introduction of diffusion models has brought significant advances to the field of audio-driven talking head generation. However, the extremely slow inference speed severely limits the practical implementation of diffusion-based talking…

Graphics · Computer Science 2025-11-18 Haotian Wang , Yuzhe Weng , Jun Du , Haoran Xu , Xiaoyan Wu , Shan He , Bing Yin , Cong Liu , Jianqing Gao , Qingfeng Liu

Cross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to-speech (TTS) model,…

Sound · Computer Science 2021-02-04 Shengkui Zhao , Hao Wang , Trung Hieu Nguyen , Bin Ma

Current video-to-audio (V2A) methods struggle in complex multi-event scenarios (video scenarios involving multiple sound sources, sound events, or transitions) due to two critical limitations. First, existing methods face challenges in…

Multimedia · Computer Science 2025-11-05 Jianxuan Yang , Xiaoran Yang , Lipan Zhang , Xinyue Guo , Zhao Wang , Gongping Huang

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

Text-to-speech synthesis (TTS) has witnessed rapid progress in recent years, where neural methods became capable of producing audios with high naturalness. However, these efforts still suffer from two types of latencies: (a) the {\em…

Computation and Language · Computer Science 2020-10-08 Mingbo Ma , Baigong Zheng , Kaibo Liu , Renjie Zheng , Hairong Liu , Kainan Peng , Kenneth Church , Liang Huang
‹ Prev 1 3 4 5 6 7 10 Next ›