中文
相关论文

相关论文: Vevo2: A Unified and Controllable Framework for Sp…

200 篇论文

High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not meet requirements for…

音频与语音处理 · 电气工程与系统科学 2021-12-21 Rongjie Huang , Feiyang Chen , Yi Ren , Jinglin Liu , Chenye Cui , Zhou Zhao

We propose a semi-supervised singing synthesizer, which is able to learn new voices from audio data only, without any annotations such as phonetic segmentation. Our system is an encoder-decoder model with two encoders, linguistic and…

声音 · 计算机科学 2020-11-06 Jordi Bonada , Merlijn Blaauw

This paper proposes a controllable singing voice synthesis system capable of generating expressive singing voice with two novel methodologies. First, a local style token module, which predicts frame-level style tokens from an input pitch…

声音 · 计算机科学 2022-04-08 Juheon Lee , Hyeong-Seok Choi , Kyogu Lee

We propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables…

声音 · 计算机科学 2025-01-24 Shuqi Dai , Yunyun Wang , Roger B. Dannenberg , Zeyu Jin

We investigate the feasibility of a singing voice synthesis (SVS) system by using a decomposed framework to improve flexibility in generating singing voices. Due to data-driven approaches, SVS performs a music score-to-waveform mapping;…

声音 · 计算机科学 2024-07-15 Lester Phillip Violeta , Taketo Akama

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthesis models currently rely on annotated audio data, but it is…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Rongjie Huang , Chunlei Zhang , Yongqi Wang , Dongchao Yang , Luping Liu , Zhenhui Ye , Ziyue Jiang , Chao Weng , Zhou Zhao , Dong Yu

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the…

声音 · 计算机科学 2023-12-29 Zhifang Guo , Jianguo Mao , Rui Tao , Long Yan , Kazushige Ouchi , Hong Liu , Xiangdong Wang

Significant strides have been made in creating voice identity representations using speech data. However, the same level of progress has not been achieved for singing voices. To bridge this gap, we suggest a framework for training singer…

声音 · 计算机科学 2024-01-11 Bernardo Torres , Stefan Lattner , Gaël Richard

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model…

Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Zhipeng Chen , Lan Yang , Yonggang Qi , Honggang Zhang , Kaiyue Pang , Ke Li , Yi-Zhe Song

Recently, denoising diffusion models have demonstrated remarkable performance among generative models in various domains. However, in the speech domain, the application of diffusion models for synthesizing time-varying audio faces…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Ji-Sang Hwang , Sang-Hoon Lee , Seong-Whan Lee

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Jiaben Chen , Xin Yan , Yihang Chen , Siyuan Cen , Zixin Wang , Qinwei Ma , Haoyu Zhen , Kaizhi Qian , Lie Lu , Chuang Gan

Automatic transcription of monophonic/polyphonic music is a challenging task due to the lack of availability of large amounts of transcribed data. In this paper, we propose a data augmentation method that converts natural speech to singing…

声音 · 计算机科学 2021-02-18 Sakya Basak , Shrutina Agarwal , Sriram Ganapathy , Naoya Takahashi

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and…

多媒体 · 计算机科学 2024-10-01 Kun Su , Xiulong Liu , Eli Shlizerman

While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single song. To address…

声音 · 计算机科学 2026-02-10 Jiatao Chen , Xing Tang , Xiaoyue Duan , Yutang Feng , Jinchao Zhang , Jie Zhou

Regenerating singing voices with altered lyrics while preserving melody consistency remains challenging, as existing methods either offer limited controllability or require laborious manual alignment. We propose YingMusic-Singer-Plus, a…

音频与语音处理 · 电气工程与系统科学 2026-04-10 Chunbo Hao , Junjie Zheng , Guobin Ma , Yuepeng Jiang , Huakang Chen , Wenjie Tian , Gongyu Chen , Zihao Chen , Lei Xie

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

声音 · 计算机科学 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin