English
Related papers

Related papers: AnyAccomp: Generalizable Accompaniment Generation …

200 papers

Audio-driven talking-head generation has achieved remarkable progress with recent models such as AniTalker, FLOAT, and Sonic. Despite their success, most existing approaches rely on a single static reference image to condition the entire…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zhicheng Zhang , Lei Wang , Yu Zhang , Yongsheng Gao

We propose Composition Sampling, a simple but effective method to generate diverse outputs for conditional generation of higher quality compared to previous stochastic decoding strategies. It builds on recently proposed plan-based neural…

Computation and Language · Computer Science 2022-03-30 Shashi Narayan , Gonçalo Simões , Yao Zhao , Joshua Maynez , Dipanjan Das , Michael Collins , Mirella Lapata

Anomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits traditional detection methods, making anomaly generation…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Xuan Tong , Yang Chang , Qing Zhao , Jiawen Yu , Boyang Wang , Junxiong Lin , Yuxuan Lin , Xinji Mai , Haoran Wang , Zeng Tao , Yan Wang , Wenqiang Zhang

There have been a number of techniques that have demonstrated the generation of multimedia data for one modality at a time using GANs, such as the ability to generate images, videos, and audio. However, so far, the task of multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Vinod K Kurmi , Vipul Bajaj , Badri N Patro , K S Venkatesh , Vinay P Namboodiri , Preethi Jyothi

Recent advancements in song generation have shown promising results in generating songs from lyrics and/or global text prompts. However, most existing systems lack the ability to model the temporally varying attributes of songs, limiting…

Sound · Computer Science 2026-05-29 Pengfei Cai , Joanna Wang , Haorui Zheng , Xu Li , Zihao Ji , Teng Ma , Zhongliang Liu , Chen Zhang , Pengfei Wan

Existing singing voice synthesis models (SVS) are usually trained on singing data and depend on either error-prone time-alignment and duration features or explicit music score information. In this paper, we propose Karaoker, a multispeaker…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-30 Panos Kakoulidis , Nikolaos Ellinas , Georgios Vamvoukakis , Konstantinos Markopoulos , June Sig Sung , Gunu Jho , Pirros Tsiakoulis , Aimilios Chalamandaris

The proliferation of generative AI has led to hyper-realistic synthetic videos, escalating misuse risks and outstripping binary real/fake detectors. We introduce SAGA (Source Attribution of Generative AI videos), the first comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Rohit Kundu , Vishal Mohanty , Hao Xiong , Shan Jia , Athula Balachandran , Amit K. Roy-Chowdhury

Melody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key…

Sound · Computer Science 2025-02-10 Wei Chen , Binzhu Sha , Jing Yang , Zhuo Wang , Fan Fan , Zhiyong Wu

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled…

Multimodal music creation requires models that can both generate audio from high-level cues and edit existing mixtures in a targeted manner. Yet most multimodal music systems are built for a single task and a fixed prompting interface,…

Electrocardiogram (ECG), a non-invasive and affordable tool for cardiac monitoring, is highly sensitive in detecting acute heart attacks. However, due to the lengthy nature of ECG recordings, numerous machine learning methods have been…

Signal Processing · Electrical Eng. & Systems 2025-03-04 Yue Wang , Xu Cao , Yaojun Hu , Haochao Ying , Hongxia Xu , Ruijia Wu , James Matthew Rehg , Jimeng Sun , Jian Wu , Jintai Chen

Recent studies in singing voice synthesis have achieved high-quality results leveraging advances in text-to-speech models based on deep neural networks. One of the main issues in training singing voice synthesis models is that they require…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-15 Soonbeom Choi , Juhan Nam

Music generation models can produce high-fidelity coherent accompaniment given complete audio input, but are limited to editing and loop-based workflows. We study real-time audio-to-audio accompaniment: as a model hears an input audio…

The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-29 Chenyu Yang , Shuai Wang , Hangting Chen , Jianwei Yu , Wei Tan , Rongzhi Gu , Yaoxun Xu , Yizhi Zhou , Haina Zhu , Haizhou Li

Singing Voice Conversion (SVC) aims to transform a source singing voice into a target singer while preserving lyrics and melody. Most existing SVC methods depend on F0 extractors to capture the lead melody from clean vocals. However, no…

Sound · Computer Science 2026-05-13 Chen Geng , Meng Chen , Ruohua Zhou , Ruolan Liu , Weifeng Zhao

Electronic music artists and sound designers have unique workflow practices that necessitate specialized approaches for developing music information retrieval and creativity support tools. Furthermore, electronic music instruments, such as…

Sound · Computer Science 2022-10-28 Olga Vechtomova , Gaurav Sahu

Recent advances in audio generation have focused on text-to-audio (T2A) and video-to-audio (V2A) tasks. However, T2A or V2A methods cannot generate holistic sounds (onscreen and off-screen). This is because T2A cannot generate sounds…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Saksham Singh Kushwaha , Yapeng Tian

Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos. Following the perspective of extracting essential signals from videos that can precisely control the mature…

Sound · Computer Science 2025-03-11 Juncheng Wang , Chao Xu , Cheng Yu , Lei Shang , Zhe Hu , Shujun Wang , Liefeng Bo

Despite significant recent progress in the area of Brain-Computer Interface (BCI), there are numerous shortcomings associated with collecting Electroencephalography (EEG) signals in real-world environments. These include, but are not…

Quantitative Methods · Quantitative Biology 2019-10-14 Nik Khadijah Nik Aznan , Amir Atapour-Abarghouei , Stephen Bonner , Jason Connolly , Noura Al Moubayed , Toby Breckon

Deep generative models have led to significant advances in cross-modal generation such as text-to-image synthesis. Training these models typically requires paired data with direct correspondence between modalities. We introduce the novel…

Computer Vision and Pattern Recognition · Computer Science 2019-08-21 Shuang Ma , Daniel McDuff , Yale Song
‹ Prev 1 3 4 5 6 7 10 Next ›