中文
相关论文

相关论文: WildMix Dataset and Spectro-Temporal Transformer M…

200 篇论文

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

声音 · 计算机科学 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

The separation of single-channel underwater acoustic signals is a challenging problem with practical significance. Few existing studies focus on the source separation problem with unknown numbers of signals, and how to evaluate the…

声音 · 计算机科学 2024-05-29 Qinggang Sun , Kejun Wang

Cinematic audio source separation (CASS), as a problem of extracting the dialogue, music, and effects stems from their mixture, is a relatively new subtask of audio source separation. To date, only one publicly available dataset exists for…

音频与语音处理 · 电气工程与系统科学 2024-08-27 Karn N. Watcharasupat , Chih-Wei Wu , Iroro Orife

Modern audio source separation techniques rely on optimizing sequence model architectures such as, 1D-CNNs, on mixture recordings to generalize well to unseen mixtures. Specifically, recent focus is on time-domain based architectures such…

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, however, the order of…

音频与语音处理 · 电气工程与系统科学 2023-03-15 Cong Han , Nima Mesgarani

Spatial transcriptomics (ST) enables transcriptome-wide profiling while preserving the spatial context of tissues, offering unprecedented opportunities to study tissue organization and cell-cell interactions in situ. Despite recent…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Wei Wang , Quoc-Toan Ly , Chong Yu , Jun Bai

Speech separation is very important in real-world applications such as human-machine interaction, hearing aids devices, and automatic meeting transcription. In recent years, a significant improvement occurred towards the solution based on…

声音 · 计算机科学 2024-08-29 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

Source separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the success of self-supervised learning models in single-channel…

音频与语音处理 · 电气工程与系统科学 2023-04-04 Yuang Li , Xianrui Zheng , Philip C. Woodland

This paper addresses the challenge of speaker separation, which remains an active research topic despite the promising results achieved in recent years. These results, however, often degrade in real recording conditions due to the presence…

声音 · 计算机科学 2024-11-14 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

Recently, supervised speech separation has made great progress. However, limited by the nature of supervised training, most existing separation methods require ground-truth sources and are trained on synthetic datasets. This ground-truth…

音频与语音处理 · 电气工程与系统科学 2022-08-09 Jiangyu Han , Yanhua Long

Music demixing is the task of separating different tracks from the given single audio signal into components, such as drums, bass, and vocals from the rest of the accompaniment. Separation of sources is useful for a range of areas,…

声音 · 计算机科学 2024-05-08 Roman Solovyev , Alexander Stempkovskiy , Tatiana Habruseva

Spatial transcriptomics (ST) has revolutionised transcriptomics analysis by preserving tissue architecture, allowing researchers to study gene expression in its native spatial context. However, despite its potential, ST still faces…

定量方法 · 定量生物学 2025-05-19 Anthony Baptista , Rosamond Nuamah , Ciro Chiappini , Anita Grigoriadis

Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models' performance. In this paper, we propose a decoupled…

声音 · 计算机科学 2020-10-29 Shuai Zhang , Jiangyan Yi , Zhengkun Tian , Ye Bai , Jianhua Tao , Zhengqi wen

With the recent advancements of data driven approaches using deep neural networks, music source separation has been formulated as an instrument-specific supervised problem. While existing deep learning models implicitly absorb the spatial…

音频与语音处理 · 电气工程与系统科学 2022-02-16 Darius Petermann , Minje Kim

Deep neural networks have become an indispensable technique for audio source separation (ASS). It was recently reported that a variant of CNN architecture called MMDenseNet was successfully employed to solve the ASS problem of estimating…

声音 · 计算机科学 2018-05-30 Naoya Takahashi , Nabarun Goswami , Yuki Mitsufuji

Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach,…

声音 · 计算机科学 2025-11-26 Genís Plaja-Roglans , Yun-Ning Hung , Xavier Serra , Igor Pereira

We propose a time-domain audio source separation method using down-sampling (DS) and up-sampling (US) layers based on a discrete wavelet transform (DWT). The proposed method is based on one of the state-of-the-art deep neural networks,…

声音 · 计算机科学 2022-12-05 Tomohiko Nakamura , Hiroshi Saruwatari

We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources…

计算机视觉与模式识别 · 计算机科学 2021-08-27 Sagnik Majumder , Ziad Al-Halah , Kristen Grauman

Transformers have drawn attention in the MIR field for their remarkable performance shown in natural language processing and computer vision. However, prior works in the audio processing domain mostly use Transformer as a temporal feature…

声音 · 计算机科学 2021-10-26 Wei-Tsung Lu , Ju-Chiang Wang , Minz Won , Keunwoo Choi , Xuchen Song

Music source separation performance has greatly improved in recent years with the advent of approaches based on deep learning. Such methods typically require large amounts of labelled training data, which in the case of music consist of…

声音 · 计算机科学 2019-09-19 Ethan Manilow , Gordon Wichern , Prem Seetharaman , Jonathan Le Roux