中文
相关论文

相关论文: UniArray: Unified Spectral-Spatial Modeling for Ar…

200 篇论文

Speech enhancement (SE) is usually required as a front end to improve the speech quality in noisy environments, while the enhanced speech might not be optimal for automatic speech recognition (ASR) systems due to speech distortion. On the…

音频与语音处理 · 电气工程与系统科学 2022-05-27 Qiu-Shi Zhu , Jie Zhang , Zi-Qiang Zhang , Li-Rong Dai

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal enhancement. However, the textbook solutions for optimal…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Reinhold Haeb-Umbach , Tomohiro Nakatani , Marc Delcroix , Christoph Boeddeker , Tsubasa Ochiai

Source separation can improve automatic speech recognition (ASR) under multi-party meeting scenarios by extracting single-speaker signals from overlapped speech. Despite the success of self-supervised learning models in single-channel…

音频与语音处理 · 电气工程与系统科学 2023-04-04 Yuang Li , Xianrui Zheng , Philip C. Woodland

This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This framework integrates models focused on three distinct tasks:…

音频与语音处理 · 电气工程与系统科学 2025-09-18 Younghoo Kwon , Dongheon Lee , Dohwan Kim , Jung-Woo Choi

ASR has been shown to achieve great performance recently. However, most of them rely on massive paired data, which is not feasible for low-resource languages worldwide. This paper investigates how to learn directly from unpaired phone…

声音 · 计算机科学 2022-08-01 Da-rong Liu , Po-chun Hsu , Yi-chen Chen , Sung-feng Huang , Shun-po Chuang , Da-yi Wu , Hung-yi Lee

Recently, Multi-Contrast MR Reconstruction (MCMR) has emerged as a hot research topic that leverages high-quality auxiliary modalities to reconstruct undersampled target modalities of interest. However, existing methods often struggle to…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Jialin Li , Yiwei Ren , Kai Pan , Dong Wei , Pujin Cheng , Xian Wu , Xiaoying Tang

Microphone array post-filters have demonstrated their ability to greatly reduce noise at the output of a beamformer. However, current techniques only consider a single source of interest, most of the time assuming stationary background…

声音 · 计算机科学 2016-03-11 Jean-Marc Valin , Jean Rouat , François Michaud

In reverberant conditions with multiple concurrent speakers, each microphone acquires a mixture signal of multiple speakers at a different location. In over-determined conditions where the microphones out-number speakers, we can narrow down…

声音 · 计算机科学 2023-10-31 Zhong-Qiu Wang , Shinji Watanabe

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omni-directional…

音频与语音处理 · 电气工程与系统科学 2024-01-18 Yang Yang , George Sung , Shao-Fu Shih , Hakan Erdogan , Chehung Lee , Matthias Grundmann

Current multimodal medical image fusion typically assumes that source images are of high quality and perfectly aligned at the pixel level. Its effectiveness heavily relies on these conditions and often deteriorates when handling misaligned…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Dayong Su , Yafei Zhang , Huafeng Li , Jinxing Li , Yu Liu

Linear superarrays (LSAs) have been proposed to address the limited steering capability of conventional linear differential microphone arrays (LDMAs) by integrating omnidirectional and directional microphones, enabling more flexible…

信号处理 · 电气工程与系统科学 2025-11-03 Yuanhang Qian , Xueqin Luo , Jilu Jin , Gongping Huang , Jingdong Chen , Jacob Benesty

Multisequence Magnetic Resonance Imaging (MRI) provides a more reliable diagnosis in clinical applications through complementary information across sequences. However, in practice, the absence of certain MR sequences is a common problem…

图像与视频处理 · 电气工程与系统科学 2025-10-21 Jihoon Cho , Jonghye Woo , Jinah Park

Target speech separation refers to extracting a target speaker's voice from an overlapped audio of simultaneous talkers. Previously the use of visual modality for target speech separation has demonstrated great potentials. This work…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Lianwu Chen , Yuexian Zou , Dong Yu

Audiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural network…

计算机视觉与模式识别 · 计算机科学 2018-05-01 Michael Wand , Ngoc Thang Vu , Juergen Schmidhuber

The data-hungry approach of supervised classification drives the interest of the researchers toward unsupervised approaches, especially for problems such as medical image segmentation, where labeled data are difficult to get. Motivated by…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Kovvuri Sai Gopal Reddy , Bodduluri Saran , A. Mudit Adityaja , Saurabh J. Shigwan , Nitin Kumar , Snehasis Mukherjee

Separating two sources from an audio mixture is an important task with many applications. It is a challenging problem since only one signal channel is available for analysis. In this paper, we propose a novel framework for singing voice…

声音 · 计算机科学 2017-11-15 Zhe-Cheng Fan , Yen-Lin Lai , Jyh-Shing Roger Jang

The recent trends of densification and centralized signal processing in radio access networks suggest that future networks may comprise ubiquitous antennas coordinated to form a network-wide gigantic array, referred to as the ubiquitous…

信息论 · 计算机科学 2015-04-01 Kaibin Huang , Jiayi Chen , Vincent K. N. Lau

This paper proposes a universal sound separation (USS) method capable of handling untrained sampling frequencies (SFs). The USS aims at separating arbitrary sources of different types and can be the key technique to realize a source…

音频与语音处理 · 电气工程与系统科学 2023-09-25 Tomohiko Nakamura , Kohei Yatabe

Using deep neural networks (DNNs) for encoding of microphone array (MA) signals to the Ambisonics spatial audio format can surpass certain limitations of established conventional methods, but existing DNN-based methods need to be trained…

音频与语音处理 · 电气工程与系统科学 2025-01-15 Mikko Heikkinen , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

Generative modeling has recently achieved remarkable success across image, video, and audio domains, demonstrating powerful capabilities for unified representation learning. Yet speech front-end tasks such as speech enhancement (SE), target…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Ziqian Wang , Zikai Liu , Yike Zhu , Xingchen Li , Boyi Kang , Jixun Yao , Xianjun Xia , Chuanzeng Huang , Lei Xie