中文
相关论文

相关论文: UniArray: Unified Spectral-Spatial Modeling for Ar…

200 篇论文

Recently, a spatially selective non-linear filter (SSF) has been proposed for target speaker extraction, using the target direction-of-arrival (DOA) as a spatial cue. Since learned intermediate features are tied to the microphone geometry,…

音频与语音处理 · 电气工程与系统科学 2026-05-19 Jiatong Li , Wiebke Middelberg , Simon Doclo

Speech separation seeks to isolate individual speech signals from a multi-talk speech mixture. Despite much progress, a system well-trained on synthetic data often experiences performance degradation on out-of-domain data, such as…

声音 · 计算机科学 2025-03-18 Wupeng Wang , Zexu Pan , Jingru Lin , Shuai Wang , Haizhou Li

The capability of the human to pay attention to both coarse and fine-grained regions has been applied to computer vision tasks. Motivated by that, we propose a collaborative learning framework in the complex domain for monaural noise…

声音 · 计算机科学 2021-06-23 Andong Li , Chengshi Zheng , Lu Zhang , Xiaodong Li

Single-cell RNA sequencing (scRNA-seq) profiles large numbers of cells but loses spatial context, whereas spatial transcriptomics (ST) preserves partial spatial structure at lower resolution. Most existing integration methods either…

机器学习 · 计算机科学 2026-05-28 Ehtesamul Azim , Muhtasim Noor Alif , Tae Hyun Hwang , Yanjie Fu , Wei Zhang

The majority of multichannel speech enhancement algorithms are two-step procedures that first apply a linear spatial filter, a so-called beamformer, and combine it with a single-channel approach for postprocessing. However, the serial…

音频与语音处理 · 电气工程与系统科学 2021-06-16 Kristina Tesch , Timo Gerkmann

This paper tackles automatically discovering phone-like acoustic units (AUD) from unlabeled speech data. Past studies usually proposed single-step approaches. We propose a two-stage approach: the first stage learns a subword-discriminative…

音频与语音处理 · 电气工程与系统科学 2021-06-08 Siyuan Feng , Piotr Żelasko , Laureano Moro-Velázquez , Odette Scharenborg

Acoustic beamforming with a microphone array represents an adequate technology for remote acoustic surveillance, as the system has no mechanical parts and it has moderate size. However, in order to accomplish real implementation, several…

声音 · 计算机科学 2017-03-08 Abdulla AlShehhi , M. Luai Hammadih , M. Sami Zitouni , Saif AlKindi , Nazar Ali , Luis Weruaga

Synthetic Aperture Radar (SAR) images are inherently corrupted by speckle noise, limiting their utility in high-precision applications. While deep learning methods have shown promise in SAR despeckling, most methods employ a single unified…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Ziqing Ma , Chang Yang , Zhichang Guo , Yao Li

Separating target speech from mixed signals containing flexible speaker quantities presents a challenging task. While existing methods demonstrate strong separation performance and noise robustness, they predominantly assume prior knowledge…

音频与语音处理 · 电气工程与系统科学 2025-07-18 Daning Zhang , Ying Wei

Audio-Visual Segmentation (AVS) aims to identify and segment sound-producing objects in videos by leveraging both visual and audio modalities. It has emerged as a significant research area in multimodal perception, enabling fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Jia Li , Yapeng Tian

Ultrasound is widely used in clinical practice due to its affordability, portability, and safety. However, current AI research often overlooks combined disease prediction and tissue segmentation. We propose UniUSNet, a universal framework…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Zehui Lin , Zhuoneng Zhang , Xindi Hu , Zhifan Gao , Xin Yang , Yue Sun , Dong Ni , Tao Tan

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

Background noise, interfering speech and room reverberation frequently distort target speech in real listening environments. In this study, we address joint speech separation and dereverberation, which aims to separate target speech from…

音频与语音处理 · 电气工程与系统科学 2020-07-15 Ke Tan , Yong Xu , Shi-Xiong Zhang , Meng Yu , Dong Yu

Accurate lesion segmentation in ultrasound images is essential for preventive screening and clinical diagnosis, yet remains challenging due to low contrast, blurry boundaries, and significant scale variations. Although existing deep…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Chen Wang , Yixin Zhu , Yongbin Zhu , Fengyuan Shi , Qi Li , Jun Wang , Zuozhu Liu , Keli Hu

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets

While existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios.…

声音 · 计算机科学 2024-07-31 Tianrui Pan , Jie Liu , Bohan Wang , Jie Tang , Gangshan Wu

In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker counting, diarization, and ASR. It handles various recording…

In this paper, we introduce an unsupervised approach for Speech Segmentation, which builds on previously researched approaches, e.g., Speaker Diarization, while being applicable to an inclusive set of acoustic-semantic distinctions, paving…

计算与语言 · 计算机科学 2025-01-08 Avishai Elmakies , Omri Abend , Yossi Adi

We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited…

Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR). On the other…

音频与语音处理 · 电气工程与系统科学 2020-08-03 Yong Xu , Meng Yu , Shi-Xiong Zhang , Lianwu Chen , Chao Weng , Jianming Liu , Dong Yu