English
Related papers

Related papers: UniArray: Unified Spectral-Spatial Modeling for Ar…

200 papers

Traditional image stitching methods estimate warps from hand-crafted geometric features, whereas recent learning-based solutions leverage semantic features from neural networks instead. These two lines of research have largely diverged…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yuan Mei , Lang Nie , Kang Liao , Yunqiu Xu , Chunyu Lin , Bin Xiao

We enhance the vanilla adversarial training method for unsupervised Automatic Speech Recognition (ASR) by a diffusion-GAN. Our model (1) injects instance noises of various intensities to the generator's output and unlabeled reference text…

Computation and Language · Computer Science 2023-03-27 Xianchao Wu

The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-03 Tianqin Zheng , Jilu Jin , Hanchen Pei , Gongping Huang , Jingdong Chen , Jacob Benesty

A method of binaural rendering from microphone array signals of arbitrary geometry is proposed. To reproduce binaural signals from microphone array recordings at a remote location, a spherical microphone array is generally used for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-19 Naoto Iijima , Shoichi Koyama , Hiroshi Saruwatari

Distant speech processing is a challenging task, especially when dealing with the cocktail party effect. Sound source separation is thus often required as a preprocessing step prior to speech recognition to improve the signal to distortion…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Francois Grondin , Jean-Samuel Lauzon , Jonathan Vincent , Francois Michaud

This paper describes an efficient unsupervised learning method for a neural source separation model that utilizes a probabilistic generative model of observed multichannel mixtures proposed for blind source separation (BSS). For this…

Sound · Computer Science 2023-06-21 Yoshiaki Bando , Yoshiki Masuyama , Aditya Arie Nugraha , Kazuyoshi Yoshii

Automatic speech recognition (ASR) has been widely researched with supervised approaches, while many low-resourced languages lack audio-text aligned data, and supervised methods cannot be applied on them. In this work, we propose a…

Computation and Language · Computer Science 2018-08-14 Yi-Chen Chen , Chia-Hao Shen , Sung-Feng Huang , Hung-yi Lee

The problem of multi-objective design of sparse MIMO arrays for better multitarget detection capabilities is considered. A novel approach for efficient utilization of the antenna design resources; namely, the number of available array…

Signal Processing · Electrical Eng. & Systems 2022-10-24 Suleyman Gokhun Tanyer , Paul Dent , Murtaza Ali , Curtis Davis , SenthinelKumar Rajagopal , Peter Driessen

Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only…

Sound · Computer Science 2025-05-28 Jiatong Shi , Hye-Jin Shim , Shinji Watanabe

Speech enhancement and separation have been a long-standing problem, especially with the recent advances using a single microphone. Although microphones perform well in constrained settings, their performance for speech separation decreases…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-15 Muhammed Zahid Ozturk , Chenshu Wu , Beibei Wang , Min Wu , K. J. Ray Liu

Traditional speech systems typically rely on separate, task-specific models for text-to-speech (TTS), automatic speech recognition (ASR), and voice conversion (VC), resulting in fragmented pipelines that limit scalability, efficiency, and…

Sound · Computer Science 2026-01-19 Runyuan Cai , Yu Lin , Yiming Wang , Chunlin Fu , Xiaodong Zeng

Continuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common application scenario is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-14 Zhuohuang Zhang , Takuya Yoshioka , Naoyuki Kanda , Zhuo Chen , Xiaofei Wang , Dongmei Wang , Sefik Emre Eskimez

Multimodal MRI offers complementary information for brain tumor segmentation, but clinical scans often lack one or more modalities, which degrades segmentation performance. In this paper, we propose UniME (Uni-Encoder Meets Multi-Encoders),…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Peibo Song , Xiaotian Xue , Jinshuo Zhang , Zihao Wang , Jinhua Liu , Shujun Fu , Fangxun Bao , Si Yong Yeo

Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-12 Jia Qi Yip , Shengkui Zhao , Yukun Ma , Chongjia Ni , Chong Zhang , Hao Wang , Trung Hieu Nguyen , Kun Zhou , Dianwen Ng , Eng Siong Chng , Bin Ma

Dysarthria is a disability that causes a disturbance in the human speech system and reduces the quality and intelligibility of a person's speech. Because of this effect, the normal speech processing systems can not work properly on impaired…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-22 Aref Farhadipour , Hadi Veisi

This paper presents a neural method for distant speech recognition (DSR) that jointly separates and diarizes speech mixtures without supervision by isolated signals. A standard separation method for multi-talker DSR is a statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yoshiaki Bando , Tomohiko Nakamura , Shinji Watanabe

In this work, we address the challenge of generalizable audio deepfake detection (ADD) across diverse speech synthesis paradigms-including conventional text-to-speech (TTS) systems and modern diffusion or flow-matching (FM) based…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-17 Farhan Sheth , Girish , Mohd Mujtaba Akhtar , Muskaan Singh

The surge of massive antenna arrays in wireless networks calls for the adoption of analog/hybrid array solutions, where multiple antenna elements are driven by a common radio front end to form a beam along a specific angle in order to…

Signal Processing · Electrical Eng. & Systems 2024-10-28 Silvio Mandelli , Marcus Henninger , Jinfeng Du

With the deployment of large antenna arrays at high-frequency bands, future wireless communication systems are likely to operate in the radiative near-field (NF). Unlike far-field beam steering, NF beams can be focused on a spatial region…

Signal Processing · Electrical Eng. & Systems 2026-03-13 Ahmed Hussain , Asmaa Abdallah , Abdulkadir Celik , Ahmed M. Eltawil

Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing methods mainly face two challenges, the difficulty in jointly…

Sound · Computer Science 2025-12-03 Xinlei Yin , Xiulian Peng , Xue Jiang , Zhiwei Xiong , Yan Lu