English
Related papers

Related papers: Can all variations within the unified mask-based b…

200 papers

Self-supervised learning through masked autoencoders has attracted great attention for remote sensing (RS) foundation model (FM) development, enabling improved representation learning across diverse sensors and downstream tasks. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Leonard Hackel , Tom Burgert , Begüm Demir

To accommodate the explosive wireless traffics, massive multiple-input multiple-output (MIMO) is regarded as one of the key enabling technologies for next-generation communication systems. In massive MIMO cellular networks, coordinated…

Information Theory · Computer Science 2023-03-27 Jungang Ge , Ying-Chang Liang , Liao Zhang , Ruizhe Long , Sumei Sun

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT)…

Sound · Computer Science 2019-01-03 Xingjian Du , Mengyao Zhu , Xuan Shi , Xinpeng Zhang , Wen Zhang , Jingdong Chen

This work presents a statistical analysis of a class of jointly optimized beamformer-assisted acoustic echo cancelers (AEC) with the beamformer (BF) implemented in the Generalized Sidelobe Canceler (GSC) form and using the least-mean square…

Statistics Theory · Mathematics 2015-03-06 Marcos H. Maruo , José C. M. Bermudez , Leonardo S. Resende

We develop an unsupervised deep learning framework for real-time scalable and generalizable downlink beamforming in multi-user multiple-input single-output (MU-MISO) systems. The proposed semi-amortized lifted learning-to-optimize (SALLO)…

Machine Learning · Computer Science 2026-04-01 Yubo Zhang , Xiao-Yang Liu , Xiaodong Wang

Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of different types, a task we…

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-12 Ke Zhang , Junjie Li , Shuai Wang , Yangjie Wei , Yi Wang , Yannan Wang , Haizhou Li

In this paper, we investigate a unified linear transceiver design with mean-square-error (MSE) as the objective function for a wide range of wireless systems. The unified design is based on an elegant mathematical programming technology…

Information Theory · Computer Science 2013-01-10 Chengwen Xing , Zesong Fei , Shaodan Ma , Jingming Kuang , Yik-Chung Wu

Resolution Enhancement Techniques (RETs) are critical to meet the demands of advanced technology nodes. Among RETs, Source Mask Optimization (SMO) is pivotal, concurrently optimizing both the source and the mask to expand the process…

Signal Processing · Electrical Eng. & Systems 2024-05-17 Guojin Chen , Hongquan He , Peng Xu , Hao Geng , Bei Yu

Generative target speaker extraction (TSE) methods often produce more natural outputs than predictive models. Recent work based on diffusion or flow matching (FM) typically relies on a small, fixed number of reverse steps with a fixed step…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-21 Tsun-An Hsieh , Minje Kim

In this paper, we propose an Expectation-Maximization-based (EM) Personalized Federated Learning (PFL) framework for multi-objective optimization (MOO) in Integrated Sensing and Communication (ISAC) systems. In contrast to standard…

Signal Processing · Electrical Eng. & Systems 2025-10-09 Zhou Ni , Sravan Reddy Chintareddy , Peiyuan Guan , Morteza Hashemi

This paper investigates the performance of Binaural Signal Matching (BSM) methods for near-field sound reproduction using a wearable glasses-mounted microphone array. BSM is a flexible, signal-independent approach for binaural rendering…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Sapir Goldring , Zamir Ben Hur , David Lou Alon , Chad McKell , Sebastian Prepelita , Boaz Rafaely

Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Riki Shimizu , Xilin Jiang , Nima Mesgarani

Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture of sources. It often uses a model conditioned on a fixed form…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-16 Chenda Li , Yao Qian , Zhuo Chen , Dongmei Wang , Takuya Yoshioka , Shujie Liu , Yanmin Qian , Michael Zeng

This work studies the beamforming design in the joint proactive eavesdropping (PE) and target sensing (TS) systems. The base station (BS) wiretaps the information transmitted by the illegal transmitter and sends the waveform for TS. The…

Information Theory · Computer Science 2025-12-30 Qian Dan , Hongjiang Lei , Ki-Hong Park , Gaofeng Pan , Mohamed-Slim Alouini

Spatial clustering techniques can achieve significant multi-channel noise reduction across relatively arbitrary microphone configurations, but have difficulty incorporating a detailed speech/noise model. In contrast, LSTM neural networks…

Sound · Computer Science 2020-12-07 Zhaoheng Ni , Felix Grezes , Viet Anh Trinh , Michael I. Mandel

In this paper we present the first on-sky results with the fibered aperture masking instrument FIRST. Its principle relies on the combination of spatial filtering and aperture masking using single-mode fibers, a novel technique that is…

Instrumentation and Methods for Astrophysics · Physics 2015-06-04 E. Huby , G. Perrin , F. Marchis , S. Lacour , T. Kotani , G. Duchêne , E. Choquet , E. L. Gates , J. M. Woillez , O. Lai , P. Fédou , C. Collin , F. Chapron , V. Arslanyan , K. J. Burns

In this paper, we consider multi-quality multicast beamforming of a video stream from a multi-antenna base station (BS) to multiple single-antenna users receiving different qualities of the same video stream, via scalable video coding…

Information Theory · Computer Science 2017-07-18 Chengjun Guo , Ying Cui , Derrick Wing Kwan Ng , Zhi Liu

Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spectrogram mask for the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-22 Hao Ma , Rujin Chen , Xiao-Lei Zhang , Ju Liu , Xuelong Li

In diffusion MRI (dMRI), a good sampling scheme is important for efficient acquisition and robust reconstruction. Diffusion weighted signal is normally acquired on single or multiple shells in q-space. Signal samples are typically…

Medical Physics · Physics 2017-09-26 Jian Cheng , Dinggang Shen , Pew-Thian Yap , Peter J. Basser