中文
相关论文

相关论文: LightSAFT: Lightweight Latent Source Aware Frequen…

200 篇论文

Multichannel speech enhancement (SE) aims to restore clean speech from noisy measurements by leveraging spatiotemporal signal features. In ad-hoc array conditions, microphone invariance (MI) requires systems to handle different microphone…

声音 · 计算机科学 2025-08-28 Haoyin Yan , Jie Zhang , Chengqian Jiang , Shuang Zhang

We propose a visually conditioned music remixing system by incorporating deep visual and audio models. The method is based on a state of the art audio-visual source separation model which performs music instrument source separation with…

声音 · 计算机科学 2020-10-29 Li-Chia Yang , Alexander Lerch

We present a neural network architecture able to efficiently detect modulation scheme in a portion of I/Q signals. This network is lighter by up to two orders of magnitude than other state-of-the-art architectures working on the same or…

机器学习 · 计算机科学 2021-11-29 Thomas Courtat , Hélion du Mas des Bourboux

In this paper, we consider the information-theoretic characterization of the set of achievable rates and distortions in a broad class of multiterminal communication scenarios with general continuous-valued sources and channels. A framework…

信息论 · 计算机科学 2022-02-24 Farhad Shirani , S. Sandeep Pradhan

In this report, we present our award-winning solutions for the Music Demixing Track of Sound Demixing Challenge 2023. First, we propose TFC-TDF-UNet v3, a time-efficient music source separation model that achieves state-of-the-art results…

声音 · 计算机科学 2023-07-24 Minseok Kim , Jun Hyung Lee , Soonyoung Jung

Patch-wise Transformer based time series forecasting achieves superior accuracy. However, this superiority relies heavily on intricate model design with massive parameters, rendering both training and inference expensive, thus preventing…

机器学习 · 计算机科学 2025-01-22 Meng Wang , Jintao Yang , Bin Yang , Hui Li , Tongxin Gong , Bo Yang , Jiangtao Cui

Heterogeneity across devices in federated learning (FL) typically refers to statistical (e.g., non-i.i.d. data distributions) and resource (e.g., communication bandwidth) dimensions. In this paper, we focus on another important dimension…

分布式、并行与集群计算 · 计算机科学 2024-01-10 Su Wang , Seyyedali Hosseinalipour , Christopher G. Brinton

Visual sound source separation aims at identifying sound components from a given sound mixture with the presence of visual cues. Prior works have demonstrated impressive results, but with the expense of large multi-stage architectures and…

计算机视觉与模式识别 · 计算机科学 2021-04-19 Lingyu Zhu , Esa Rahtu

In this paper, we address the intricate issue of RF signal separation by presenting a novel adaptation of the WaveNet architecture that introduces learnable dilation parameters, significantly enhancing signal separation in dense RF…

信号处理 · 电气工程与系统科学 2024-02-16 Yu Tian , Ahmed Alhammadi , Abdullah Quran , Abubakar Sani Ali

In recent years, how to strike a good trade-off between accuracy and inference speed has become the core issue for real-time semantic segmentation applications, which plays a vital role in real-world scenarios such as autonomous driving…

计算机视觉与模式识别 · 计算机科学 2021-07-19 Guangwei Gao , Guoan Xu , Yi Yu , Jin Xie , Jian Yang , Dong Yue

Flow maps enable high-quality image generation in a single forward pass. However, unlike iterative diffusion models, their lack of an explicit sampling trajectory impedes incorporating external constraints for conditional generation and…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Abbas Mammadov , So Takao , Bohan Chen , Ricardo Baptista , Morteza Mardani , Yee Whye Teh , Julius Berner

Data-driven models for audio source separation such as U-Net or Wave-U-Net are usually models dedicated to and specifically trained for a single task, e.g. a particular instrument isolation. Training them for various tasks at once commonly…

音频与语音处理 · 电气工程与系统科学 2019-11-22 Gabriel Meseguer-Brocal , Geoffroy Peeters

This paper proposes DroFiT (Drone Frequency lightweight Transformer for speech enhancement, a single microphone speech enhancement network for severe drone self-noise. DroFit integrates a frequency-wise Transformer with a full/sub-band…

音频与语音处理 · 电气工程与系统科学 2026-03-10 Jeongmin Lee , Chanhong Jeon , Hyungjoo Seo , Taewook Kang

Diffusion models excel at generating high-quality outputs but face challenges in data-scarce domains, where exhaustive retraining or costly paired data are often required. To address these limitations, we propose Latent Aligned Diffusion…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Xuqin Wang , Tao Wu , Yanfeng Zhang , Lu Liu , Dong Wang , Mingwei Sun , Yongliang Wang , Niclas Zeller , Daniel Cremers

This paper proposes StrEBM, a structured latent energy-based model for source-wise structured representation learning. The framework is motivated by a broader goal of promoting identifiable and decoupled latent organization by assigning…

机器学习 · 统计学 2026-04-21 Yuan-Hao Wei

Existing satellite remote sensing change detection (CD) methods often crop original large-scale bi-temporal image pairs into small patch pairs and then use pixel-level CD methods to fairly process all the patch pairs. However, due to the…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Lihui Xue , Zhihao Wang , Xueqian Wang , Gang Li

Diffusion-based foundation models have recently garnered much attention in the field of generative modeling due to their ability to generate images of high quality and fidelity. Although not straightforward, their recent application to the…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Nikos Kostagiolas , Pantelis Georgiades , Yannis Panagakis , Mihalis A. Nicolaou

Music source separation with deep neural networks typically relies only on amplitude features. In this paper we show that additional phase features can improve the separation performance. Using the theoretical relationship between STFT…

We consider transmission of a continuous amplitude source over an L-block Rayleigh fading $M_t \times M_r$ MIMO channel when the channel state information is only available at the receiver. Since the channel is not ergodic, Shannon's…

信息论 · 计算机科学 2007-11-09 Deniz Gunduz , Elza Erkip

Despite the rapid evolution of semantic segmentation for land cover classification in high-resolution remote sensing imagery, integrating multiple data modalities such as Digital Surface Model (DSM), RGB, and Near-infrared (NIR) remains a…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Tong Wang , Guanzhou Chen , Xiaodong Zhang , Chenxi Liu , Xiaoliang Tan , Jiaqi Wang , Chanjuan He , Wenlin Zhou