中文
相关论文

相关论文: Audio declipping performance enhancement via cross…

200 篇论文

Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the…

声音 · 计算机科学 2026-01-26 Thomas Serre , Mathieu Fontaine , Éric Benhaim , Slim Essid

The signal demixing problem seeks to separate a superposition of multiple signals into its constituent components. This paper studies a two-stage approach that first decompresses and subsequently deconvolves the noisy and undersampled…

信息检索 · 计算机科学 2022-05-25 Zhenan Fan , Halyun Jeong , Babhru Joshi , Michael P. Friedlander

The practice of speculative decoding, whereby inference is probabilistically supported by a smaller, cheaper, ``drafter'' model, has become a standard technique for systematically reducing the decoding time of large language models. This…

计算与语言 · 计算机科学 2025-10-03 Jameson Sandler , Ahmet Üstün , Marco Romanelli , Sara Hooker , Ferdinando Fioretto

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

In this paper, we analyze the uniqueness of the sparse time frequency decomposition and investigate the efficiency of the nonlinear matching pursuit method. Under the assumption of scale separation, we show that the sparse time frequency…

信息论 · 计算机科学 2015-05-08 Chunguang Liu , Thomas Y. Hou , Zuoqiang Shi

An efficient despeckling method using a quantum-inspired adaptive threshold function is presented for reducing noise of ultrasound images. In the first step, the ultrasound image is decorrelated by an spectrum equalization procedure due to…

计算机视觉与模式识别 · 计算机科学 2018-07-10 Hamid Reza Shahdoosti

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

The characterization of multicomponent signals with a particular emphasis on musical and communication signals is one of the problems studied in the dissertation. In order to provide an efficient analysis of the multicomponent signals, the…

信号处理 · 电气工程与系统科学 2019-03-01 Andjela Draganic

Departing from traditional communication theory where decoding algorithms are assumed to perform without error, a system where noise perturbs both computational devices and communication channels is considered here. This paper studies…

信息论 · 计算机科学 2010-05-31 Lav R. Varshney

A divide and conquer strategy for enhancement of noisy speeches in adverse environments involving lower levels of SNR is presented in this paper, where the total system of speech enhancement is divided into two separate steps. The first…

音频与语音处理 · 电气工程与系统科学 2018-02-09 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

Autoregressive (AR) modeling is invaluable in signal processing, in particular in speech and audio fields. Attempts in the literature can be found that regularize or constrain either the time-domain signal values or the AR coefficients,…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Ondřej Mokrý , Pavel Rajmic

Self-supervised learning (SSL) has recently shown remarkable results in closing the gap between supervised and unsupervised learning. The idea is to learn robust features that are invariant to distortions of the input data. Despite its…

声音 · 计算机科学 2023-03-08 Bac Nguyen , Stefan Uhlich , Fabien Cardinaux

A new technique is presented for producing images from interferometric data. The method, ``smear fitting'', makes the constraints necessary for interferometric imaging double as a model, with uncertainties, of the sky brightness…

天体物理学 · 物理学 2009-11-11 Robert I. Reid

Previous DCASE challenges contributed to an increase in the performance of acoustic scene classification systems. State-of-the-art classifiers demand significant processing capabilities and memory which is challenging for…

音频与语音处理 · 电气工程与系统科学 2021-12-10 Nagashree K. S. Rao , Nils Peters

Deconvolution of the telescope Point Spread Function (PSF) is necessary for even moderate dynamic range imaging with interferometric telescopes. The process of deconvolution can be treated as a search for a model image such that the…

天体物理学 · 物理学 2009-11-10 S. Bhatnagar , T. J. Cornwell

Self-supervised representation learning approaches have grown in popularity due to the ability to train models on large amounts of unlabeled data and have demonstrated success in diverse fields such as natural language processing, computer…

机器学习 · 计算机科学 2023-02-06 John Harvill , Jarred Barber , Arun Nair , Ramin Pishehvar

Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in vision-language representation learning, powering diverse downstream tasks and serving as the default vision backbone in multimodal large language models (MLLMs).…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Chuan Qin , Constantin Venhoff , Sonia Joseph , Fanyi Xiao , Stefan Scherer

Channel uncertainty and co-channel interference are two major challenges in the design of wireless systems such as future generation cellular networks. This paper studies receiver design for a wireless channel model with both time-varying…

信息论 · 计算机科学 2009-10-15 Yan Zhu , Dongning Guo , Michael L. Honig

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-channel SE, such as…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Mohamed Elminshawi , Wolfgang Mack , Srikanth Raj Chetupalli , Soumitro Chakrabarty , Emanuël A. P. Habets

Clipping is a common nonlinear distortion that occurs whenever the input or output of an audio system exceeds the supported range. This phenomenon undermines not only the perception of speech quality but also downstream processes utilizing…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Jayeon Yi , Junghyun Koo , Kyogu Lee