中文
相关论文

相关论文: PHALAR: Phasors for Learned Musical Audio Represen…

200 篇论文

Models of sensory processing and learning in the cortex need to efficiently assign credit to synapses in all areas. In deep learning, a known solution is error backpropagation, which however requires biologically implausible weight…

神经元与认知 · 定量生物学 2024-02-05 Kevin Max , Laura Kriener , Garibaldi Pineda García , Thomas Nowotny , Ismael Jaras , Walter Senn , Mihai A. Petrovici

Selecting appropriate inductive biases is an essential step in the design of machine learning models, especially when working with audio, where even short clips may contain millions of samples. To this end, we propose the combolutional…

声音 · 计算机科学 2025-08-06 Cameron Churchwell , Minje Kim , Paris Smaragdis

In this paper, we tackle the task of musical stem retrieval. Given a musical mix, it consists in retrieving a stem that would fit with it, i.e., that would sound pleasant if played together. To do so, we introduce a new method based on…

声音 · 计算机科学 2025-02-25 Alain Riou , Antonin Gagneré , Gaëtan Hadjeres , Stefan Lattner , Geoffroy Peeters

We describe a new algorithm to solve a particular phase retrieval problem, that has wide applications in audio processing: the reconstruction of a function from its scalogram, that is from the modulus of its wavelet transform. It is a…

最优化与控制 · 数学 2017-04-11 Irène Waldspurger

Lens flare removal is challenging due to the large spatial extent of flare artifacts and their entanglement with scene structures, while existing methods heavily rely on large-scale paired data. We propose a semi-supervised flare removal…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Xiyu Zhu , Wei Wang , Kui Jiang , Zhengguo Li

Phase retrieval approaches based on DL provide a framework to obtain phase information from an intensity hologram or diffraction pattern in a robust manner and in real time. However, current DL architectures applied to the phase problem…

图像与视频处理 · 电气工程与系统科学 2021-07-07 Yuhe Zhang , Mike Andreas Noack , Patrik Vagovic , Kamel Fezzaa , Francisco Garcia-Moreno , Tobias Ritschel , Pablo Villanueva-Perez

In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and text encoder together. Secondly, we show that using this…

声音 · 计算机科学 2022-06-30 Andrew Koh , Eng Siong Chng

We propose Fast Language-Audio Pre-training (FLAP), a self-supervised approach that efficiently and effectively learns aligned audio and language representations through masking, contrastive learning and reconstruction. For efficiency, FLAP…

声音 · 计算机科学 2023-11-06 Ching-Feng Yeh , Po-Yao Huang , Vasu Sharma , Shang-Wen Li , Gargi Gosh

Learning rich visual representations using contrastive self-supervised learning has been extremely successful. However, it is still a major question whether we could use a similar approach to learn superior auditory representations. In this…

声音 · 计算机科学 2020-10-20 Haider Al-Tahan , Yalda Mohsenzadeh

In many areas of imaging science, it is difficult to measure the phase of linear measurements. As such, one often wishes to reconstruct a signal from intensity measurements, that is, perform phase retrieval. In several applications the…

信息论 · 计算机科学 2015-06-16 Afonso S. Bandeira , Dustin G. Mixon

Compared to traditional electrodynamic loudspeakers, the parametric array loudspeaker (PAL) offers exceptional directivity for audio applications but suffers from significant nonlinear distortions due to its inherent intricate demodulation…

音频与语音处理 · 电气工程与系统科学 2025-07-11 Mengtong Li , Tao Zhuang , Kai Chen , Jia-Xin Zhong , Jing Lu

Vision-based fall analysis has advanced rapidly, but a key bottleneck remains: visually similarmotions can correspond to very different physical outcomes because small differences in contactmechanics and protective responses are hard to…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Xianqi Zhang

This paper proposes a novel algorithm for image phase retrieval, i.e., for recovering complex-valued images from the amplitudes of noisy linear combinations (often the Fourier transform) of the sought complex images. The algorithm is…

信号处理 · 电气工程与系统科学 2018-10-19 Joshin P. Krishnan , José M. Bioucas-Dias , Vladimir Katkovnik

In text-audio retrieval (TAR) tasks, due to the heterogeneity of contents between text and audio, the semantic information contained in the text is only similar to certain frames within the audio. Yet, existing works aggregate the entire…

声音 · 计算机科学 2023-03-31 Yifei Xin , Dongchao Yang , Yuexian Zou

In a variety of fields, in particular those involving imaging and optics, we often measure signals whose phase is missing or has been irremediably distorted. Phase retrieval attempts to recover the phase information of a signal from the…

信息论 · 计算机科学 2019-10-02 Gilles Baechler , Miranda Kreković , Juri Ranieri , Amina Chebira , Yue M. Lu , Martin Vetterli

Tremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization - a hallmark of foundation models in other computer vision tasks -…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Bowen Wen , Matthew Trepte , Joseph Aribido , Jan Kautz , Orazio Gallo , Stan Birchfield

Several automatic approaches for objective music performance assessment (MPA) have been proposed in the past, however, existing systems are not yet capable of reliably predicting ratings with the same accuracy as professional judges. This…

声音 · 计算机科学 2021-08-16 Pavan Seshadri , Alexander Lerch

Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less consideration of cross-lingual alignment during training. Although standardized contrastive…

计算与语言 · 计算机科学 2026-04-15 Seungyoon Lee , Minhyuk Kim , Seongtae Hong , Youngjoon Jang , Dongsuk Oh , Heuiseok Lim

While log-amplitude mel-spectrogram has widely been used as the feature representation for processing speech based on deep learning, the effectiveness of another aspect of speech spectrum, i.e., phase information, was shown recently for…

声音 · 计算机科学 2022-05-02 Shunsuke Hidaka , Kohei Wakamiya , Tokihiko Kaburagi

Sparse modeling is one of the efficient techniques for imaging that allows recovering lost information. In this paper, we present a novel iterative phase-retrieval algorithm using a sparse representation of the object amplitude and phase.…

计算机视觉与模式识别 · 计算机科学 2011-08-17 Artem Migukin , Vladimir Katkovnik , Jaakko Astola
‹ 上一页 1 2 3 10 下一页 ›