中文
相关论文

相关论文: Learning Temporal Resolution in Spectrogram for Au…

200 篇论文

The goal of this thesis was to develop a procedure to optimize the mirror suppression, which reduces the dynamic range of the Fast Fourier Transform spectrometers (FFTS), and implement this procedure in the Field Programmable Gate Array…

天体物理仪器与方法 · 物理学 2022-08-23 Gerrit Grutzeck

Most of existing audio fingerprinting systems have limitations to be used for high-specific audio retrieval at scale. In this work, we generate a low-dimensional representation from a short unit segment of audio, and couple this fingerprint…

声音 · 计算机科学 2021-02-11 Sungkyun Chang , Donmoon Lee , Jeongsoo Park , Hyungui Lim , Kyogu Lee , Karam Ko , Yoonchang Han

We propose a method for automatic local time-adaptation of the spectrogram of audio signals: it is based on the decomposition of a signal within a Gabor multi-frame through the STFT operator. The sparsity of the analysis in every individual…

声音 · 计算机科学 2011-09-29 M. Liuni , A. Röbel , M. Romito , X. Rodet

Audio deepfake detection is increasingly important as synthetic speech becomes more realistic and accessible. Recent methods, including those using graph neural networks (GNNs) to model frequency and temporal dependencies, show strong…

声音 · 计算机科学 2026-01-13 Falih Gozi Febrinanto , Kristen Moore , Chandra Thapa , Jiangang Ma , Vidya Saikrishna

This study investigated the waveform representation for audio signal classification. Recently, many studies on audio waveform classification such as acoustic event detection and music genre classification have been published. Most studies…

音频与语音处理 · 电气工程与系统科学 2019-09-19 Masaki Okawa , Takuya Saito , Naoki Sawada , Hiromitsu Nishizaki

We introduce a novel, training-free method for sampling differentiable representations (diffreps) using pretrained diffusion models. Rather than merely mode-seeking, our method achieves sampling by "pulling back" the dynamics of the…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Yash Savani , Marc Finzi , J. Zico Kolter

The widespread application of audio and video communication technology make the compressed audio data flowing over the Internet, and make it become an important carrier for covert communication. There are many steganographic schemes emerged…

多媒体 · 计算机科学 2019-02-27 Yanzhen Ren , Dengkai Liu , Qiaochu Xiong , Jianming Fu , Lina Wang

This paper introduces a novel technique for reconstructing the phase of modified spectrograms of audio signals. From the analysis of mixtures of sinusoids we obtain relationships between phases of successive time frames in the…

声音 · 计算机科学 2016-05-25 Paul Magron , Roland Badeau , Bertrand David

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

音频与语音处理 · 电气工程与系统科学 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson

We seek to develop simultaneous segmentation and classification of notes from audio recordings in presence of outliers. The selected architecture for modeling time series is hierarchical linear dynamical system (HLDS). We propose a novel…

声音 · 计算机科学 2022-03-01 Leila Kalantari , Jose Principe , Kathryn E. Sieving

Representation learning frameworks in unlabeled time series have been proposed for medical signal processing. Despite the numerous excellent progresses have been made in previous works, we observe the representation extracted for the time…

信号处理 · 电气工程与系统科学 2024-01-12 Luyuan Xie , Cong Li , Xin Zhang , Shengfang Zhai , Yuejian Fang , Qingni Shen , Zhonghai Wu

Computer musicians refer to mesostructures as the intermediate levels of articulation between the microstructure of waveshapes and the macrostructure of musical forms. Examples of mesostructures include melody, arpeggios, syncopation,…

声音 · 计算机科学 2023-01-25 Cyrus Vahidi , Han Han , Changhong Wang , Mathieu Lagrange , György Fazekas , Vincent Lostanlen

Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform…

声音 · 计算机科学 2024-05-31 Adam Sorrenti

While deep learning has reduced the prevalence of manual feature extraction, transformation of data via feature engineering remains essential for improving model performance, particularly for underwater acoustic signals. The methods by…

Automated respiratory audio analysis promises scalable, non-invasive disease screening, yet progress is limited by scarce labeled data and costly expert annotation. Zero-shot inference eliminates task-specific supervision, but existing…

声音 · 计算机科学 2026-04-15 Tsai-Ning Wang , Herman Teun den Dekker , Lin-Lin Chen , Neil Zeghidour , Aaqib Saeed

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

声音 · 计算机科学 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

Environmental sound classification systems often do not perform robustly across different sound classification tasks and audio signals of varying temporal structures. We introduce a multi-stream convolutional neural network with temporal…

声音 · 计算机科学 2019-01-28 Xinyu Li , Venkata Chebiyyam , Katrin Kirchhoff

In this report we describe an ongoing line of research for solving single-channel source separation problems. Many monaural signal decomposition techniques proposed in the literature operate on a feature space consisting of a time-frequency…

声音 · 计算机科学 2015-04-29 Pablo Sprechmann , Joan Bruna , Yann LeCun

Photoacoustic (PA) image reconstruction involves acoustic inversion that necessitates the specification of the speed of sound (SoS) within the medium of propagation. Due to the lack of information on the spatial distribution of the SoS…

图像与视频处理 · 电气工程与系统科学 2024-06-05 Mengjie Shi , Tom Vercauteren , Wenfeng Xia

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given region for spatial audio recorded by a microphone array. The…

音频与语音处理 · 电气工程与系统科学 2025-09-12 Jinzheng Zhao , Yong Xu , Haohe Liu , Davide Berghi , Xinyuan Qian , Qiuqiang Kong , Junqi Zhao , Mark D. Plumbley , Wenwu Wang