English
Related papers

Related papers: SPICE: Self-supervised Pitch Estimation

200 papers

In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances. This is…

Computation and Language · Computer Science 2020-02-06 Alexander H. Liu , Tao Tu , Hung-yi Lee , Lin-shan Lee

With the exponential growth of video content, the need for automated video highlight detection to extract key moments or highlights from lengthy videos has become increasingly pressing. This technology has the potential to enhance user…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Zahidul Islam , Sujoy Paul , Mrigank Rochan

Machine learning has achieved impressive performance in tomographic reconstruction, but supervised training requires paired measurements and ground-truth images that are often unavailable. This has motivated self-supervised approaches,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Markus Haltmeier , Lukas Neumann , Nadja Gruber , Gyeongha Hwang

Deep-learning metrics have recently demonstrated extremely good performance to match image patches for stereo reconstruction. However, training such metrics requires large amount of labeled stereo images, which can be difficult or costly to…

Computer Vision and Pattern Recognition · Computer Science 2016-12-06 Stepan Tulyakov , Anton Ivanov , Francois Fleuret

Accurate real depth annotations are difficult to acquire, needing the use of special devices such as a LiDAR sensor. Self-supervised methods try to overcome this problem by processing video or stereo sequences, which may not always be…

Computer Vision and Pattern Recognition · Computer Science 2020-09-04 Adrian Lopez-Rodriguez , Krystian Mikolajczyk

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Scott Wisdom , Efthymios Tzinis , Hakan Erdogan , Ron J. Weiss , Kevin Wilson , John R. Hershey

Dense depth estimation from a single image is a key problem in computer vision, with exciting applications in a multitude of robotic tasks. Initially viewed as a direct regression problem, requiring annotated labels as supervision at…

Computer Vision and Pattern Recognition · Computer Science 2019-11-20 Vitor Guizilini , Jie Li , Rares Ambrus , Sudeep Pillai , Adrien Gaidon

We propose an objective measurement method for pitch extractors' responses to frequency-modulated signals. The method simultaneously measures the linear and the non-linear time-invariant responses and random and time-varying responses. It…

A crucial aspect for the successful deployment of audio-based models "in-the-wild" is the robustness to the transformations introduced by heterogeneous acquisition conditions. In this work, we propose a method to perform one-shot microphone…

Sound · Computer Science 2020-10-20 Zalán Borsos , Yunpeng Li , Beat Gfeller , Marco Tagliasacchi

We present a method for audio denoising that combines processing done in both the time domain and the time-frequency domain. Given a noisy audio clip, the method trains a deep neural network to fit this signal. Since the fitting is only…

Sound · Computer Science 2020-06-11 Michael Michelashvili , Lior Wolf

Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not…

Sound · Computer Science 2025-03-04 Sripathi Sridhar , Mark Cartwright

Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-13 Benjamin van Niekerk , Marc-André Carbonneau , Herman Kamper

We apply transfer learning to the task of phoneme segmentation and demonstrate the utility of representations learned in self-supervised pre-training for the task. Our model extends transformer-style encoders with strategically placed…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-04 Luke Strgar , David Harwath

We introduce two unsupervised source separation methods, which involve self-supervised training from single-channel two-source speech mixtures. Our first method, mixture permutation invariant training (MixPIT), enables learning a neural…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-11 Ertuğ Karamatlı , Serap Kırbız

Accurate facial estimation is crucial for realistic digital human animation, and ARKit blendshape coefficients offer an interpretable representation by mapping facial motions to semantic animation controls. However, learning high-quality…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Zejian Kang , Xuanyang Xu , Wentao Yang , Kai Zheng , Yuanchen Fei , Hongyuan Zou , Hui Shan , Shuo Yang , Xiangru Huang

Distance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-27 Michael Neri , Archontis Politis , Daniel Krause , Marco Carli , Tuomas Virtanen

Sound source tracking is commonly performed using classical array-processing algorithms, while machine-learning approaches typically rely on precise source position labels that are expensive or impractical to obtain. This paper introduces a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-12 Luan Vinícius Fiorio , Ivana Nikoloska , Bruno Defraene , Alex Young , Johan David , Ronald M. Aarts

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sudha Krishnamurthy

This paper addresses the problem of key phrase extraction from sentences. Existing state-of-the-art supervised methods require large amounts of annotated data to achieve good performance and generalization. Collecting labeled data is,…

Computation and Language · Computer Science 2019-04-09 Jue Wang , Ke Chen , Lidan Shou , Sai Wu , Sharad Mehrotra

Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the application of CLIP to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Sooyoung Park , Arda Senocak , Joon Son Chung