English
Related papers

Related papers: A Hybrid Approach to Audio-to-Score Alignment

200 papers

Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework that leverages "shots" as the fundamental units of video…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Junyu Xie , Tengda Han , Max Bain , Arsha Nagrani , Eshika Khandelwal , Gül Varol , Weidi Xie , Andrew Zisserman

Modern text-to-speech synthesis pipelines typically involve multiple processing stages, each of which is designed or learnt independently from the rest. In this work, we take on the challenging task of learning to synthesise speech from…

Sound · Computer Science 2021-03-18 Jeff Donahue , Sander Dieleman , Mikołaj Bińkowski , Erich Elsen , Karen Simonyan

This study addresses the task of performing robust and reliable time-delay estimation in signals in noisy and reverberating environments. In contrast to the popular signal processing based methods, this paper proposes to transform the input…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Akshay Raina , Vipul Arora

This paper introduces a novel approach for streaming openvocabulary keyword spotting (KWS) with text-based keyword enrollment. For every input frame, the proposed method finds the optimal alignment ending at the frame using connectionist…

Sound · Computer Science 2024-09-27 Sichen Jin , Youngmoon Jung , Seungjin Lee , Jaeyoung Roh , Changwoo Han , Hoonyoung Cho

The proliferation and ubiquity of temporal data across many disciplines has sparked interest for similarity, classification and clustering methods specifically designed to handle time series data. A core issue when dealing with time series…

Machine Learning · Computer Science 2023-09-26 Iñigo Martinez

In this paper we present a new method for text-independent speaker verification that combines segmental dynamic time warping (SDTW) and the d-vector approach. The d-vectors, generated from a feed forward deep neural network trained to…

Sound · Computer Science 2018-06-27 Mohamed Adel , Mohamed Afify , Akram Gaballah

The automated creation of accurate musical notation from an expressive human performance is a fundamental task in computational musicology. To this end, we present an end-to-end deep learning approach that constructs detailed musical scores…

Sound · Computer Science 2024-10-02 Tim Beyer , Angela Dai

Over the past one hundred years, the classic teaching methodology of "see one, do one, teach one" has governed the surgical education systems worldwide. With the advent of Operation Room 2.0, recording video, kinematic and many other types…

Computer Vision and Pattern Recognition · Computer Science 2019-07-23 Hassan Ismail Fawaz , Germain Forestier , Jonathan Weber , François Petitjean , Lhassane Idoumghar , Pierre-Alain Muller

Audio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single…

Sound · Computer Science 2024-03-18 Qian Wang , Jia-Chen Gu , Zhen-Hua Ling

In the domain of music production and audio processing, the implementation of automatic pitch correction of the singing voice, also known as Auto-Tune, has significantly transformed the landscape of vocal performance. While auto-tuning…

Sound · Computer Science 2024-03-11 Mahyar Gohari , Paolo Bestagini , Sergio Benini , Nicola Adami

Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects that accurately align with the corresponding audio. However, existing methods often face temporal misalignment, where audio cues and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Kexin Li , Zongxin Yang , Yi Yang , Jun Xiao

Temporal alignment of multiple signals through time warping is crucial in many fields, such as classification within speech recognition or robot motion learning. Almost all related works are limited to data in Euclidean space. Although an…

Robotics · Computer Science 2025-07-15 Julian Richter , Christopher A. Erdös , Christian Scheurer , Jochen J. Steil , Niels Dehio

The goal of this paper is twofold. First, we introduce DALI, a large and rich multimodal dataset containing 5358 audio tracks with their time-aligned vocal melody notes and lyrics at four levels of granularity. The second goal is to explain…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Gabriel Meseguer-Brocal , Alice Cohen-Hadria , Geoffroy Peeters

Time series averaging in dynamic time warping (DTW) spaces has been successfully applied to improve pattern recognition systems. This article proposes and analyzes subgradient methods for the problem of finding a sample mean in DTW spaces.…

Computer Vision and Pattern Recognition · Computer Science 2017-01-24 David Schultz , Brijnesh Jain

Temporal data are naturally everywhere, especially in the digital era that sees the advent of big data and internet of things. One major challenge that arises during temporal data analysis and mining is the comparison of time series or…

Machine Learning · Computer Science 2017-11-15 Saeid Soheily-Khah , Pierre-François Marteau

Audio-Visual Speech Recognition (AVSR) seeks to model, and thereby exploit, the dynamic relationship between a human voice and the corresponding mouth movements. A recently proposed multimodal fusion strategy, AV Align, based on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-20 George Sterpu , Christian Saam , Naomi Harte

Automated Text Scoring (ATS) provides a cost-effective and consistent alternative to human marking. However, in order to achieve good performance, the predictive features of the system need to be manually engineered by human experts. We…

Computation and Language · Computer Science 2017-07-18 Dimitrios Alikaniotis , Helen Yannakoudakis , Marek Rei

Generating sound effects with controllable variations is a challenging task, traditionally addressed using sophisticated physical models that require in-depth knowledge of signal processing parameters and algorithms. In the era of…

Sound · Computer Science 2024-12-30 Yunyi Liu , Craig Jin

Reconstructing the 3D model of a physical object typically requires us to align the depth scans obtained from different camera poses into the same coordinate system. Solutions to this global alignment problem usually proceed in two steps.…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Xiangru Huang , Zhenxiao Liang , Xiaowei Zhou , Yao Xie , Leonidas Guibas , Qixing Huang

The recent development of Audio-based Distributional Semantic Models (ADSMs) enables the computation of audio and lexical vector representations in a joint acoustic-semantic space. In this work, these joint representations are applied to…

Information Retrieval · Computer Science 2016-12-28 Giannis Karamanolakis , Elias Iosif , Athanasia Zlatintsi , Aggelos Pikrakis , Alexandros Potamianos
‹ Prev 1 8 9 10 Next ›