English
Related papers

Related papers: Real Time Vowel Tremolo Detection Using Low Level …

200 papers

We present a model for capturing musical features and creating novel sequences of music, called the Convolutional Variational Recurrent Neural Network. To generate sequential data, the model uses an encoder-decoder architecture with latent…

Sound · Computer Science 2018-10-09 Eunjeong Stella Koh , Shlomo Dubnov , Dustin Wright

LSTM-based speaker verification usually uses a fixed-length local segment randomly truncated from an utterance to learn the utterance-level speaker embedding, while using the average embedding of all segments of a test utterance to verify…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-05 Bin Liu , Shuai Nie , Yaping Zhang , Shan Liang , Wenju Liu

Despite recent advancements in speech generation with text prompt providing control over speech style, voice attributes in synthesized speech remain elusive and challenging to control. This paper introduces a novel task: voice attribute…

Sound · Computer Science 2024-12-03 Zhengyan Sheng , Yang Ai , Li-Juan Liu , Jia Pan , Zhen-Hua Ling

While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work has indicated that…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 Nick Mehlman , Thomas Thebaud , Dani Byrd , Shri Narayanan

Music is a mysterious language that conveys feeling and thoughts via different tones and timbre. For better understanding of timbre in music, we chose music data of 6 representative instruments, analysed their timbre features and classified…

Sound · Computer Science 2022-07-15 Zishuo Zhao , Haoyun Wang

This paper focuses on the weakly-supervised audio-visual video parsing task, which aims to recognize all events belonging to each modality and localize their temporal boundaries. This task is challenging because only overall labels…

Computer Vision and Pattern Recognition · Computer Science 2022-08-02 Haoyue Cheng , Zhaoyang Liu , Hang Zhou , Chen Qian , Wayne Wu , Limin Wang

This paper addresses the problem of automatic detection of voice pathologies directly from the speech signal. For this, we investigate the use of the glottal source estimation as a means to detect voice disorders. Three sets of features are…

Sound · Computer Science 2020-01-06 Thomas Drugman , Thomas Dubuisson , Thierry Dutoit

Sound event detection (SED) is the task of identifying sound events along with their onset and offset times. A recent, convolutional neural networks based SED method, proposed the usage of depthwise separable (DWS) and time-dilated…

Sound · Computer Science 2020-07-13 Konstantinos Drossos , Stylianos I. Mimilakis , Tuomas Virtanen

Continuously learning a variety of audio-video semantics over time is crucial for audio-related reasoning tasks in our ever-evolving world. However, this is a nontrivial problem and poses two critical challenges: sparse spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jaewoo Lee , Jaehong Yoon , Wonjae Kim , Yunji Kim , Sung Ju Hwang

Dementia, a progressive neurodegenerative disorder, affects memory, reasoning, and daily functioning, creating challenges for individuals and healthcare systems. Early detection is crucial for timely interventions that may slow disease…

Neurons and Cognition · Quantitative Biology 2025-03-04 Sahar Sinene Mehdoui , Abdelhamid Bouzid , Daniel Sierra-Sosa , Adel Elmaghraby

A wavelet-based changepoint method is proposed that determines when the variability of the noise in a sequence of functional profiles goes out-of-control from a known, fixed value. The functional portion of the profiles are allowed to come…

Methodology · Statistics 2015-08-20 Vladimir J. Geneus , Eric Chicken , Jordan Cuevas , Joseph J. Pignatiello

Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consistency due to limited natural language understanding and data…

While most music generation models use textual or parametric conditioning (e.g. tempo, harmony, musical genre), we propose to condition a language model based music generation system with audio input. Our exploration involves two distinct…

Sound · Computer Science 2024-07-31 Simon Rouard , Yossi Adi , Jade Copet , Axel Roebel , Alexandre Défossez

A main challenge in applying deep learning to music processing is the availability of training data. One potential solution is Multi-task Learning, in which the model also learns to solve related auxiliary tasks on additional datasets to…

Sound · Computer Science 2018-04-06 Daniel Stoller , Sebastian Ewert , Simon Dixon

Current models for audio--sheet music retrieval via multimodal embedding space learning use convolutional neural networks with a fixed-size window for the input audio. Depending on the tempo of a query performance, this window captures more…

Sound · Computer Science 2018-09-18 Matthias Dorfer , Jan Hajič , Gerhard Widmer

This article presents an interactive system for stage acoustics experimentation including considerations for hearing one's own and others' instruments. The quality of real-time auralization systems for psychophysical experiments on music…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-30 Ernesto Accolti , Lukas Aspöck , Manuj Yadav , Michael Vorländer

Temporal action detection aims to locate and classify actions in untrimmed videos. While recent works focus on designing powerful feature processors for pre-trained representations, they often overlook the inherent noise and redundancy…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Xinnan Zhu , Yicheng Zhu , Tixin Chen , Wentao Wu , Yuanjie Dang

Autoregressive large vision--language models (LVLMs) interface video and language by projecting video features into the LLM's embedding space as continuous visual token embeddings. However, it remains unclear where temporal evidence is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yiming Zhang , Zhuokai Zhao , Chengzhang Yu , Kun Wang , Zhendong Chu , Qiankun Li , Zihan Chen , Yang Liu , Zenghui Ding , Yining Sun , Qingsong Wen

Most studies on speaker verification systems focus on long-duration utterances, which are composed of sufficient phonetic information. However, the performances of these systems are known to degrade when short-duration utterances are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-05 Seung-bin Kim , Jee-weon Jung , Hye-jin Shim , Ju-ho Kim , Ha-Jin Yu

Multi-modal contrastive learning techniques in the audio-text domain have quickly become a highly active area of research. Most works are evaluated with standard audio retrieval and classification benchmarks assuming that (i) these models…

Sound · Computer Science 2023-03-21 Ho-Hsiang Wu , Oriol Nieto , Juan Pablo Bello , Justin Salamon