中文
相关论文

相关论文: Real Time Vowel Tremolo Detection Using Low Level …

200 篇论文

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras and microphones are…

音频与语音处理 · 电气工程与系统科学 2024-06-05 Davide Berghi , Philip J. B. Jackson

Music is a form of expression that often requires interaction between players. If one wishes to interact in such a musical way with a computer, it is necessary for the machine to be able to interpret the input given by the human to find its…

声音 · 计算机科学 2022-09-01 Filippo Carnovalini , Antonio Rodà

We target the problem of developing new low-complexity networks for the sound event detection task. Our goal is to meticulously analyze the performance-complexity trade-off, aiming to be competitive with the large state-of-the-art models,…

声音 · 计算机科学 2025-06-13 Tobias Morocutti , Florian Schmid , Jonathan Greif , Francesco Foscarin , Gerhard Widmer

The automatic identification and analysis of pronunciation errors, known as Mispronunciation Detection and Diagnosis (MDD) plays a crucial role in Computer Aided Pronunciation Learning (CAPL) tools such as Second-Language (L2) learning or…

音频与语音处理 · 电气工程与系统科学 2023-11-14 Mostafa Shahin , Julien Epps , Beena Ahmed

Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by…

声音 · 计算机科学 2023-10-18 Fernando López , Jordi Luque , Carlos Segura , Pablo Gómez

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, and temporal…

In the field of music information retrieval, the task of simultaneously identifying the presence or absence of multiple musical instruments in a polyphonic recording remains a hard problem. Previous works have seen some success in improving…

音频与语音处理 · 电气工程与系统科学 2020-06-23 Karn Watcharasupat , Siddharth Gururani , Alexander Lerch

We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phrase. We first demonstrate that by including more audio…

音频与语音处理 · 电气工程与系统科学 2021-03-03 Siddharth Sigtia , John Bridle , Hywel Richards , Pascal Clark , Erik Marchi , Vineet Garg

Connecting large libraries of digitized audio recordings to their corresponding sheet music images has long been a motivation for researchers to develop new cross-modal retrieval systems. In recent years, retrieval systems based on…

信息检索 · 计算机科学 2019-06-27 Stefan Balke , Matthias Dorfer , Luis Carvalho , Andreas Arzt , Gerhard Widmer

Outbound AI calling systems must distinguish voicemail greetings from live human answers in real time to avoid wasted agent interactions and dropped calls. We present a lightweight approach that extracts 15 temporal features from the speech…

声音 · 计算机科学 2026-04-14 Kumar Saurav

Audio-Visual Event Localization (AVEL) is the task of temporally localizing and classifying \emph{audio-visual events}, i.e., events simultaneously visible and audible in a video. In this paper, we solve AVEL in a weakly-supervised setting,…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Kalyan Ramakrishnan

We introduce a distinctive real-time, causal, neural network-based active speaker detection system optimized for low-power edge computing. This system drives a virtual cinematography module and is deployed on a commercial device. The system…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Ilya Gurvich , Ido Leichter , Dharmendar Reddy Palle , Yossi Asher , Alon Vinnikov , Igor Abramovski , Vishak Gopal , Ross Cutler , Eyal Krupka

In most current approaches of speech processing, information is extracted from the magnitude spectrum. However recent perceptual studies have underlined the importance of the phase component. The goal of this paper is to investigate the…

声音 · 计算机科学 2020-01-03 Thomas Drugman , Thomas Dubuisson , Thierry Dutoit

In the domain of music and sound processing, pitch extraction plays a pivotal role. Our research presents a specialized convolutional neural network designed for pitch extraction, particularly from the human singing voice in acapella…

声音 · 计算机科学 2023-12-19 Jeremy Cochoy

Audio Event Detection is an important task for content analysis of multimedia data. Most of the current works on detection of audio events is driven through supervised learning approaches. We propose a weakly supervised learning framework…

声音 · 计算机科学 2016-06-14 Anurag Kumar , Bhiksha Raj

Lack of large-scale note-level labeled data is the major obstacle to singing transcription from polyphonic music. We address the issue by using pseudo labels from vocal pitch estimation models given unlabeled data. The proposed method first…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Sangeun Kum , Jongpil Lee , Keunhyoung Luke Kim , Taehyoung Kim , Juhan Nam

This article develops a general detection theory for speech analysis based on time-varying autoregressive models, which themselves generalize the classical linear predictive speech analysis framework. This theory leads to a computationally…

应用统计 · 统计学 2011-08-25 Daniel Rudoy , Thomas F. Quatieri , Patrick J. Wolfe

Articulatory features can provide interpretable and flexible controls for the synthesis of human vocalizations by allowing the user to directly modify parameters like vocal strain or lip position. To make this manipulation through…

声音 · 计算机科学 2023-07-11 David Südholt , Mateo Cámara , Zhiyuan Xu , Joshua D. Reiss

Vocal education in the music field is difficult to quantify due to the individual differences in singers' voices and the different quantitative criteria of singing techniques. Deep learning has great potential to be applied in music…

音频与语音处理 · 电气工程与系统科学 2024-11-01 Zhenyi Hou , Xu Zhao , Kejie Ye , Xinyu Sheng , Shanggerile Jiang , Jiajing Xia , Yitao Zhang , Chenxi Ban , Daijun Luo , Jiaxing Chen , Yan Zou , Yuchao Feng , Guangyu Fan , Xin Yuan

There are different algorithms for vocal fold pathology diagnosis. These algorithms usually have three stages which are Feature Extraction, Feature Reduction and Classification. While the third stage implies a choice of a variety of machine…

机器学习 · 计算机科学 2013-02-08 Vahid Majidnezhad , Igor Kheidorov