English
Related papers

Related papers: Machine Learning Framework for Audio-Based Content…

200 papers

There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised temporal…

Computer Vision and Pattern Recognition · Computer Science 2018-11-13 Bruno Korbar , Du Tran , Lorenzo Torresani

Objective evaluation of audio processed with Time-Scale Modification (TSM) remains an open problem. Recently, a dataset of time-scaled audio with subjective quality labels was published and used to create an initial objective measure of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-07 Timothy Roberts , Kuldip K. Paliwal

Emotional aspects play an important part in our interaction with music. However, modelling these aspects in MIR systems have been notoriously challenging since emotion is an inherently abstract and subjective experience, thus making it…

Sound · Computer Science 2019-07-09 Shreyan Chowdhury , Andreu Vall , Verena Haunschmid , Gerhard Widmer

Music prediction tasks range from predicting tags given a song or clip of audio, predicting the name of the artist, or predicting related songs given a song, clip, artist name or tag. That is, we are interested in every semantic…

Machine Learning · Computer Science 2015-03-19 Jason Weston , Samy Bengio , Philippe Hamel

We propose a novel objective evaluation metric for synthesized audio in text-to-audio (TTA), aiming to improve the performance of TTA models. In TTA, subjective evaluation of the synthesized sound is an important, but its implementation…

Sound · Computer Science 2025-07-02 Minoru Kishi , Ryosuke Sakai , Shinnosuke Takamichi , Yusuke Kanamori , Yuki Okamoto

Modeling temporal characteristics plays a significant role in the representation learning of audio waveform. We propose Contrastive Long-form Language-Audio Pretraining (\textbf{CoLLAP}) to significantly extend the perception window for…

Sound · Computer Science 2024-10-04 Junda Wu , Warren Li , Zachary Novack , Amit Namburi , Carol Chen , Julian McAuley

A controversial issue in artificial intelligence is human emotion recognition. This paper presents a fuzzy parallel cascades (FPC) model for predicting the continuous subjective appraisal of the emotional content of music by time-varying…

Human-Computer Interaction · Computer Science 2019-10-24 Fatemeh Hasanzadeh , Mohsen Annabestani , Sahar Moghimi

In this paper, we propose and investigate the use of neural audio codec language models for the automatic generation of sample-based musical instruments based on text or reference audio prompts. Our approach extends a generative audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Shahan Nercessian , Johannes Imort , Ninon Devis , Frederik Blang

Sentiment analysis, also called opinion mining, is the field of study that analyzes people's opinions,sentiments, attitudes and emotions. Songs are important to sentiment analysis since the songs and mood are mutually dependent on each…

Computation and Language · Computer Science 2018-06-12 Gangula Rama Rohit Reddy , Radhika Mamidi

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and music, particularly…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Yujia Sun , Zeyu Zhao , Korin Richmond , Yuanchao Li

Deep learning has successfully shown excellent performance in learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal structures of…

Multimedia · Computer Science 2019-08-13 Donghuo Zeng , Yi Yu , Keizo Oyama

In ideal human computer interaction (HCI), the colloquial form of a language would be preferred by most users, since it is the form used in their day-to-day conversations. However, there is also an undeniable necessity to preserve the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 M. Nanmalar , S. Johanan Joysingh , P. Vijayalakshmi , T. Nagarajan

This paper focuses on improving the accuracy of noise audio recordings. High-quality audio recording, extraction using the mel frequency cepstral coefficients (MFCC) method produces high accuracy. While the low-quality is because of noise,…

Sound · Computer Science 2022-01-03 Roy Rudolf Huizen , Florentina Tatrin Kurniati

There is a vast amount of data generated every second due to the rapidly growing technology in the current world. This area of research attempts to determine the feelings or opinions of people on social media posts. The dataset we used was…

Computation and Language · Computer Science 2023-01-24 Keshav Kapur , Rajitha Harikrishnan

This chapter describes a number of signal-processing and statistical-modeling techniques that are commonly used to calculate likelihood ratios in human-supervised automatic approaches to forensic voice comparison. Techniques described…

In speech-related classification tasks, frequency-domain acoustic features such as logarithmic Mel-filter bank coefficients (FBANK) and cepstral-domain acoustic features such as Mel-frequency cepstral coefficients (MFCC) are often used.…

Sound · Computer Science 2022-06-20 Yikang Wang , Hiromitsu Nishizaki

Segmenting audio into homogeneous sections such as music and speech helps us understand the content of audio. It is useful as a pre-processing step to index, store, and modify audio recordings, radio broadcasts and TV programmes. Deep…

Citation sentimet analysis is one of the little studied tasks for scientometric analysis. For citation analysis, we developed eight datasets comprising citation sentences, which are manually annotated by us into three sentiment polarities…

Computation and Language · Computer Science 2020-05-12 Vishal Vyas , Kumar Ravi , Vadlamani Ravi , V. Uma , Srirangaraj Setlur , Venu Govindaraju

Speech Emotion Recognition (SER) has emerged as a critical component of the next generation human-machine interfacing technologies. In this work, we propose a new dual-level model that predicts emotions based on both MFCC features and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-24 Jianyou Wang , Michael Xue , Ryan Culhane , Enmao Diao , Jie Ding , Vahid Tarokh

Spoken content, such as online videos and podcasts, often spans multiple topics, which makes automatic topic segmentation essential for user navigation and downstream applications. However, current methods do not fully leverage acoustic…

Computation and Language · Computer Science 2026-02-09 Steffen Freisinger , Philipp Seeberger , Tobias Bocklet , Korbinian Riedhammer
‹ Prev 1 3 4 5 6 7 10 Next ›