English
Related papers

Related papers: Machine Learning Framework for Audio-Based Content…

200 papers

We propose a neural audio generative model, MDCTNet, operating in the perceptually weighted domain of an adaptive modified discrete cosine transform (MDCT). The architecture of the model captures correlations in both time and frequency…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-12 Grant Davidson , Mark Vinton , Per Ekstrand , Cong Zhou , Lars Villemoes , Lie Lu

An algorithm involving Mel-Frequency Cepstral Coefficients (MFCCs) is provided to perform signal feature extraction for the task of speaker accent recognition. Then different classifiers are compared based on the MFCC feature. For each…

Sound · Computer Science 2015-02-02 Zichen Ma , Ernest Fokoue

A music mashup combines audio elements from two or more songs to create a new work. To reduce the time and effort required to make them, researchers have developed algorithms that predict the compatibility of audio elements. Prior work has…

Sound · Computer Science 2021-03-29 Jiawen Huang , Ju-Chiang Wang , Jordan B. L. Smith , Xuchen Song , Yuxuan Wang

In music, short-term features such as pitch and tempo constitute long-term semantic features such as melody and narrative. A music genre classification (MGC) system should be able to analyze these features. In this research, we propose a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-04 Jungwoo Heo , Hyun-seo Shin , Ju-ho Kim , Chan-yeong Lim , Ha-Jin Yu

Audio Sentiment Analysis is a popular research area which extends the conventional text-based sentiment analysis to depend on the effectiveness of acoustic features extracted from speech. However, current progress on audio sentiment…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-01 Feiyang Chen , Ziqian Luo

In this study, we investigate the feasibility of utilizing state-of-the-art image perceptual metrics for evaluating audio signals by representing them as spectrograms. The encouraging outcome of the proposed approach is based on the…

Sound · Computer Science 2023-08-31 Tashi Namgyal , Alexander Hepburn , Raul Santos-Rodriguez , Valero Laparra , Jesus Malo

This paper presents a novel approach to perform sentiment analysis of news videos, based on the fusion of audio, textual and visual clues extracted from their contents. The proposed approach aims at contributing to the semiodiscoursive…

Computation and Language · Computer Science 2016-04-12 Moisés H. R. Pereira , Flávio L. C. Pádua , Adriano C. M. Pereira , Fabrício Benevenuto , Daniel H. Dalip

Emotions lie on a broad continuum and treating emotions as a discrete number of classes limits the ability of a model to capture the nuances in the continuum. The challenge is how to describe the nuances of emotions and how to enable a…

Sound · Computer Science 2022-11-16 Hira Dhamyal , Benjamin Elizalde , Soham Deshmukh , Huaming Wang , Bhiksha Raj , Rita Singh

In this work, we tackle a problem of speech emotion classification. One of the issues in the area of affective computation is that the amount of annotated data is very limited. On the other hand, the number of ways that the same emotion can…

Computation and Language · Computer Science 2018-04-02 Egor Lakomkin , Cornelius Weber , Stefan Wermter

Expectations can substantially influence perception. Predictive coding is a theory of sensory processing that aims to explain the neural mechanisms underlying the effect of expectations in sensory processing. Its main assumption is that…

Neurons and Cognition · Quantitative Biology 2022-11-24 Jasmin Stein , Katharina von Kriegstein , Alejandro Tabas

As a YouTube channel grows, each video can potentially collect enormous amounts of comments that provide direct feedback from the viewers. These comments are a major means of understanding viewer expectations and improving channel…

Information Retrieval · Computer Science 2023-06-05 Rhitabrat Pokharel , Dixit Bhatta

Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio, different people may perceive the same audio differently, resulting in caption disparities (i.e., one audio may correlate to…

Sound · Computer Science 2022-04-19 Yiming Zhang , Hong Yu , Ruoyi Du , Zhanyu Ma , Yuan Dong

Sentiment analysis (SA) using code-mixed data from social media has several applications in opinion mining ranging from customer satisfaction to social campaign analysis in multilingual societies. Advances in this area are impeded by the…

Computation and Language · Computer Science 2016-11-03 Ameya Prabhu , Aditya Joshi , Manish Shrivastava , Vasudeva Varma

Sentiment analysis using deep learning and pre-trained language models (PLMs) has gained significant traction due to their ability to capture rich contextual representations. However, existing approaches often underperform in scenarios…

Computation and Language · Computer Science 2025-11-05 Peter Atandoh , Jie Zou , Weikang Guo , Jiwei Wei , Zheng Wang

Musical audio is generally composed of three physical properties: frequency, time and magnitude. Interestingly, human auditory periphery also provides neural codes for each of these dimensions to perceive music. Inspired by these intrinsic…

Sound · Computer Science 2021-06-16 Shuai Yu , Xiaoheng Sun , Yi Yu , Wei Li

We consider the task of multimodal music mood prediction based on the audio signal and the lyrics of a track. We reproduce the implementation of traditional feature engineering based approaches and propose a new model based on deep…

Information Retrieval · Computer Science 2018-09-21 Rémi Delbouys , Romain Hennequin , Francesco Piccoli , Jimena Royo-Letelier , Manuel Moussallam

Speech emotion recognition has evolved from research to practical applications. Previous studies of emotion recognition from speech have focused on developing models on certain datasets like IEMOCAP. The lack of data in the domain of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-21 Bagus Tris Atmaja , Akira Sasou

Online education platforms have experienced explosive growth over the past decade, generating massive volumes of user-generated content in the form of reviews, ratings, and behavioral logs. These heterogeneous signals provide unprecedented…

Graphics · Computer Science 2026-04-14 Arman Bekov , Azamat Nurgali

In the last decade, video blogs (vlogs) have become an extremely popular method through which people express sentiment. The ubiquitousness of these videos has increased the importance of multimodal fusion models, which incorporate video and…

Computer Vision and Pattern Recognition · Computer Science 2018-07-04 Nathaniel Blanchard , Daniel Moreira , Aparna Bharati , Walter J. Scheirer

Sentiment analysis or opinion mining has become an open research domain after proliferation of Internet and Web 2.0 social media. People express their attitudes and opinions on social media including blogs, discussion forums, tweets, etc.…

Information Retrieval · Computer Science 2013-09-17 Anuj sharma , Shubhamoy Dey