中文
相关论文

相关论文: Transformer-based Sequence Labeling for Audio Clas…

200 篇论文

Speaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we…

声音 · 计算机科学 2022-06-28 Zhiyun Fan , Linhao Dong , Meng Cai , Zejun Ma , Bo Xu

Music genre classification is an area that utilizes machine learning models and techniques for the processing of audio signals, in which applications range from content recommendation systems to music recommendation systems. In this…

声音 · 计算机科学 2024-05-27 Keoikantse Mogonediwa

In this paper, we show that ImageNet-Pretrained standard deep CNN models can be used as strong baseline networks for audio classification. Even though there is a significant difference between audio Spectrogram and standard ImageNet image…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Kamalesh Palanisamy , Dipika Singhania , Angela Yao

Acoustic Scene Classification (ASC) is one of the core research problems in the field of Computational Sound Scene Analysis. In this work, we present SubSpectralNet, a novel model which captures discriminative features by incorporating…

声音 · 计算机科学 2019-02-26 Sai Samarth R Phaye , Emmanouil Benetos , Ye Wang

In this paper, we describe our contribution to Task 2 of the DCASE 2018 Audio Challenge. While it has become ubiquitous to utilize an ensemble of machine learning methods for classification tasks to obtain better predictive performance, the…

声音 · 计算机科学 2018-11-28 Marcel Lederle , Benjamin Wilhelm

End-to-end neural network based approaches to audio modelling are generally outperformed by models trained on high-level data representations. In this paper we present preliminary work that shows the feasibility of training the first layers…

声音 · 计算机科学 2017-12-04 Tycho Max Sylvester Tax , Jose Luis Diez Antich , Hendrik Purwins , Lars Maaløe

In this paper, we present a deep learning framework applied for Acoustic Scene Classification (ASC), the task of classifying scene contexts from environmental input sounds. An ASC system generally comprises of two main steps, referred to as…

声音 · 计算机科学 2020-05-27 Dat Ngo , Hao Hoang , Anh Nguyen , Tien Ly , Lam Pham

Deep Learning (DL) algorithms have shown impressive performance in diverse domains. Among them, audio has attracted many researchers over the last couple of decades due to some interesting patterns--particularly in classification of audio…

声音 · 计算机科学 2022-06-16 Muhammad Turab , Teerath Kumar , Malika Bendechache , Takfarinas Saber

This study investigates the extent to which Mel-Frequency Cepstral Coefficients (MFCCs) capture first language (L1) transfer in extended second language (L2) English speech. Speech samples from Mandarin and American English L1 speakers were…

音频与语音处理 · 电气工程与系统科学 2025-04-21 Peyman Jahanbin

The development of models for learning music similarity and feature extraction from audio media files is an increasingly important task for the entertainment industry. This work proposes a novel music classification model based on metric…

声音 · 计算机科学 2019-09-19 Angelo C. Mendes da Silva , Mauricio A. Nunes , Raul Fonseca Neto

Environmental Sound Classification (ESC) is a rapidly evolving field that recently demonstrated the advantages of application of visual domain techniques to the audio-related tasks. Previous studies indicate that the domain-specific…

声音 · 计算机科学 2021-04-26 Andrey Guzhov , Federico Raue , Jörn Hees , Andreas Dengel

Researches have shown accent classification can be improved by integrating semantic information into pure acoustic approach. In this work, we combine phonetic knowledge, such as vowels, with enhanced acoustic features to build an improved…

声音 · 计算机科学 2016-02-25 Zhenhao Ge

We present an end-to-end method for transforming audio from one style to another. For the case of speech, by conditioning on speaker identities, we can train a single model to transform words spoken by multiple people into multiple target…

声音 · 计算机科学 2018-06-08 Albert Haque , Michelle Guo , Prateek Verma

In this work we propose approaches to effectively transfer knowledge from weakly labeled web audio data. We first describe a convolutional neural network (CNN) based framework for sound event detection and classification using weakly…

声音 · 计算机科学 2018-09-10 Anurag Kumar , Maksim Khadkevich , Christian Fugen

Consumer-grade music recordings such as those captured by mobile devices typically contain distortions in the form of background noise, reverb, and microphone-induced EQ. This paper presents a deep learning approach to enhance low-quality…

声音 · 计算机科学 2022-04-29 Nikhil Kandpal , Oriol Nieto , Zeyu Jin

This paper proposes a speech emotion recognition method based on speech features and speech transcriptions (text). Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCC) help retain emotion-related low-level…

音频与语音处理 · 电气工程与系统科学 2019-06-14 Suraj Tripathi , Abhay Kumar , Abhiram Ramesh , Chirag Singh , Promod Yenigalla

Audio-based equipment condition monitoring suffers from a lack of standardized methodologies for algorithm selection, hindering reproducible research. This paper addresses this gap by introducing a comprehensive framework for the systematic…

机器学习 · 计算机科学 2026-03-20 Srijesh Pillai , Yodhin Agarwal , Zaheeruddin Ahmed

We introduce in this work an efficient approach for audio scene classification using deep recurrent neural networks. An audio scene is firstly transformed into a sequence of high-level label tree embedding feature vectors. The vector…

声音 · 计算机科学 2017-06-06 Huy Phan , Philipp Koch , Fabrice Katzberg , Marco Maass , Radoslaw Mazur , Alfred Mertins

Spectrograms have been widely used in Convolutional Neural Networks based schemes for acoustic scene classification, such as the STFT spectrogram and the MFCC spectrogram, etc. They have different time-frequency characteristics,…

计算机视觉与模式识别 · 计算机科学 2018-09-06 Weiping Zheng , Zhenyao Mo , Xiaotao Xing , Gansen Zhao

Transformers have set new benchmarks in audio processing tasks, leveraging self-attention mechanisms to capture complex patterns and dependencies within audio data. However, their focus on pairwise interactions limits their ability to…

声音 · 计算机科学 2025-01-30 Shubhr Singh , Emmanouil Benetos , Huy Phan , Dan Stowell