中文
相关论文

相关论文: Boosting Local Spectro-Temporal Features for Speec…

200 篇论文

Acoustic scenes are rich and redundant in their content. In this work, we present a spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network to learn from patterns that are discriminative while…

声音 · 计算机科学 2019-07-01 Huy Phan , Oliver Y. Chén , Lam Pham , Philipp Koch , Maarten De Vos , Ian McLoughlin , Alfred Mertins

In image based feature descriptor design, local information from image patches are extracted using iterative scanning operations which cause high computational costs. In order to avoid such scanning operations, we present matrix…

计算机视觉与模式识别 · 计算机科学 2021-02-09 Zainab Alhakeem , Se-In Jang

This paper introduces and motivates the use of hybrid robust feature extraction technique for spoken language identification (LID) system. The speech recognizers use a parametric form of a signal to get the most important distinguishable…

声音 · 计算机科学 2010-03-31 Pawan Kumar , Astik Biswas , A . N. Mishra , Mahesh Chandra

The aim of this paper is to discuss the use of Haar scattering networks, which is a very simple architecture that naturally supports a large number of stacked layers, yet with very few parameters, in a relatively broad set of pattern…

信号处理 · 电气工程与系统科学 2018-11-30 Fernando Fernandes Neto , Alemayehu Admasu Solomon , Rodrigo de Losso , Claudio Garcia , Pedro Delano Cavalcanti

There is an abundant literature on face detection due to its important role in many vision applications. Since Viola and Jones proposed the first real-time AdaBoost based face detector, Haar-like features have been adopted as the method of…

计算机视觉与模式识别 · 计算机科学 2010-09-30 Sakrapee Paisitkriangkrai , Chunhua Shen , Jian Zhang

Sound scene geotagging is a new topic of research which has evolved from acoustic scene classification. It is motivated by the idea of audio surveillance. Not content with only describing a scene in a recording, a machine which can locate…

音频与语音处理 · 电气工程与系统科学 2021-10-12 Helen L. Bear , Veronica Morfi , Emmanouil Benetos

This article develops a general detection theory for speech analysis based on time-varying autoregressive models, which themselves generalize the classical linear predictive speech analysis framework. This theory leads to a computationally…

应用统计 · 统计学 2011-08-25 Daniel Rudoy , Thomas F. Quatieri , Patrick J. Wolfe

Phoneme boundary detection plays an essential first step for a variety of speech processing applications such as speaker diarization, speech science, keyword spotting, etc. In this work, we propose a neural architecture coupled with a…

音频与语音处理 · 电气工程与系统科学 2020-02-18 Felix Kreuk , Yaniv Sheena , Joseph Keshet , Yossi Adi

Detecting auditory attention based on brain signals enables many everyday applications, and serves as part of the solution to the cocktail party effect in speech processing. Several studies leverage the correlation between brain signals and…

人机交互 · 计算机科学 2024-10-28 Siqi Cai , Pengcheng Sun , Tanja Schultz , Haizhou Li

This work introduces a feature extracted from stereophonic/binaural audio signals aiming to represent a measure of perceived quality degradation in processed spatial auditory scenes. The feature extraction technique is based on a simplified…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Pablo M. Delgado , Jürgen Herre

The purpose of speech emotion recognition system is to classify speakers utterances into different emotional states such as disgust, boredom, sadness, neutral and happiness. Speech features that are commonly used in speech emotion…

计算与语言 · 计算机科学 2014-06-25 Imen Trabelsi , Dorra Ben Ayed , Noureddine Ellouze

Most Web page classification models typically apply the bag of words (BOW) model to represent the feature space. The original BOW representation, however, is unable to recognize semantic relationships between terms. One possible solution is…

机器学习 · 计算机科学 2010-04-28 Wongkot Sriurai , Phayung Meesad , Choochart Haruechaiyasak

Previous studies support the idea of merging auditory-based Gabor features with deep learning architectures to achieve robust automatic speech recognition, however, the cause behind the gain of such combination is still unknown. We believe…

计算与语言 · 计算机科学 2017-02-15 Angel Mario Castro Martinez , Sri Harish Mallidi , Bernd T. Meyer

Automatic speech recognition systems usually rely on spectral-based features, such as MFCC of PLP. These features are extracted based on prior knowledge such as, speech perception or/and speech production. Recently, convolutional neural…

机器学习 · 计算机科学 2015-04-17 Dimitri Palaz , Mathew Magimai Doss , Ronan Collobert

The aim of this paper is to provide and numerically test in the presence of measurement noise a procedure for target classification in wave imaging based on comparing frequency-dependent distribution descriptors with precomputed ones in a…

偏微分方程分析 · 数学 2018-06-21 Lorenzo Baldassari

We further exploit the representational power of Haar wavelet and present a novel low-level face representation named Shape Primitives Histogram (SPH) for face recognition. Since human faces exist abundant shape features, we address the…

计算机视觉与模式识别 · 计算机科学 2014-07-23 Sheng Huang , Dan Yang , Haopeng Zhang , Luwen Huangfu , Xiaohong Zhang

We propose to implicitly learn to extract geo-temporal image features, which are mid-level features related to when and where an image was captured, by explicitly optimizing for a set of location and time estimation tasks. To train our…

计算机视觉与模式识别 · 计算机科学 2019-09-18 Menghua Zhai , Tawfiq Salem , Connor Greenwell , Scott Workman , Robert Pless , Nathan Jacobs

Feature learning forms the cornerstone for tackling challenging learning problems in domains such as speech, computer vision and natural language processing. In this paper, we consider a novel class of matrix and tensor-valued features,…

机器学习 · 计算机科学 2015-04-21 Majid Janzamin , Hanie Sedghi , Anima Anandkumar

Chord recognition systems typically comprise an acoustic model that predicts chords for each audio frame, and a temporal model that casts these predictions into labelled chord segments. However, temporal models have been shown to only…

声音 · 计算机科学 2018-08-17 Filip Korzeniowski , Gerhard Widmer

In this paper, we propose a solution for improving the quality of temporal sound localization. We employ a multimodal fusion approach to combine visual and audio features. High-quality visual features are extracted using a state-of-the-art…

声音 · 计算机科学 2024-07-03 Yurui Huang , Yang Yang , Shou Chen , Xiangyu Wu , Qingguo Chen , Jianfeng Lu