中文
相关论文

相关论文: Acoustic scene classification using multi-layer te…

200 篇论文

In this paper, we describe in detail our systems for DCASE 2020 Task 4. The systems are based on the 1st-place system of DCASE 2019 Task 4, which adopts weakly-supervised framework with an attention-based embedding-level pooling module and…

声音 · 计算机科学 2020-11-03 Yuxin Huang , Liwei Lin , Shuo Ma , Xiangdong Wang , Hong Liu , Yueliang Qian , Min Liu , Kazushige Ouch

In speaker verification, the extraction of voice representations is mainly based on the Residual Neural Network (ResNet) architecture. ResNet is built upon convolution layers which learn filters to capture local spatial patterns along all…

音频与语音处理 · 电气工程与系统科学 2021-09-14 Mickael Rouvier , Pierre-Michel Bousquet

Semantic segmentation is a fundamental task in computer vision, which can be considered as a per-pixel classification problem. Recently, although fully convolutional neural network (FCN) based approaches have made remarkable progress in…

计算机视觉与模式识别 · 计算机科学 2018-04-24 Chen-Wei Xie , Hong-Yu Zhou , Jianxin Wu

Inspired by the mammal's auditory localization pathway, in this paper we propose a pure spiking neural network (SNN) based computational model for precise sound localization in the noisy real-world environment, and implement this algorithm…

音频与语音处理 · 电气工程与系统科学 2020-07-08 Zihan Pan , Malu Zhang , Jibin Wu , Haizhou Li

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a…

Convolutional Neural Networks (CNNs) are effective models for reducing spectral variations and modeling spectral correlations in acoustic features for automatic speech recognition (ASR). Hybrid speech recognition systems incorporating CNNs…

In this paper, we propose a speaker verification method by an Attentive Multi-scale Convolutional Recurrent Network (AMCRN). The proposed AMCRN can acquire both local spatial information and global sequential information from the input…

音频与语音处理 · 电气工程与系统科学 2023-06-02 Yanxiong Li , Zhongjie Jiang , Wenchang Cao , Qisheng Huang

In this work, we propose to extend a state-of-the-art multi-source localization system based on a convolutional recurrent neural network and Ambisonics signals. We significantly improve the performance of the baseline network by changing…

声音 · 计算机科学 2021-05-06 Pierre-Amaury Grumiaux , Srdan Kitic , Laurent Girin , Alexandre Guérin

Acoustic Scene Classification (ASC) identifies an environment based on an audio signal. This paper explores ASC in low-resource conditions and proposes a novel model, DS-FlexiNet, which combines depthwise separable convolutions from…

音频与语音处理 · 电气工程与系统科学 2025-04-29 Zhi Chen , Yun-Fei Shao , Yong Ma , Mingsheng Wei , Le Zhang , Wei-Qiang Zhang

This technical report describes the IOA team's submission for TASK1A of DCASE2019 challenge. Our acoustic scene classification (ASC) system adopts a data augmentation scheme employing generative adversary networks. Two major classifiers, 1D…

音频与语音处理 · 电气工程与系统科学 2019-07-17 Hangting Chen , Zuozhen Liu , Zongming Liu , Pengyuan Zhang , Yonghong Yan

In this study, we propose advancing all-neural speech recognition by directly incorporating attention modeling within the Connectionist Temporal Classification (CTC) framework. In particular, we derive new context vectors using time…

计算与语言 · 计算机科学 2018-03-16 Amit Das , Jinyu Li , Rui Zhao , Yifan Gong

Current deep learning based video classification architectures are typically trained end-to-end on large volumes of data and require extensive computational resources. This paper aims to exploit audio-visual information in video…

计算机视觉与模式识别 · 计算机科学 2020-12-21 Feiyan Hu , Eva Mohedano , Noel O'Connor , Kevin McGuinness

When domain experts are needed to perform data annotation for complex machine-learning tasks, reducing annotation effort is crucial in order to cut down time and expenses. For cases when there are no annotations available, one approach is…

机器学习 · 计算机科学 2022-06-22 Einari Vaaras , Manu Airaksinen , Okko Räsänen

When recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels.…

音频与语音处理 · 电气工程与系统科学 2022-11-04 Sofoklis Kakouros , Themos Stafylakis , Ladislav Mosner , Lukas Burget

This paper explores the impact of dimensionality reduction and pooling methods for Environmental Sound Classification (ESC) using lightweight CNNs. We evaluate Sparse Salient Region Pooling (SSRP) and its variants, SSRP-Basic (SSRP-B) and…

信号处理 · 电气工程与系统科学 2025-11-14 Parinaz Binandeh Dehaghani , Danilo Pena , A. Pedro Aguiar

Deep Convolutional Neural Networks (CNN) have exhibited superior performance in many visual recognition tasks including image classification, object detection, and scene label- ing, due to their large learning capacity and resistance to…

计算机视觉与模式识别 · 计算机科学 2016-10-12 Miao Sun , Tony X. Han , Xun Xu , Ming-Chang Liu , Ahmad Khodayari-Rostamabad

This paper proposes to use low-level spatial features extracted from multichannel audio for sound event detection. We extend the convolutional recurrent neural network to handle more than one type of these multichannel features by learning…

声音 · 计算机科学 2017-06-09 Sharath Adavanne , Pasi Pertilä , Tuomas Virtanen

In this paper, we describe in detail the system we submitted to DCASE2019 task 4: sound event detection (SED) in domestic environments. We employ a convolutional neural network (CNN) with an embedding-level attention pooling module to solve…

音频与语音处理 · 电气工程与系统科学 2019-09-16 Liwei Lin , Xiangdong Wang , Hong Liu , Yueliang Qian

Continuous sign language recognition (cSLR) is a public significant task that transcribes a sign language video into an ordered gloss sequence. It is important to capture the fine-grained gloss-level details, since there is no explicit…

计算机视觉与模式识别 · 计算机科学 2021-07-28 Pan Xie , Zhi Cui , Yao Du , Mengyi Zhao , Jianwei Cui , Bin Wang , Xiaohui Hu

A major advantage of a deep convolutional neural network (CNN) is that the focused receptive field size is increased by stacking multiple convolutional layers. Accordingly, the model can explore the long-range dependency of features from…

声音 · 计算机科学 2020-06-17 Xugang Lu , Peng Shen , Sheng Li , Yu Tsao , Hisashi Kawai
‹ 上一页 1 8 9 10 下一页 ›