中文
相关论文

相关论文: A new Speech Feature Fusion method with cross gate…

200 篇论文

We introduce PGF-Net (Progressive Gated-Fusion Network), a novel deep learning framework designed for efficient and interpretable multimodal sentiment analysis. Our framework incorporates three primary innovations. Firstly, we propose a…

机器学习 · 计算机科学 2025-08-25 Bin Wen , Tien-Ping Tan

The joint training framework for speech enhancement and recognition methods have obtained quite good performances for robust end-to-end automatic speech recognition (ASR). However, these methods only utilize the enhanced feature as the…

声音 · 计算机科学 2020-11-10 Cunhang Fan , Jiangyan Yi , Jianhua Tao , Zhengkun Tian , Bin Liu , Zhengqi Wen

While attention-based approaches have shown considerable progress in enhancing image fusion and addressing the challenges posed by long-range feature dependencies, their efficacy in capturing local features is compromised by the lack of…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Jingjing Liu , Li Zhang , Xiaoyang Zeng , Wanquan Liu , Jianhua Zhang

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

计算机与社会 · 计算机科学 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

In this paper, we first present a new variant of Gaussian restricted Boltzmann machine (GRBM) called multivariate Gaussian restricted Boltzmann machine (MGRBM), with its definition and learning algorithm. Then we propose using a learned…

计算与语言 · 计算机科学 2013-09-25 Xin Zheng , Zhiyong Wu , Helen Meng , Weifeng Li , Lianhong Cai

In this paper, we propose an effective feature extraction algorithm, called Multi-Subregion based Correlation Filter Bank (MS-CFB), for robust face recognition. MS-CFB combines the benefits of global-based and local-based feature extraction…

计算机视觉与模式识别 · 计算机科学 2016-03-25 Yan Yan , Hanzi Wang , David Suter

Convolutional Neural Networks (CNN) have been used in Automatic Speech Recognition (ASR) to learn representations directly from the raw signal instead of hand-crafted acoustic features, providing a richer and lossless input signal. Recent…

声音 · 计算机科学 2020-02-12 Paul-Gauthier Noé , Titouan Parcollet , Mohamed Morchid

This paper will describe a novel approach to the cocktail party problem that relies on a fully convolutional neural network (FCN) architecture. The FCN takes noisy audio data as input and performs nonlinear, filtering operations to produce…

声音 · 计算机科学 2018-07-24 Frank Longueira , Sam Keene

In this paper, we propose a novel Convolutional Neural Network (CNN) architecture for learning multi-scale feature representations with good tradeoffs between speed and accuracy. This is achieved by using a multi-branch network, which has…

计算机视觉与模式识别 · 计算机科学 2019-08-01 Chun-Fu Chen , Quanfu Fan , Neil Mallinar , Tom Sercu , Rogerio Feris

Audio-visual speech enhancement system is regarded to be one of promising solutions for isolating and enhancing speech of desired speaker. Conventional methods focus on predicting clean speech spectrum via a naive convolution neural network…

音频与语音处理 · 电气工程与系统科学 2022-09-28 Xinmeng Xu , Jianjun Hao

In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding layer. To understand…

音频与语音处理 · 电气工程与系统科学 2018-09-13 Suwon Shon , Hao Tang , James Glass

Speech Emotion Recognition (SER) is the use of machines to detect the emotional state of humans based on the speech, which is gaining importance in natural human-computer interaction. Speech is a very valuable source of information, as…

Deep Convolutional Neural Networks (CNNs) are capable of learning unprecedentedly effective features from images. Some researchers have struggled to enhance the parameters' efficiency using grouped convolution. However, the relation between…

计算机视觉与模式识别 · 计算机科学 2017-06-22 Yujia Chen , Ce Li

The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local…

声音 · 计算机科学 2022-02-16 Tianchi Liu , Rohan Kumar Das , Kong Aik Lee , Haizhou Li

In this paper, we propose to use deep 3-dimensional convolutional networks (3D CNNs) in order to address the challenge of modelling spectro-temporal dynamics for speech emotion recognition (SER). Compared to a hybrid of Convolutional Neural…

计算与语言 · 计算机科学 2017-08-18 Jaebok Kim , Khiet P. Truong , Gwenn Englebienne , Vanessa Evers

Feature selection is a preprocessing step which plays a crucial role in the domain of machine learning and data mining. Feature selection methods have been shown to be effctive in removing redundant and irrelevant features, improving the…

机器学习 · 计算机科学 2021-06-01 Xiongshi Deng , Min Li , Lei Wang , Qikang Wan

This paper introduces a novel convolutional neural networks (CNN) framework tailored for end-to-end audio deep learning models, presenting advancements in efficiency and explainability. By benchmarking experiments on three standard speech…

声音 · 计算机科学 2024-05-06 Linh Vu , Thu Tran , Wern-Han Lim , Raphael Phan

The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on…

音频与语音处理 · 电气工程与系统科学 2025-04-03 Tianqin Zheng , Jilu Jin , Hanchen Pei , Gongping Huang , Jingdong Chen , Jacob Benesty

Emotion recognition from speech is a challenging task. Re-cent advances in deep learning have led bi-directional recur-rent neural network (Bi-RNN) and attention mechanism as astandard method for speech emotion recognition, extractingand…

声音 · 计算机科学 2021-06-09 Zixuan Peng , Yu Lu , Shengfeng Pan , Yunfeng Liu

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform…

音频与语音处理 · 电气工程与系统科学 2022-11-22 Takenori Yoshimura , Shinji Takaki , Kazuhiro Nakamura , Keiichiro Oura , Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda