中文
相关论文

相关论文: Improved Feature Extraction Network for Neuro-Orie…

200 篇论文

Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel approach that…

音频与语音处理 · 电气工程与系统科学 2024-12-24 Anup Singh , Kris Demuynck , Vipul Arora

Speech Emotion Recognition (SER) task has known significant improvements over the last years with the advent of Deep Neural Networks (DNNs). However, even the most successful methods are still rather failing when adaptation to specific…

音频与语音处理 · 电气工程与系统科学 2021-04-16 Clément Le Moine , Nicolas Obin , Axel Roebel

In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetric Cross Attention (ACA) to replace temporal pooling. ACA is…

FullSubNet has shown its promising performance on speech enhancement by utilizing both fullband and subband information. However, the relationship between fullband and subband in FullSubNet is achieved by simply concatenating the output of…

声音 · 计算机科学 2022-11-11 Jun Chen , Wei Rao , Zilin Wang , Zhiyong Wu , Yannan Wang , Tao Yu , Shidong Shang , Helen Meng

Extracting the desired speech from a mixture is a meaningful and challenging task. The end-to-end DNN-based methods, though attractive, face the problem of generalization. In this paper, we explore a sequential approach for target speech…

音频与语音处理 · 电气工程与系统科学 2020-11-02 Zhaoyi Gu , Lele Liao , Kai Chen , Jing Lu

Inner speech recognition has gained enormous interest in recent years due to its applications in rehabilitation, developing assistive technology, and cognitive assessment. However, since language and speech productions are a complex…

Given an audio clip and a reference face image, the goal of the talking head generation is to generate a high-fidelity talking head video. Although some audio-driven methods of generating talking head videos have made some achievements in…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Jianrong Wang , Yaxin Zhao , Li Liu , Tianyi Xu , Qi Li , Sen Li

3D speech enhancement can effectively improve the auditory experience and plays a crucial role in augmented reality technology. However, traditional convolutional-based speech enhancement methods have limitations in extracting dynamic voice…

音频与语音处理 · 电气工程与系统科学 2023-11-21 Han Yin , Jisheng Bai , Mou Wang , Siwei Huang , Yafei Jia , Jianfeng Chen

A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetically discriminative/speaker discriminative DNNs as feature extractors for speaker verification has shown…

计算与语言 · 计算机科学 2017-01-04 Shi-Xiong Zhang , Zhuo Chen , Yong Zhao , Jinyu Li , Yifan Gong

Recently, attention mechanisms have been applied successfully in neural network-based speaker verification systems. Incorporating the Squeeze-and-Excitation block into convolutional neural networks has achieved remarkable performance.…

音频与语音处理 · 电气工程与系统科学 2022-07-12 Mufan Sang , John H. L. Hansen

Target speaker extraction aims at extracting the target speaker from a mixture of multiple speakers exploiting auxiliary information about the target speaker. In this paper, we consider a complete time-domain target speaker extraction…

音频与语音处理 · 电气工程与系统科学 2022-05-30 Ragini Sinha , Marvin Tammen , Christian Rollwage , Simon Doclo

We propose a mixed deep neural network strategy, incorporating parallel combination of Convolutional (CNN) and Recurrent Neural Networks (RNN), cascaded with deep autoencoders and fully connected layers towards automatic identification of…

机器学习 · 计算机科学 2019-04-10 Pramit Saha , Sidney Fels

In this study, we propose an ensemble learning framework for electroencephalogram-based overt speech classification, leveraging denoising diffusion probabilistic models with varying convolutional kernel sizes. The ensemble comprises three…

声音 · 计算机科学 2024-11-15 Soowon Kim , Ha-Na Jo , Eunyeong Ko

In response to the increasing interest in human--machine communication across various domains, this paper introduces a novel approach called iPhonMatchNet, which addresses the challenge of barge-in scenarios, wherein user speech overlaps…

音频与语音处理 · 电气工程与系统科学 2023-12-15 Yong-Hyeok Lee , Namhyun Cho

This paper introduces a lightweight deep learning model for real-time speech enhancement, designed to operate efficiently on resource-constrained devices. The proposed model leverages a compact architecture that facilitates rapid inference…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Shuubham Ojha , Felix Gervits , Carol Espy-Wilson

Extracting correlation features between codes-words with high computational efficiency is crucial to steganalysis of Voice over IP (VoIP) streams. In this paper, we utilized attention mechanisms, which have recently attracted enormous…

多媒体 · 计算机科学 2020-04-16 Hao Yang , ZhongLiang Yang , YongJian Bao , Sheng Liu , YongFeng Huang

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts…

机器学习 · 计算机科学 2022-11-08 Itai Gat , Hagai Aronowitz , Weizhong Zhu , Edmilson Morais , Ron Hoory

The reliability of using fully convolutional networks (FCNs) has been successfully demonstrated by recent studies in many speech applications. One of the most popular variants of these FCNs is the `U-Net', which is an encoder-decoder…

音频与语音处理 · 电气工程与系统科学 2021-11-10 Vinay Kothapally , Wei Xia , Shahram Ghorbani , John H. L. Hansen , Wei Xue , Jing Huang

Conformer and Mamba have achieved strong performance in speech modeling but face limitations in speaker diarization. Mamba is efficient but struggles with local details and nonlinear patterns. Conformer's self-attention incurs high memory…

声音 · 计算机科学 2026-01-28 Zhen Liao , Gaole Dai , Mengqiao Chen , Wenqing Cheng , Wei Xu

Despite remarkable progress, automatic speaker verification (ASV) systems typically lack the transparency required for high-accountability applications. Motivated by how human experts perform forensic speaker comparison (FSC), we propose a…

音频与语音处理 · 电气工程与系统科学 2026-04-07 Yi Ma , Shuai Wang , Tianchi Liu , Haizhou Li
‹ 上一页 1 8 9 10 下一页 ›