中文
相关论文

相关论文: Duality Temporal-channel-frequency Attention Enhan…

200 篇论文

Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications,…

声音 · 计算机科学 2025-10-31 Rinku Sebastian , Simon O'Keefe , Martin Trefzer

Neural network-based speaker recognition has achieved significant improvement in recent years. A robust speaker representation learns meaningful knowledge from both hard and easy samples in the training set to achieve good performance.…

音频与语音处理 · 电气工程与系统科学 2022-10-31 Ruijie Tao , Kong Aik Lee , Zhan Shi , Haizhou Li

Dynamic Facial Expression Recognition (DFER) facilitates the understanding of psychological intentions through non-verbal communication. Existing methods struggle to manage irrelevant information, such as background noise and redundant…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Meng-zhu Li , Quanxing Zha , Hongjun Wu

Human action recognition has become an important research focus in computer vision due to the wide range of applications where it is used. 3D Resnet-based CNN models, particularly MC3, R3D, and R(2+1)D, have different convolutional filters…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Mohammad Rasras , Iuliana Marin , Serban Radu , Irina Mocanu

Cognitive Language Processing (CLP), situated at the intersection of Natural Language Processing (NLP) and cognitive science, plays a progressively pivotal role in the domains of artificial intelligence, cognitive intelligence, and brain…

机器学习 · 计算机科学 2024-06-06 Weiguo Chen , Changjian Wang , Kele Xu , Yuan Yuan , Yanru Bai , Dongsong Zhang

Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageous under some adverse…

音频与语音处理 · 电气工程与系统科学 2022-07-19 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in…

声音 · 计算机科学 2025-08-12 Cunhang Fan , Sheng Zhang , Jingjing Zhang , Enrui Liu , Xinhui Li , Gangming Zhao , Zhao Lv

Speech enhancement in the time domain is becoming increasingly popular in recent years, due to its capability to jointly enhance both the magnitude and the phase of speech. In this work, we propose a dense convolutional network (DCN) with…

音频与语音处理 · 电气工程与系统科学 2021-03-09 Ashutosh Pandey , DeLiang Wang

We propose a learnable mel-frequency cepstral coefficient (MFCC) frontend architecture for deep neural network (DNN) based automatic speaker verification. Our architecture retains the simplicity and interpretability of MFCC-based features…

声音 · 计算机科学 2021-02-23 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

Attention modules for Convolutional Neural Networks (CNNs) are an effective method to enhance performance on multiple computer-vision tasks. While existing methods appropriately model channel-, spatial- and self-attention, they primarily…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Shantanu Jaiswal , Basura Fernando , Cheston Tan

Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and…

声音 · 计算机科学 2022-04-05 Ruiteng Zhang , Jianguo Wei , Wenhuan Lu , Lin Zhang , Yantao Ji , Junhai Xu , Xugang Lu

Deep neural networks (DNNs) have shown promising results for acoustic echo cancellation (AEC). But the DNN-based AEC models let through all near-end speakers including the interfering speech. In light of recent studies on personalized…

声音 · 计算机科学 2022-07-01 Shimin Zhang , Ziteng Wang , Yukai Ju , Yihui Fu , Yueyue Na , Qiang Fu , Lei Xie

Affective Computing has recently attracted the attention of the research community, due to its numerous applications in diverse areas. In this context, the emergence of video-based data allows to enrich the widely used spatial features with…

计算机视觉与模式识别 · 计算机科学 2021-02-19 Decky Aspandi , Federico Sukno , Björn Schuller , Xavier Binefa

Speaker recognition systems based on deep speaker embeddings have achieved significant performance in controlled conditions according to the results obtained for early NIST SRE (Speaker Recognition Evaluation) datasets. From the practical…

Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment…

音频与语音处理 · 电气工程与系统科学 2022-10-25 Chang Zeng , Xin Wang , Erica Cooper , Xiaoxiao Miao , Junichi Yamagishi

Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Sarah Ibrahimi , Xiaohang Sun , Pichao Wang , Amanmeet Garg , Ashutosh Sanan , Mohamed Omar

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio recordings and the speaker annotation is provided only at the…

声音 · 计算机科学 2018-06-25 Martin Karu , Tanel Alumäe

This paper proposes a multi-task learning network with phoneme-aware and channel-wise attentive learning strategies for text-dependent Speaker Verification (SV). In the proposed structure, the frame-level multi-task learning along with the…

声音 · 计算机科学 2021-06-28 Yan Liu , Zheng Li , Lin Li , Qingyang Hong

To reap the promising benefits of massive multiple-input multiple-output (MIMO) systems, accurate channel state information (CSI) is required through channel estimation. However, due to the complicated wireless propagation environment and…

信号处理 · 电气工程与系统科学 2024-11-04 Binggui Zhou , Xi Yang , Shaodan Ma , Feifei Gao , Guanghua Yang

Auditory attention is a selective type of hearing in which people focus their attention intentionally on a specific source of a sound or spoken words whilst ignoring or inhibiting other auditory stimuli. In some sense, the auditory…

机器学习 · 计算机科学 2021-10-26 Mahak Kothari , Shreyansh Joshi , Adarsh Nandanwar , Aadetya Jaiswal , Veeky Baths
‹ 上一页 1 8 9 10 下一页 ›