中文
相关论文

相关论文: DMF2Mel: A Dynamic Multiscale Fusion Network for E…

200 篇论文

Learning an effective speaker representation is crucial for achieving reliable performance in speaker verification tasks. Speech signals are high-dimensional, long, and variable-length sequences containing diverse information at each…

音频与语音处理 · 电气工程与系统科学 2023-08-25 Wei Xia , John H. L. Hansen

Event-based semantic segmentation explores the potential of event cameras, which offer high dynamic range and fine temporal resolution, to achieve robust scene understanding in challenging environments. Despite these advantages, the task…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Zhijiang Li , Haoran He

Decoding speech from stereo-electroencephalography (sEEG) signals has emerged as a promising direction for brain-computer interfaces (BCIs). Its clinical applicability, however, is limited by the inherent non-stationarity of neural signals,…

人机交互 · 计算机科学 2025-09-30 Suli Wang , Yang-yang Li , Siqi Cai , Haizhou Li

3D face alignment of monocular images is a crucial process in the recognition of faces with disguise.3D face reconstruction facilitated by alignment can restore the face structure which is helpful in detcting disguise interference.This…

计算机视觉与模式识别 · 计算机科学 2019-09-02 Lei Jiang Xiao-Jun Wu Josef Kittler

Meta-learning facilitates few-shot hyperspectral target detection (HTD), but adapting deep backbones remains challenging. Full-parameter fine-tuning is inefficient and prone to overfitting, and existing methods largely ignore the…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Luqi Gong , Qixin Xie , Yue Chen , Ziqiang Chen , Fanda Fan , Shuai Zhao , Chao Li

Feature-level fusion shows promise in collaborative perception (CP) through balanced performance and communication bandwidth trade-off. However, its effectiveness critically relies on input feature quality. The acquisition of high-quality…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Chengchang Tian , Jianwei Ma , Yan Huang , Zhanye Chen , Honghao Wei , Hui Zhang , Wei Hong

FullSubNet is our recently proposed real-time single-channel speech enhancement network that achieves outstanding performance on the Deep Noise Suppression (DNS) Challenge dataset. A number of variants of FullSubNet have been proposed, but…

音频与语音处理 · 电气工程与系统科学 2023-03-08 Xiang Hao , Xiaofei Li

Softmax attention is the principle backbone of foundation models for various artificial intelligence applications, yet its quadratic complexity in sequence length can limit its inference throughput in long-context settings. To address this…

机器学习 · 计算机科学 2024-12-10 Jerome Sieber , Carmen Amo Alonso , Alexandre Didier , Melanie N. Zeilinger , Antonio Orvieto

Long-term time series forecasting (LTSF) is hampered by the challenge of modeling complex dependencies that span multiple temporal scales and frequency resolutions. Existing methods, including Transformer and MLP-based models, often…

机器学习 · 计算机科学 2025-09-22 Qianyang Li , Xingjun Zhang , Shaoxun Wang , Jia Wei

Accurate medical image segmentation remains challenging due to blurred lesion boundaries (LBA), loss of high-frequency details (LHD), and difficulty in modeling long-range anatomical structures (DC-LRSS). Vision Mamba employs…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Ze Rong , ZiYue Zhao , Zhaoxin Wang , Lei Ma

Motivated by the attention mechanism of the human visual system and recent developments in the field of machine translation, we introduce our attention-based and recurrent sequence to sequence autoencoders for fully unsupervised…

音频与语音处理 · 电气工程与系统科学 2020-05-20 Shahin Amiriparian , Pawel Winokurow , Vincent Karas , Sandra Ottl , Maurice Gerczuk , Björn W. Schuller

Complex systems such as aircraft engines, turbines, and industrial machinery often operate under dynamically changing conditions. These varying operating conditions can substantially influence degradation behavior and make prognostic…

机器学习 · 计算机科学 2026-04-14 Yuqi Su , Xiaolei Fang

Ophthalmic diseases pose a significant global health burden. However, traditional diagnostic methods and existing monocular image-based deep learning approaches often overlook the pathological correlations between the two eyes. In practical…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Guohao Huo , Zibo Lin , Zitong Wang , Ruiting Dai , Hao Tang

Resting-state functional magnetic resonance imaging (rs-fMRI) is a noninvasive technique pivotal for understanding human neural mechanisms of intricate cognitive processes. Most rs-fMRI studies compute a single static functional…

It becomes urgent to design effective anti-spoofing algorithms for vulnerable automatic speaker verification systems due to the advancement of high-quality playback devices. Current studies mainly treat anti-spoofing as a binary…

计算机视觉与模式识别 · 计算机科学 2023-01-19 Yongqiang Dou , Haocheng Yang , Maolin Yang , Yanyan Xu , Dengfeng Ke

Deep neural networks face several challenges in hyperspectral image classification, including insufficient utilization of joint spatial-spectral information, gradient vanishing with increasing depth, and overfitting. To enhance feature…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Guandong Li , Mengxia Ye

Frequency modulation features capture the fine structure of speech formants that constitute beneficial and supplementary to the traditional energy-based cepstral features. Improvements have been demonstrated mainly in GMM-HMM systems for…

声音 · 计算机科学 2019-09-04 Isidoros Rodomagoulakis , Petros Maragos

This paper presents a novel neural model - Dynamic Fusion Network (DFN), for machine reading comprehension (MRC). DFNs differ from most state-of-the-art models in their use of a dynamic multi-strategy attention process, in which passages,…

计算与语言 · 计算机科学 2018-02-28 Yichong Xu , Jingjing Liu , Jianfeng Gao , Yelong Shen , Xiaodong Liu

Hearing aids (HAs) are widely used to provide personalized speech enhancement (PSE) services, improving the quality of life for individuals with hearing loss. However, HA performance significantly declines in noisy environments as it treats…

音频与语音处理 · 电气工程与系统科学 2025-09-10 Ye Ni , Ruiyu Liang , Xiaoshuai Hao , Jiaming Cheng , Qingyun Wang , Chengwei Huang , Cairong Zou , Wei Zhou , Weiping Ding , Björn W. Schuller

Understanding sentiment in complex textual expressions remains a fundamental challenge in affective computing. To address this, we propose a Dynamic Fusion Learning Model (DyFuLM), a multimodal framework designed to capture both…

计算与语言 · 计算机科学 2025-12-02 Ruohan Zhou , Jiachen Yuan , Churui Yang , Wenzheng Huang , Guoyan Zhang , Shiyao Wei , Jiazhen Hu , Ning Xin , Md Maruf Hasan