中文
相关论文

相关论文: An Efficient End-to-End Transformer with Progressi…

200 篇论文

Humans are sophisticated at reading interlocutors' emotions from multimodal signals, such as speech contents, voice tones and facial expressions. However, machines might struggle to understand various emotions due to the difficulty of…

人工智能 · 计算机科学 2022-12-21 Feng Qiu , Wanzeng Kong , Yu Ding

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

Recently, referring image segmentation has aroused widespread interest. Previous methods perform the multi-modal fusion between language and vision at the decoding side of the network. And, linguistic feature interacts with visual feature…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Guang Feng , Zhiwei Hu , Lihe Zhang , Huchuan Lu

Speech emotion recognition is a challenging research topic that plays a critical role in human-computer interaction. Multimodal inputs further improve the performance as more emotional information is used. However, existing studies learn…

声音 · 计算机科学 2023-02-28 Weidong Chen , Xiaofeng Xing , Xiangmin Xu , Jichen Yang , Jianxin Pang

Recently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments. One approach is the attention-based encoder-decoder framework that learns a mapping…

计算与语言 · 计算机科学 2017-02-02 Suyoun Kim , Takaaki Hori , Shinji Watanabe

In this work, we present a hybrid CTC/Attention model based on a ResNet-18 and Convolution-augmented transformer (Conformer), that can be trained in an end-to-end manner. In particular, the audio and visual encoders learn to extract…

计算机视觉与模式识别 · 计算机科学 2021-02-15 Pingchuan Ma , Stavros Petridis , Maja Pantic

Cross-modal transformers have demonstrated superiority in various vision tasks by effectively integrating different modalities. This paper first critiques prior token exchange methods which replace less informative tokens with inter-modal…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Ding Jia , Jianyuan Guo , Kai Han , Han Wu , Chao Zhang , Chang Xu , Xinghao Chen

Emotion Recognition in Conversations (ERC) is hard because discriminative evidence is sparse, localized, and often asynchronous across modalities. We center ERC on emotion hotspots and present a unified model that detects per-utterance…

计算与语言 · 计算机科学 2025-10-13 Yu Liu , Hanlei Shi , Haoxun Li , Yuqing Sun , Yuxuan Ding , Linlin Gong , Leyuan Qu , Taihao Li

Existing top-performance autonomous driving systems typically rely on the multi-modal fusion strategy for reliable scene understanding. This design is however fundamentally restricted due to overlooking the modality-specific strengths and…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Zeyu Yang , Nan Song , Wei Li , Xiatian Zhu , Li Zhang , Philip H. S. Torr

With the rapid advancement of deep learning, attention mechanisms have become indispensable in electroencephalography (EEG) signal analysis, significantly enhancing Brain-Computer Interface (BCI) applications. This paper presents a…

信号处理 · 电气工程与系统科学 2025-07-08 Jiyuan Wang , Weishan Ye , Jialin He , Li Zhang , Gan Huang , Zhuliang Yu , Zhen Liang

Spoken languages often utilise intonation, rhythm, intensity, and structure, to communicate intention, which can be interpreted differently depending on the rhythm of speech of their utterance. These speech acts provide the foundation of…

End-to-end task-oriented dialog systems usually suffer from the challenge of incorporating knowledge bases. In this paper, we propose a novel yet simple end-to-end differentiable model called memory-to-sequence (Mem2Seq) to address this…

计算与语言 · 计算机科学 2018-05-22 Andrea Madotto , Chien-Sheng Wu , Pascale Fung

End-to-end learning framework is useful for building dialog systems for its simplicity in training and efficiency in model updating. However, current end-to-end approaches only consider user semantic inputs in learning and under-utilize…

计算与语言 · 计算机科学 2019-07-04 Weiyan Shi , Zhou Yu

In the pathway toward Artificial General Intelligence (AGI), understanding human's affection is essential to enhance machine's cognition abilities. For achieving more sensual human-AI interaction, Multimodal Affective Computing (MAC) in…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Ronghao Lin , Ying Zeng , Sijie Mai , Haifeng Hu

In this paper, we make the explicit connection between image segmentation methods and end-to-end diarization methods. From these insights, we propose a novel, fully end-to-end diarization model, EEND-M2F, based on the Mask2Former…

声音 · 计算机科学 2024-01-24 Marc Härkönen , Samuel J. Broughton , Lahiru Samarakoon

Transformer-based models have significantly improved performance across a range of multimodal understanding tasks, such as visual question answering and action recognition. However, multimodal Transformers significantly suffer from a…

机器学习 · 计算机科学 2024-02-26 Sungjin Park , Edward Choi

Emotion recognition has become a popular topic of interest, especially in the field of human computer interaction. Previous works involve unimodal analysis of emotion, while recent efforts focus on multi-modal emotion recognition from…

计算与语言 · 计算机科学 2019-03-11 Chan Woo Lee , Kyu Ye Song , Jihoon Jeong , Woo Yong Choi

This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a…

声音 · 计算机科学 2024-01-29 Shreyan Chowdhury , Gerhard Widmer

Multimodal sentiment analysis enhances conventional sentiment analysis, which traditionally relies solely on text, by incorporating information from different modalities such as images, text, and audio. This paper proposes a novel…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Taoxu Zhao , Meisi Li , Kehao Chen , Liye Wang , Xucheng Zhou , Kunal Chaturvedi , Mukesh Prasad , Ali Anaissi , Ali Braytee

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy…

多媒体 · 计算机科学 2024-10-01 Mengying Ge , Mingyang Li , Dongkai Tang , Pengbo Li , Kuo Liu , Shuhao Deng , Songbai Pu , Long Liu , Yang Song , Tao Zhang