English
Related papers

Related papers: TECO: Improving Multimodal Intent Recognition with…

200 papers

Composed image retrieval which combines a reference image and a text modifier to identify the desired target image is a challenging task, and requires the model to comprehend both vision and language modalities and their interactions.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Shu Zhao , Huijuan Xu

Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Haoshuo Zhang , Yufei Bo , Hongwei Zhang , Meixia Tao

Referring Expression Comprehension (REC) aims to localize specified entities or regions in an image based on natural language descriptions. While existing methods handle single-entity localization, they often ignore complex inter-entity…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Yizhi Hu , Zezhao Tian , Xingqun Qi , Chen Su , Bingkun Yang , Junhui Yin , Muyi Sun , Man Zhang , Zhenan Sun

Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding functions that map an instance to a single point in the shared…

Computer Vision and Pattern Recognition · Computer Science 2019-07-18 Yale Song , Mohammad Soleymani

Multimodal Entity Linking (MEL) aims to link ambiguous mentions within multimodal contexts to associated entities in a multimodal knowledge base. Existing approaches to MEL introduce multimodal interaction and fusion mechanisms to bridge…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Zhiwei Hu , Víctor Gutiérrez-Basulto , Zhiliang Xiang , Ru Li , Jeff Z. Pan

Multimodal intent recognition aims to infer human intents by jointly modeling various modalities, playing a pivotal role in real-world dialogue systems. However, current methods struggle to model hierarchical semantics underlying complex…

Multimedia · Computer Science 2026-03-05 Qianrui Zhou , Hua Xu , Yunjin Gu , Yifan Wang , Songze Li , Hanlei Zhang

Image-sentence retrieval has attracted extensive research attention in multimedia and computer vision due to its promising application. The key issue lies in jointly learning the visual and textual representation to accurately estimate…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Xuri Ge , Fuhai Chen , Songpei Xu , Fuxiang Tao , Joemon M. Jose

We propose a framework for multimodal sentiment analysis and emotion recognition using convolutional neural network-based feature extraction from text and visual modalities. We obtain a performance improvement of 10% over the state of the…

Multimedia · Computer Science 2017-08-01 Erik Cambria , Devamanyu Hazarika , Soujanya Poria , Amir Hussain , R. B. V. Subramaanyam

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize…

Information Retrieval · Computer Science 2024-12-17 Zelong Sun , Dong Jing , Guoxing Yang , Nanyi Fei , Zhiwu Lu

Multimodal emotion recognition in conversation (MERC) and multimodal emotion-cause pair extraction (MECPE) have recently garnered significant attention. Emotions are the expression of affect or feelings; responses to specific events, or…

Computation and Language · Computer Science 2024-10-10 Guimin Hu , Zhihong Zhu , Daniel Hershcovich , Lijie Hu , Hasti Seifi , Jiayuan Xie

Referring image segmentation aims at segmenting the foreground masks of the entities that can well match the description given in the natural language expression. Previous approaches tackle this problem using implicit feature interaction…

Computer Vision and Pattern Recognition · Computer Science 2020-10-02 Shaofei Huang , Tianrui Hui , Si Liu , Guanbin Li , Yunchao Wei , Jizhong Han , Luoqi Liu , Bo Li

Emotion Prediction in Conversation (EPC) aims to forecast the emotions of forthcoming utterances by utilizing preceding dialogues. Previous EPC approaches relied on simple context modeling for emotion extraction, overlooking fine-grained…

Multimedia · Computer Science 2024-08-09 Haoxiang Shi , Ziqi Liang , Jun Yu

Previous studies on multimodal fake news detection mainly focus on the alignment and integration of cross-modal features, as well as the application of text-image consistency. However, they overlook the semantic enhancement effects of large…

Multimedia · Computer Science 2025-07-21 Peican Zhu , Yubo Jing , Le Cheng , Bin Chen , Xiaodong Cui , Lianwei Wu , Keke Tang

Semantic communication focuses on transmitting task-relevant semantic information, aiming for intent-oriented communication. While existing systems improve efficiency by extracting key semantics, they still fail to deeply understand and…

Information Theory · Computer Science 2025-08-14 Peigen Ye , Jingpu Duan , Hongyang Du , Yulan Guo

Multimodal sentiment analysis aims to identify the emotions expressed by individuals through visual, language, and acoustic cues. However, most existing research assume that all modalities are available during both training and testing,…

Sound · Computer Science 2026-04-21 Weide Liu , Huijing Zhan

Linguistic knowledge has brought great benefits to scene text recognition by providing semantics to refine character sequences. However, since linguistic knowledge has been applied individually on the output sequence, previous methods have…

Computer Vision and Pattern Recognition · Computer Science 2022-08-16 Byeonghu Na , Yoonsik Kim , Sungrae Park

Identifying user intent from mobile UI operation trajectories is critical for advancing UI understanding and enabling task automation agents. While Multimodal Large Language Models (MLLMs) excel at video understanding tasks, their real-time…

Artificial Intelligence · Computer Science 2025-12-23 Zhe Yang , Xiaoshuang Sheng , Zhengnan Zhang , Jidong Wu , Zexing Wang , Xin He , Shenghua Xu , Guanjing Xiong

The goal of Text-to-Image Person Retrieval (TIPR) is to retrieve specific person images according to the given textual descriptions. A primary challenge in this task is bridging the substantial representational gap between visual and…

Computation and Language · Computer Science 2025-01-20 Delong Liu , Haiwen Li , Zhicheng Zhao , Yuan Dong

The objective of image captioning models is to bridge the gap between the visual and linguistic modalities by generating natural language descriptions that accurately reflect the content of input images. In recent years, researchers have…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Alessandro Nicolosi , Rita Cucchiara

Conversational multimodal understanding aims to infer the meaning or label of the current utterance from its preceding dialogue context together with textual, acoustic, and visual signals. Existing methods mainly strengthen contextual…

Multimedia · Computer Science 2026-04-29 Zhaoyan Pan , Hengyang Zhou , Xiangdong Li , Yuning Wang , Ye Lou , Jiatong Pan , Ji Zhou , Wei Zhang