English
Related papers

Related papers: TriSPrompt: A Hierarchical Soft Prompt Model for M…

200 papers

The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite its promise, real-world datasets often contain noisy labels…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Bikang Pan , Qun Li , Xiaoying Tang , Wei Huang , Zhen Fang , Feng Liu , Jingya Wang , Jingyi Yu , Ye Shi

The increasing proliferation of misinformation and its alarming impact have motivated both industry and academia to develop approaches for misinformation detection and fact checking. Recent advances on large language models (LLMs) have…

Computation and Language · Computer Science 2024-07-22 Sahar Tahmasebi , Eric Müller-Budack , Ralph Ewerth

Multimodal emotion and intent recognition is essential for automated human-computer interaction, It aims to analyze users' speech, text, and visual information to predict their emotions or intent. One of the significant challenges is that…

Artificial Intelligence · Computer Science 2025-07-09 Wei Zhang , Juan Chen , Yanbo J. Wang , En Zhu , Xuan Yang , Yiduo Wang

The field of Multimodal Sentiment Analysis (MSA) has recently witnessed an emerging direction seeking to tackle the issue of data incompleteness. Recognizing that the language modality typically contains dense sentiment information, we…

Computation and Language · Computer Science 2024-11-04 Haoyu Zhang , Wenbin Wang , Tianshu Yu

Visible-modal object tracking gives rise to a series of downstream multi-modal tracking tributaries. To inherit the powerful representations of the foundation model, a natural modus operandi for multi-modal tracking is full fine-tuning on…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Jiawen Zhu , Simiao Lai , Xin Chen , Dong Wang , Huchuan Lu

The proliferation of rumors on social media has become a major concern due to its ability to create a devastating impact. Manually assessing the veracity of social media messages is a very time-consuming task that can be much helped by…

Computer Vision and Pattern Recognition · Computer Science 2023-02-13 Abderrazek Azri , Cécile Favre , Nouria Harbi , Jérôme Darmont , Camille Noûs

Multimodal Large Language Models (MLLMs) demonstrate remarkable performance across a wide range of domains, with increasing emphasis on enhancing their zero-shot generalization capabilities for unseen tasks across various modalities.…

The development of accurate and scalable cross-modal image-text retrieval methods, where queries from one modality (e.g., text) can be matched to archive entries from another (e.g., remote sensing image) has attracted great attention in…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Georgii Mikriukov , Mahdyar Ravanbakhsh , Begüm Demir

In multimodal sentiment analysis, collecting text data is often more challenging than video or audio due to higher annotation costs and inconsistent automatic speech recognition (ASR) quality. To address this challenge, our study has…

Computation and Language · Computer Science 2025-03-25 Yuzhe Weng , Haotian Wang , Tian Gao , Kewei Li , Shutong Niu , Jun Du

Hierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex label hierarchy. Recently, the pretrained language models (PLM)have been widely adopted in HTC through a fine-tuning paradigm.…

Computation and Language · Computer Science 2022-10-11 Zihan Wang , Peiyi Wang , Tianyu Liu , Binghuai Lin , Yunbo Cao , Zhifang Sui , Houfeng Wang

Multi-sensor perception is crucial to ensure the reliability and accuracy in autonomous driving system, while multi-object tracking (MOT) improves that by tracing sequential movement of dynamic objects. Most current approaches for…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Wenwei Zhang , Hui Zhou , Shuyang Sun , Zhe Wang , Jianping Shi , Chen Change Loy

Multimodal sentiment analysis has gained significant attention due to the proliferation of multimodal content on social media. However, existing studies in this area rely heavily on large-scale supervised data, which is time-consuming and…

Computation and Language · Computer Science 2023-08-02 Xiaocui Yang , Shi Feng , Daling Wang , Pengfei Hong , Soujanya Poria

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yilin Ye , Shishi Xiao , Xingchen Zeng , Wei Zeng

Multimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather…

Machine Learning · Computer Science 2025-11-18 Yihua Wang , Qi Jia , Cong Xu , Feiyu Chen , Yuhan Liu , Haotian Zhang , Liang Jin , Lu Liu , Zhichun Wang

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Wangyuan Zhu , Jun Yu

Exploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the singlemodality scenario. However multimodal datasets often suffer from problems such as data misalignment and label inconsistencies, where…

Computer Vision and Pattern Recognition · Computer Science 2017-09-29 Sarah Taghavi Namin , Mohammad Najafi , Mathieu Salzmann , Lars Petersson

Misinformation undermines individual knowledge and affects broader societal narratives. Despite growing interest in the research community in multi-modal misinformation detection, existing methods exhibit limitations in capturing semantic…

Computation and Language · Computer Science 2024-10-22 Swarang Joshi , Siddharth Mavani , Joel Alex , Arnav Negi , Rahul Mishra , Ponnurangam Kumaraguru

Exploiting the foundation models (e.g., CLIP) to build a versatile keypoint detector has gained increasing attention. Most existing models accept either the text prompt (e.g., ``the nose of a cat''), or the visual prompt (e.g., support…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Changsheng Lu , Zheyuan Liu , Piotr Koniusz

Missing modalities are a common challenge in real-world multimodal learning scenarios, occurring during both training and testing. Existing methods for managing missing modalities often require the design of separate prompts for each…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Zhe Chen , Xun Lin , Yawen Cui , Zitong Yu

With the emergence of large pre-trained vison-language model like CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning tries to probe the beneficial information for…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Yinghui Xing , Qirui Wu , De Cheng , Shizhou Zhang , Guoqiang Liang , Peng Wang , Yanning Zhang