中文
相关论文

相关论文: MuVaC: A Variational Causal Framework for Multimod…

200 篇论文

Major Adverse Cardiovascular Events (MACE) remain the leading cause of mortality globally, as reported in the Global Disease Burden Study 2021. Opportunistic screening leverages data collected from routine health check-ups and multimodal…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Jialu Pi , Juan Maria Farina , Rimita Lahiri , Jiwoong Jeong , Archana Gurudu , Hyung-Bok Park , Chieh-Ju Chao , Chadi Ayoub , Reza Arsanjani , Imon Banerjee

Multimodal Sentiment Analysis (MSA) aims to predict sentiment from language, acoustic, and visual data in videos. However, imbalanced unimodal performance often leads to suboptimal fused representations. Existing approaches typically adopt…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Dingkang Yang , Mingcheng Li , Xuecheng Wu , Zhaoyu Chen , Kaixun Jiang , Keliang Liu , Peng Zhai , Lihua Zhang

Multimodal large language models have recently shown promising progress in visual mathematical reasoning. However, their performance is often limited by a critical yet underexplored bottleneck: inaccurate visual perception. Through…

人工智能 · 计算机科学 2026-03-10 Peijin Xie , Zhen Xu , Bingquan Liu , Baoxun Wang

We present a transformer-based sarcasm detection model that accounts for the context from the entire conversation thread for more robust predictions. Our model uses deep transformer layers to perform multi-head attentions among the target…

计算与语言 · 计算机科学 2020-05-26 Xiangjue Dong , Changmao Li , Jinho D. Choi

Sarcasm is the use of words usually used to either mock or annoy someone, or for humorous purposes. Sarcasm is largely used in social networks and microblogging websites, where people mock or censure in a way that makes it difficult even…

计算与语言 · 计算机科学 2023-02-07 Alif Tri Handoyo , Hidayaturrahman , Derwin Suhartono

The rapid emergence of multimodal deepfakes (visual and auditory content are manipulated in concert) undermines the reliability of existing detectors that rely solely on modality-specific artifacts or cross-modal inconsistencies. In this…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Yuxuan Du , Zhendong Wang , Yuhao Luo , Caiyong Piao , Zhiyuan Yan , Hao Li , Li Yuan

Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Yue Jiang , Haiwei Xue , Minghao Han , Mingcheng Li , Xiaolu Hou , Dingkang Yang , Lihua Zhang , Xu Zheng

Multimodal sentiment analysis (MSA) draws increasing attention with the availability of multimodal data. The boost in performance of MSA models is mainly hindered by two problems. On the one hand, recent MSA works mostly focus on learning…

机器学习 · 计算机科学 2021-11-17 Ying Zeng , Sijie Mai , Haifeng Hu

In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile. Our analysis reveals that the fundamental limitation lies in an inductive gap, models…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Haoyu Wang , Haonan Wang , Yuyan Chen , Jun Chen , Gang Liu , Qian Wang , Jiahong Yan , Yanghua Xiao

Social media has evolved into a complex multimodal environment where text, images, and other signals interact to shape nuanced meanings, often concealing harmful intent. Identifying such intent, whether sarcasm, hate speech, or…

人工智能 · 计算机科学 2025-09-09 Rui Lu , Jinhe Bi , Yunpu Ma , Feng Xiao , Yuntao Du , Yijun Tian

Human emotions are complex, with sarcasm being a subtle and distinctive form. Despite progress in sarcasm research, sarcasm generation remains underexplored, primarily due to the overreliance on textual modalities and the neglect of visual…

计算与语言 · 计算机科学 2025-07-15 Changli Wang , Rui Wu , Fang Yin

Causal mediation analysis (CMA) is a powerful method to dissect the total effect of a treatment into direct and mediated effects within the potential outcome framework. This is important in many scientific applications to identify the…

机器学习 · 计算机科学 2023-06-14 Ziyang Jiang , Yiling Liu , Michael H. Klein , Ahmed Aloui , Yiman Ren , Keyu Li , Vahid Tarokh , David Carlson

We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand the…

声音 · 计算机科学 2025-11-21 Suhwan Choi , Kyu Won Kim , Myungjoo Kang

Multimodal emotion recognition in conversations (MERC) requires integrating multimodal signals while being robust to noise and modeling contextual reasoning. Existing approaches often emphasize fusion but overlook uncertainty in noisy…

计算与语言 · 计算机科学 2026-04-03 Yiqiang Cai , Chengyan Wu , Bolei Ma , Bo Chen , Yun Xue , Julia Hirschberg , Ziwei Gong

The rapid growth of social media has led to the widespread dissemination of fake news across multiple content forms, including text, images, audio, and video. Traditional unimodal detection methods fall short in addressing complex…

多媒体 · 计算机科学 2025-04-15 Moyang Liu , Kaiying Yan , Yukun Liu , Ruibo Fu , Zhengqi Wen , Xuefei Liu , Chenxing Li

Sarcasm is a term that refers to the use of words to mock, irritate, or amuse someone. It is commonly used on social media. The metaphorical and creative nature of sarcasm presents a significant difficulty for sentiment analysis systems…

计算与语言 · 计算机科学 2022-10-21 Amirhossein Abaskohi , Arash Rasouli , Tanin Zeraati , Behnam Bahrak

Memes are a powerful tool for communication over social media. Their affinity for evolving across politics, history, and sociocultural phenomena makes them an ideal communication vehicle. To comprehend the subtle message conveyed within a…

计算与语言 · 计算机科学 2023-05-30 Shivam Sharma , Ramaneswaran S , Udit Arora , Md. Shad Akhtar , Tanmoy Chakraborty

Suicide remains a pressing global public health concern. While social media platforms offer opportunities for early risk detection through online conversation trees, existing approaches face two major limitations: (1) They rely on…

计算与语言 · 计算机科学 2026-03-02 Jun Li , Xiangmeng Wang , Haoyang Li , Yifei Yan , Shijie Zhang , Hong Va Leong , Ling Feng , Nancy Xiaonan Yu , Qing Li

A large number of annotated video-caption pairs are required for training video captioning models, resulting in high annotation costs. Active learning can be instrumental in reducing these annotation requirements. However, active learning…

计算机视觉与模式识别 · 计算机科学 2022-12-22 Gyanendra Das , Xavier Thomas , Anant Raj , Vikram Gupta

As an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years. However, previous approaches either (i) use separately pre-trained visual and textual models,…

计算机视觉与模式识别 · 计算机科学 2022-04-22 Yan Ling , Jianfei Yu , Rui Xia