English
Related papers

Related papers: DyKen-Hyena: Dynamic Kernel Generation via Cross-M…

200 papers

Multimodal sentiment analysis utilizes multiple heterogeneous modalities for sentiment classification. The recent multimodal fusion schemes customize LSTMs to discover intra-modal dynamics and design sophisticated attention mechanisms to…

Artificial Intelligence · Computer Science 2020-10-19 Sunny Verma , Jiwei Wang , Zhefeng Ge , Rujia Shen , Fan Jin , Yang Wang , Fang Chen , Wei Liu

Multimodal emotion recognition (MER) extracts emotions from multimodal data, including visual, speech, and text inputs, playing a key role in human-computer interaction. Attention-based fusion methods dominate MER research, achieving strong…

Artificial Intelligence · Computer Science 2025-06-03 Jiajun He , Jinyi Mi , Tomoki Toda

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Dingkang Yang , Mingcheng Li , Linhao Qu , Kun Yang , Peng Zhai , Song Wang , Lihua Zhang

Multi behavior recommendation leverages multiple types of user-item interactions to address data sparsity and cold-start issues,providing personalized services in domains such as healthcare and ecommerce.Most existing methods utilize graph…

Information Retrieval · Computer Science 2025-12-02 Ruiqi Luo , Ran Jin , Kaixi Hu , Xiaohui Tao , Lin Li

Recent advances in multimodal AI have enabled progress in detecting synthetic and out-of-context content. However, existing efforts largely overlook the intent behind AI-generated images. To fill this gap, we introduce S-HArM, a multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Anastasios Skoularikis , Stefanos-Iordanis Papadopoulos , Symeon Papadopoulos , Panagiotis C. Petrantonakis

Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully leveraging these complementary sources of information requires…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Inseok Jeon , Suhwan Cho , Minhyeok Lee , Seunghoon Lee , Minseok Kang , Jungho Lee , Chaewon Park , Donghyeong Kim , Sangyoun Lee

We address the problem of referring image segmentation that aims to generate a mask for the object specified by a natural language expression. Many recent works utilize Transformer to extract features for the target object by aggregating…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Chang Liu , Henghui Ding , Yulun Zhang , Xudong Jiang

Existing periodic activation-based implicit neural representation (INR) networks, such as SIREN and FINER, suffer from hidden feature redundancy, where neurons within a layer capture overlapping frequency components due to the use of a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Mohammed Alsakabi , Wael Mobeirek , John M. Dolan , Ozan K. Tonguz

Multimodal emotion recognition in conversation (MERC) requires representations that effectively integrate signals from multiple modalities. These signals include modality-specific cues, information shared across modalities, and interactions…

Machine Learning · Computer Science 2026-01-22 Anh-Tuan Mai , Cam-Van Thi Nguyen , Duc-Trong Le

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Kaihang Pan , Wang Lin , Zhongqi Yue , Tenglong Ao , Liyu Jia , Wei Zhao , Juncheng Li , Siliang Tang , Hanwang Zhang

The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In current AV-LLMs, audio and video features are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Chaeyoung Jung , Youngjoon Jang , Jongmin Choi , Joon Son Chung

Contrastive learning has proven effective in training sequential recommendation models by incorporating self-supervised signals from augmented views. Most existing methods generate multiple views from the same interaction sequence through…

Information Retrieval · Computer Science 2025-04-24 Yuanpeng Qu , Hajime Nobuhara

Multimodal emotion and intent recognition is essential for automated human-computer interaction, It aims to analyze users' speech, text, and visual information to predict their emotions or intent. One of the significant challenges is that…

Artificial Intelligence · Computer Science 2025-07-09 Wei Zhang , Juan Chen , Yanbo J. Wang , En Zhu , Xuan Yang , Yiduo Wang

Conventional machine learning methods are predominantly designed to predict outcomes based on a single data type. However, practical applications may encompass data of diverse types, such as text, images, and audio. We introduce…

Effectively leveraging multimodal data such as various images, laboratory tests and clinical information is gaining traction in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on…

Image and Video Processing · Electrical Eng. & Systems 2023-11-28 Yingying Fang , Shuang Wu , Sheng Zhang , Chaoyan Huang , Tieyong Zeng , Xiaodan Xing , Simon Walsh , Guang Yang

Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Shuhong Zheng , Zhipeng Bao , Ruoyu Zhao , Martial Hebert , Yu-Xiong Wang

Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Zhiwei Chen , Yupeng Hu , Zhiheng Fu , Zixu Li , Jiale Huang , Qinlei Huang , Yinwei Wei

Emotion Recognition in Conversations (ERC) has considerable prospects for developing empathetic machines. For multimodal ERC, it is vital to understand context and fuse modality information in conversations. Recent graph-based fusion…

Computation and Language · Computer Science 2022-03-07 Dou Hu , Xiaolong Hou , Lingwei Wei , Lianxin Jiang , Yang Mo

Understanding sentiment in complex textual expressions remains a fundamental challenge in affective computing. To address this, we propose a Dynamic Fusion Learning Model (DyFuLM), a multimodal framework designed to capture both…

Computation and Language · Computer Science 2025-12-02 Ruohan Zhou , Jiachen Yuan , Churui Yang , Wenzheng Huang , Guoyan Zhang , Shiyao Wei , Jiazhen Hu , Ning Xin , Md Maruf Hasan

Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge that lies in inductive biases induced by distinct supervision…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Renjie Lu , Xulong Zhang , Xiaoyang Qu , Shangfei Wang , Jianzong Wang