中文
相关论文

相关论文: FlexMUSE: Multimodal Unification and Semantics Enh…

200 篇论文

Modality fusion is a cornerstone of multimodal learning, enabling information integration from diverse data sources. However, vanilla fusion methods are limited by (1) inability to account for heterogeneous interactions between modalities…

机器学习 · 计算机科学 2025-05-27 Jiayi Xin , Sukwon Yun , Jie Peng , Inyoung Choi , Jenna L. Ballard , Tianlong Chen , Qi Long

In recent years, there has been significant progress in semantic communication systems empowered by deep learning techniques. It has greatly improved the efficiency of information transmission. Nevertheless, traditional semantic…

信号处理 · 电气工程与系统科学 2025-12-01 Zengle Zhu , Rongqing Zhang , Xiang Cheng , Liuqing Yang

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Jiadong Pan , Liang Li , Yuxin Peng , Yu-Ming Tang , Shuohuan Wang , Yu Sun , Hua Wu , Qingming Huang , Haifeng Wang

Sentiment analysis and emotion recognition in videos are challenging tasks, given the diversity and complexity of the information conveyed in different modalities. Developing a highly competent framework that effectively addresses the…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Prasad Chaudhari , Aman Kumar , Chandravardhan Singh Raghaw , Mohammad Zia Ur Rehman , Nagendra Kumar

Learning generative models that span multiple data modalities, such as vision and language, is often motivated by the desire to learn more useful, generalisable representations that faithfully capture common underlying factors between the…

机器学习 · 统计学 2019-11-11 Yuge Shi , N. Siddharth , Brooks Paige , Philip H. S. Torr

Multimodal clinical prediction faces three challenges: multiple foundation models (FMs) with complementary strengths per modality, pervasive missing modalities at training and test time, and sample-specific variation in modality…

机器学习 · 计算机科学 2026-05-19 Seungik Cho , Anqi Li , Wei Qiu

Multimodal Sarcasm Explanation (MuSE) is a new yet challenging task, which aims to generate a natural language sentence for a multimodal social post (an image as well as its caption) to explain why it contains sarcasm. Although the existing…

计算与语言 · 计算机科学 2023-06-30 Liqiang Jing , Xuemeng Song , Kun Ouyang , Mengzhao Jia , Liqiang Nie

Despite rapid progress in multimodal large language models (MLLMs) and emerging omni-modal architectures, current benchmarks remain limited in scope and integration, suffering from incomplete modality coverage, restricted interaction to…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Yue Jiang , Dingkang Yang , Minghao Han , Jinghang Han , Zizhi Chen , Yizhou Liu , Mingcheng Li , Peng Zhai , Lihua Zhang

Multimodal sentiment analysis in videos is a key task in many real-world applications, which usually requires integrating multimodal streams including visual, verbal and acoustic behaviors. To improve the robustness of multimodal fusion,…

计算机视觉与模式识别 · 计算机科学 2022-06-20 Lianyang Ma , Yu Yao , Tao Liang , Tongliang Liu

Multimodal stock trading volume movement prediction with stock-related news is one of the fundamental problems in the financial area. Existing multimodal works that train models from scratch face the problem of lacking universal knowledge…

计算与语言 · 计算机科学 2023-09-12 Ruibo Chen , Zhiyuan Zhang , Yi Liu , Ruihan Bao , Keiko Harimoto , Xu Sun

Due to the severe lack of labeled data, existing methods of medical visual question answering usually rely on transfer learning to obtain effective image feature representation and use cross-modal fusion of visual and linguistic features to…

多媒体 · 计算机科学 2021-05-04 Haifan Gong , Guanqi Chen , Sishuo Liu , Yizhou Yu , Guanbin Li

Multimodal Large Language Models (MLLMs) have advanced in integrating diverse modalities but frequently suffer from hallucination. A promising solution to mitigate this issue is to generate text with citations, providing a transparent chain…

计算与语言 · 计算机科学 2025-05-21 Caiyu Hu , Yikai Zhang , Tinghui Zhu , Yiwei Ye , Yanghua Xiao

In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Zhiqi Ai , Zhiyong Chen , Shugong Xu

Multi-modal knowledge graph completion (MMKGC) aims to discover missing facts in multi-modal knowledge graphs (MMKGs) by leveraging both structural relationships and diverse modality information of entities. Existing MMKGC methods follow…

计算与语言 · 计算机科学 2026-04-20 Zhiqiang Liu , Yichi Zhang , Mengshu Sun , Lei Liang , Wen Zhang

Creating a meaningful representation by fusing single modalities (e.g., text, images, or audio) is the core concept of multimodal learning. Although several techniques for building multimodal representations have been proven successful,…

机器学习 · 计算机科学 2025-08-08 Maciej Pawłowski , Anna Wróblewska , Sylwia Sysko-Romańczuk

We introduce a new task, MultiMedia Event Extraction (M2E2), which aims to extract events and their arguments from multimedia documents. We develop the first benchmark and collect a dataset of 245 multimedia news articles with extensively…

多媒体 · 计算机科学 2020-05-07 Manling Li , Alireza Zareian , Qi Zeng , Spencer Whitehead , Di Lu , Heng Ji , Shih-Fu Chang

Humans perceive the world through multimodal cues to understand and interact with the environment. Learning a scene representation for multiple modalities enhances comprehension of the physical world. However, modality conflicts, arising…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Zhifeng Gu , Bing Wang

Multi-Modal Entity Alignment (MMEA) aims to retrieve equivalent entities from different Multi-Modal Knowledge Graphs (MMKGs), a critical information retrieval task. Existing studies have explored various fusion paradigms and consistency…

多媒体 · 计算机科学 2025-05-16 Taoyu Su , Jiawei Sheng , Duohe Ma , Xiaodong Li , Juwei Yue , Mengxiao Song , Yingkai Tang , Tingwen Liu

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g., depth, edges) are insufficient: they specify layout but…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Dianbing Xi , Jiepeng Wang , Yuanzhi Liang , Xi Qiu , Jialun Liu , Hao Pan , Yuchi Huo , Rui Wang , Haibin Huang , Chi Zhang , Xuelong Li