English
Related papers

Related papers: Rethinking Multimodal Content Moderation from an A…

200 papers

With the flourishing of social media platforms, vision-language pre-training (VLP) recently has received great attention and many remarkable progresses have been achieved. The success of VLP largely benefits from the information…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Zhiyuan Ma , Jianjun Li , Guohui Li , Kaiyan Huang

Existed pre-training methods either focus on single-modal tasks or multi-modal tasks, and cannot effectively adapt to each other. They can only utilize single-modal data (i.e. text or image) or limited multi-modal data (i.e. image-text…

Computation and Language · Computer Science 2022-03-15 Wei Li , Can Gao , Guocheng Niu , Xinyan Xiao , Hao Liu , Jiachen Liu , Hua Wu , Haifeng Wang

News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances, existing methods struggle with three key challenges: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Xiaoxing You , Qiang Huang , Lingyu Li , Chi Zhang , Xiaopeng Liu , Min Zhang , Jun Yu

Deep multimodal learning has shown remarkable success by leveraging contrastive learning to capture explicit one-to-one relations across modalities. However, real-world data often exhibits shared relations beyond simple pairwise…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Raja Kumar , Raghav Singhal , Pranamya Kulkarni , Deval Mehta , Kshitij Jadhav

Multimodal machine learning has gained significant attention in recent years due to its potential for integrating information from multiple modalities to enhance learning and decision-making processes. However, it is commonly observed that…

Machine Learning · Computer Science 2025-09-12 Sahiti Yerramilli , Jayant Sravan Tamarapalli , Jonathan Francis , Eric Nyberg

Effectively leveraging multimodal data such as various images, laboratory tests and clinical information is gaining traction in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on…

Image and Video Processing · Electrical Eng. & Systems 2023-11-28 Yingying Fang , Shuang Wu , Sheng Zhang , Chaoyan Huang , Tieyong Zeng , Xiaodan Xing , Simon Walsh , Guang Yang

Multimodal recommender systems (MRS) improve recommendation performance by integrating complementary semantic information from multiple modalities. However, the assumption of complete multimodality rarely holds in practice due to missing…

Information Retrieval · Computer Science 2025-10-16 Huilin Chen , Miaomiao Cai , Fan Liu , Zhiyong Cheng , Richang Hong , Meng Wang

Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding. However, their capacity to comprehend the dynamic changes of objects within a shared spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Kewei Wei , Bocheng Hu , Jie Cao , Xiaohan Chen , Zhengxi Lu , Wubing Xia , Weili Xu , Jiaao Wu , Junchen He , Mingyu Jia , Ciyun Zhao , Ye Sun , Yizhi Li , Zhonghan Zhao , Jian Zhang , Gaoang Wang

Recent advancements in multimodal large language models (MLLMs) have demonstrated considerable potential for comprehensive 3D scene understanding. However, existing approaches typically utilize only one or a limited subset of 3D modalities,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Yue Zhang , Yingzhao Jian , Hehe Fan , Yi Yang , Roger Zimmermann

Learning from multiple modalities, such as audio and video, offers opportunities for leveraging complementary information, enhancing robustness, and improving contextual understanding and performance. However, combining such modalities…

Multimedia · Computer Science 2024-10-15 Konstantinos Kontras , Christos Chatzichristos , Matthew Blaschko , Maarten De Vos

We present a new pre-training strategy called M$^{3}$3D ($\underline{M}$ulti-$\underline{M}$odal $\underline{M}$asked $\underline{3D}$) built based on Multi-modal masked autoencoders that can leverage 3D priors and learned cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Muhammad Abdullah Jamal , Omid Mohareri

Misinformation on the web increasingly appears in multimodal forms, combining text, images, and OCR-rendered content in ways that amplify harm to public trust and vulnerable communities. While prior fact-checking systems often rely on…

Computation and Language · Computer Science 2026-01-14 Aditya Kishore , Gaurav Kumar , Jasabanta Patro

With the rapid growth of social media platforms, users are sharing billions of multimedia posts containing audio, images, and text. Researchers have focused on building autonomous systems capable of processing such multimedia data to solve…

Computer Vision and Pattern Recognition · Computer Science 2023-03-13 Muhammad Saad Saeed , Shah Nawaz , Muhammad Haris Khan , Muhammad Zaigham Zaheer , Karthik Nandakumar , Muhammad Haroon Yousaf , Arif Mahmood

In the current context where online platforms have been effectively weaponized in a variety of geo-political events and social issues, Internet memes make fair content moderation at scale even more difficult. Existing work on meme…

Artificial Intelligence · Computer Science 2023-04-10 Abhinav Kumar Thakur , Filip Ilievski , Hông-Ân Sandlin , Zhivar Sourati , Luca Luceri , Riccardo Tommasini , Alain Mermoud

Vision-language models (VLMs) have been proven effective for detecting multi-modal misinformation on social platforms, especially in zero-shot settings with unavailable or delayed annotations. However, a single VLM's capacity falls short in…

Multimedia · Computer Science 2026-03-04 Wei Jiang , Tong Chen , Wei Yuan , Quoc Viet Hung Nguyen , Hongzhi Yin

Cross-modal generalization aims to learn a shared discrete representation space from multimodal pairs, enabling knowledge transfer across unannotated modalities. However, achieving a unified representation for all modality pairs requires…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Yan Xia , Hai Huang , Minghui Fang , Zhou Zhao

Controllable captioning is essential for precise multimodal alignment and instruction following, yet existing models often lack fine-grained control and reliable evaluation protocols. To address this gap, we present the AnyCap Project, an…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Yiming Ren , Zhiqiang Lin , Yu Li , Gao Meng , Weiyun Wang , Junjie Wang , Zicheng Lin , Jifeng Dai , Yujiu Yang , Wenhai Wang , Ruihang Chu

Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by a generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Zhicheng Zhang , Wuyou Xia , Chenxi Zhao , Zhou Yan , Xiaoqiang Liu , Yongjie Zhu , Wenyu Qin , Pengfei Wan , Di Zhang , Jufeng Yang

Learning joint embedding space for various modalities is of vital importance for multimodal fusion. Mainstream modality fusion approaches fail to achieve this goal, leaving a modality gap which heavily affects cross-modal fusion. In this…

Computer Vision and Pattern Recognition · Computer Science 2020-12-11 Sijie Mai , Haifeng Hu , Songlong Xing

Multimodal sentiment analysis has recently gained popularity because of its relevance to social media posts, customer service calls and video blogs. In this paper, we address three aspects of multimodal sentiment analysis; 1. Cross modal…

Computation and Language · Computer Science 2020-03-03 Ayush Kumar , Jithendra Vepa
‹ Prev 1 8 9 10 Next ›