English
Related papers

Related papers: GOMA: Toward Structure-Driven Multimodal Alignment…

200 papers

Logistics optimization nowadays is becoming one of the hottest areas in the AI community. In the past year, significant advancements in the domain were achieved by representing the problem in a form of graph. Another promising area of…

Machine Learning · Computer Science 2022-05-26 Zangir Iklassov , Dmitrii Medvedev

Multimodal recommendation systems have attracted increasing attention for their improved performance by leveraging items' multimodal information. Prior methods often build modality-specific item-item semantic graphs from raw modality…

Information Retrieval · Computer Science 2025-08-11 Xiaoxiong Zhang , Xin Zhou , Zhiwei Zeng , Dusit Niyato , Zhiqi Shen

Benefiting from the powerful expressive capability of graphs, graph-based approaches have achieved impressive performance in various biomedical applications. Most existing methods tend to define the adjacency matrix among samples manually…

Machine Learning · Computer Science 2021-07-02 Shuai Zheng , Zhenfeng Zhu , Zhizhe Liu , Zhenyu Guo , Yang Liu , Yao Zhao

In this paper, we introduce an Adaptive Graph Signal Processing with Dynamic Semantic Alignment (AGSP DSA) framework to perform robust multimodal data fusion over heterogeneous sources, including text, audio, and images. The requested…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 KV Karthikeya , Ashok Kumar Das , Shantanu Pal , Vivekananda Bhat K , Arun Sekar Rajasekaran

We introduce CLARGA, a general-purpose multimodal fusion architecture for multimodal representation learning that works with any number and type of modalities without changing the underlying framework. Given a supervised dataset, CLARGA can…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Santosh Patapati

Recent urbanization has coincided with the enrichment of geotagged data, such as street view and point-of-interest (POI). Region embedding enhanced by the richer data modalities has enabled researchers and city administrators to understand…

Machine Learning · Computer Science 2021-05-07 Tianyuan Huang , Zhecheng Wang , Hao Sheng , Andrew Y. Ng , Ram Rajagopal

Human perception integrates multiple modalities, such as vision, hearing, and language, into a unified understanding of the surrounding reality. While recent multimodal models have achieved significant progress by aligning pairs of…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Giordano Cicchetti , Eleonora Grassucci , Luigi Sigillo , Danilo Comminiello

In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing these features is essential for robust performance,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Dexuan Ding , Lei Wang , Liyun Zhu , Tom Gedeon , Piotr Koniusz

Spatial multi-modal omics technology, highlighted by Nature Methods as an advanced biological technique in 2023, plays a critical role in resolving biological regulatory processes with spatial context. Recently, graph neural networks based…

Genomics · Quantitative Biology 2024-12-19 Xinlei Huang , Zhiqi Ma , Dian Meng , Yanran Liu , Shiwei Ruan , Qingqiang Sun , Xubin Zheng , Ziyue Qiao

Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal reasoning, they typically treat emotion categories as…

Machine Learning · Computer Science 2026-05-20 Zeheng Wang , Bo Zhao , Yijie Zhu , Zhishu Liu , Hui Ma , Ruixin Zhang , Shouhong Ding , Qianyu Xie , Zitong Yu

Audiovisual data is everywhere in this digital age, which raises higher requirements for the deep learning models developed on them. To well handle the information of the multi-modal data is the key to a better audiovisual modal. We observe…

Sound · Computer Science 2023-09-27 Meng Liu , Ke Liang , Dayu Hu , Hao Yu , Yue Liu , Lingyuan Meng , Wenxuan Tu , Sihang Zhou , Xinwang Liu

Pretrained unimodal encoders incorporate rich semantic information into embedding space structures. To be similarly informative, multi-modal encoders typically require massive amounts of paired data for alignment and training. We introduce…

Machine Learning · Computer Science 2023-10-10 Dustin Klebe , Tal Shnitzer , Mikhail Yurochkin , Leonid Karlinsky , Justin Solomon

Development of multimodal interactive systems is hindered by the lack of rich, multimodal (text, images) conversational data, which is needed in large quantities for LLMs. Previous approaches augment textual dialogues with retrieved images,…

Computation and Language · Computer Science 2024-10-04 Hossein Aboutalebi , Hwanjun Song , Yusheng Xie , Arshit Gupta , Justin Sun , Hang Su , Igor Shalyminov , Nikolaos Pappas , Siffi Singh , Saab Mansour

Retrieval-Augmented Generation (RAG) has emerged as the dominant paradigm for grounding large language model outputs in verifiable evidence. However, as modern AI agents transition from static knowledge bases to continuous multimodal…

Machine Learning · Computer Science 2025-11-05 Rohan Wandre , Yash Gajewar , Namrata Patel , Vivek Dhalkari

Recently, multimodal graph learning (MGL) has garnered significant attention for integrating diverse modality information and structured context to support various network applications. However, real-world graphs are often isolated due to…

Machine Learning · Computer Science 2026-05-14 Sirui Zhang , Haonan Wang , Xunkai Li , Zekai Chen , Shumeng Li , Hongchao Qin , Rong-Hua Li , Guoren Wang

The rapid development of Multimodal Large Language Models (MLLMs) has enabled the integration of multiple modalities, including texts and images, within the large language model (LLM) framework. However, texts and images are usually…

Artificial Intelligence · Computer Science 2025-03-11 Yi Fang , Bowen Jin , Jiacheng Shen , Sirui Ding , Qiaoyu Tan , Jiawei Han

Graph Anomaly Detection (GAD) aims to identify atypical graph entities, such as nodes, edges, or substructures, that deviate significantly from the majority. While existing text-rich approaches typically integrate structural context into…

Computation and Language · Computer Science 2026-05-20 Wen Shi , Zhe Wang , Huafei Huang , Qing Qing , Ziqi Xu , Qixin Zhang , Xikun Zhang , Renqiang Luo , Feng Xia

Semantics has enabled 3D scene understanding and affordance-driven object interaction. However, robots operating in real-world environments face a critical limitation: they cannot anticipate how objects move. Long-horizon mobile…

In multimodal graph learning, graph structures that integrate information from multiple sources, such as vision and text, can more comprehensively model complex entity relationships. However, the continuous growth of their data scale poses…

Machine Learning · Computer Science 2026-02-10 Lian Shen , Zhendan Chen , Meijia Song , Yinhui jiang , Ziming Su , Juan Liu , Xiangrong Liu

A Tone Mapping Operator (TMO) is required to render images with a High Dynamic Range (HDR) on media with limited dynamic capabilities. TMOs compress the dynamic range with the aim of preserving the visually perceptual cues of the scene.…

Image and Video Processing · Electrical Eng. & Systems 2024-11-11 Abhishek Goswami , Erwan Bernard , Wolf Hauser , Frederic Dufaux , Rafal Mantiuk