中文
相关论文

相关论文: Multimodal Transformers are Hierarchical Modal-wis…

200 篇论文

Understanding structure-property relationships in complex materials requires integrating complementary measurements across multiple length scales. Here we propose an interpretable "multimodal" machine learning framework that unifies…

材料科学 · 物理学 2026-02-03 Shun Muroga , Hideaki Nakajima , Taiyo Shimizu , Kazufumi Kobashi , Kenji Hata

Recently, table structure recognition has achieved impressive progress with the help of deep graph models. Most of them exploit single visual cues of tabular elements or simply combine visual cues with other modalities via early fusion to…

计算机视觉与模式识别 · 计算机科学 2022-03-11 Hao Liu , Xin Li , Bing Liu , Deqiang Jiang , Yinsong Liu , Bo Ren

Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Bingyu Li , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

Whole slide image (WSI) classification is an essential task in computational pathology. Despite the recent advances in multiple instance learning (MIL) for WSI classification, accurate classification of WSIs remains challenging due to the…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Saisai Ding , Jun Wang , Juncheng Li , Jun Shi

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Wangyuan Zhu , Jun Yu

Multimodal Sentiment Analysis (MSA) endeavors to understand human sentiment by leveraging language, visual, and acoustic modalities. Despite the remarkable performance exhibited by previous MSA approaches, the presence of inherent…

多媒体 · 计算机科学 2025-05-09 Weize Quan , Yunfei Feng , Ming Zhou , Yunzhen Zhao , Tong Wang , Dong-Ming Yan

Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data…

Multimodal Sentiment Analysis (MSA) integrates diverse modalities(text, audio, and video) to comprehensively analyze and understand individuals' emotional states. However, the real-world prevalence of incomplete data poses significant…

计算与语言 · 计算机科学 2025-01-13 Xincheng Wang , Liejun Wang , Yinfeng Yu , Xinxin Jiao

Multiple Instance Learning (MIL) and transformers are increasingly popular in histopathology Whole Slide Image (WSI) classification. However, unlike human pathologists who selectively observe specific regions of histopathology tissues under…

计算机视觉与模式识别 · 计算机科学 2023-07-18 Conghao Xiong , Hao Chen , Joseph J. Y. Sung , Irwin King

Benefiting from the powerful expressive capability of graphs, graph-based approaches have achieved impressive performance in various biomedical applications. Most existing methods tend to define the adjacency matrix among samples manually…

机器学习 · 计算机科学 2021-07-02 Shuai Zheng , Zhenfeng Zhu , Zhizhe Liu , Zhenyu Guo , Yang Liu , Yao Zhao

Recent research in time series forecasting has explored integrating multimodal features into models to improve accuracy. However, the accuracy of such methods is constrained by three key challenges: inadequate extraction of fine-grained…

机器学习 · 计算机科学 2025-10-21 Shule Hao , Junpeng Bao , Wenli Li

This paper investigates the optimal selection and fusion of feature encoders across multiple modalities and combines these in one neural network to improve sentiment detection. We compare different fusion methods and examine the impact of…

计算与语言 · 计算机科学 2024-06-04 Zehui Wu , Ziwei Gong , Jaywon Koo , Julia Hirschberg

In digital pathology, the multiple instance learning (MIL) strategy is widely used in the weakly supervised histopathology whole slide image (WSI) classification task where giga-pixel WSIs are only labeled at the slide level. However,…

图像与视频处理 · 电气工程与系统科学 2024-03-28 Zhan Shi , Jingwei Zhang , Jun Kong , Fusheng Wang

Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and visual data. However, existing methods often suffer from spurious correlations both within…

机器学习 · 计算机科学 2026-05-21 Menghua Jiang , Yuxia Lin , Baoliang Chen , Haifeng Hu , Yuncheng Jiang , Sijie Mai

Multimodal Sentiment Analysis (MSA) utilizes multimodal data to infer the users' sentiment. Previous methods focus on equally treating the contribution of each modality or statically using text as the dominant modality to conduct…

计算与语言 · 计算机科学 2024-10-08 Xinyu Feng , Yuming Lin , Lihua He , You Li , Liang Chang , Ya Zhou

Forecasting high-resolution land subsidence is a critical yet challenging task due to its complex, non-linear dynamics. While standard architectures like ConvLSTM often fail to model long-range dependencies, we argue that a more fundamental…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Wendong Yao , Binhua Huang , Soumyabrata Dev

Multivariate time series analysis has long been one of the key research topics in the field of artificial intelligence. However, analyzing complex time series data remains a challenging and unresolved problem due to its high dimensionality,…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Hao Si , Xiao Wang , Fan Zhang , Xiaoya Zhou , Dengdi Sun , Wanli Lyu , Qingquan Yang , Jin Tang

Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure). By integrating RGB with modalities such as thermal and depth, multi-modal fusion increases…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Xiaofan Yang , Yubin Liu , Wei Pan , Guoqing Chu , Junming Zhang , Jie Zhao , Zhuoqi Man , Xuanming Cao

Cancer survival prediction requires integrating pathological Whole Slide Images (WSIs) and genomic profiles, a challenging task due to the inherent heterogeneity and the complexity of modeling both inter- and intra-modality interactions.…

图像与视频处理 · 电气工程与系统科学 2025-07-08 Mingxin Liu , Chengfei Cai , Jun Li , Pengbo Xu , Jinze Li , Jiquan Ma , Jun Xu

Since the introduction of the Transformer architecture for large language models, the softmax-based attention layer has faced increasing scrutinity due to its quadratic-time computational complexity. Attempts have been made to replace it…

机器学习 · 计算机科学 2026-02-02 Robert Forchheimer