English
Related papers

Related papers: Context-Gated Cross-Modal Perception with Visual M…

200 papers

Medical visual question answering (Med-VQA) is a crucial multimodal task in clinical decision support and telemedicine. Recent methods fail to fully leverage domain-specific medical knowledge, making it difficult to accurately associate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xianyao Zheng , Hong Yu , Hui Cui , Changming Sun , Xiangyu Li , Ran Su , Leyi Wei , Jia Zhou , Junbo Wang , Qiangguo Jin

Multimodal semantic segmentation has emerged as a powerful paradigm for enhancing scene understanding by leveraging complementary information from multiple sensing modalities (e.g., RGB, depth, and thermal). However, existing cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Guoan Xu , Yang Xiao , Guangwei Gao , Dongchen Zhu , Guo-Jun Qi , Wenjing Jia

Multi-modal brain tumor segmentation remains challenging for practical deployment due to the high computational costs of mainstream models. In this work, we propose GMLN-BTS, a Graph-based Multi-modal interaction Lightweight Network for…

Image and Video Processing · Electrical Eng. & Systems 2026-03-06 Guohao Huo , Ruiting Dai , Zitong Wang , Junxin Kong , Hao Tang

Recent studies suggest that Visual Language Models (VLMs) hold great potential for tasks such as automated medical diagnosis. However, processing complex three-dimensional (3D) multimodal medical images poses significant challenges -…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Hao Wu , Hui Li , Yiyun Su

In data-scarce scenarios, deep learning models often overfit to noise and irrelevant patterns, which limits their ability to generalize to unseen samples. To address these challenges in medical image segmentation, we introduce Diff-UMamba,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Dhruv Jain , Romain Modzelewski , Romain Herault , Clement Chatelain , Eva Torfeh , Sebastien Thureau

Recent advancements in State Space Models, notably Mamba, have demonstrated superior performance over the dominant Transformer models, particularly in reducing the computational complexity from quadratic to linear. Yet, difficulties in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Fei Xie , Weijia Zhang , Zhongdao Wang , Chao Ma

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Sheng Lu , Hao Chen , Rui Yin , Juyan Ba , Yu Zhang , Yuanzhe Li

EEG-based emotion recognition struggles with capturing multi-scale spatiotemporal dynamics and ensuring computational efficiency for real-time applications. Existing methods often oversimplify temporal granularity and spatial hierarchies,…

Signal Processing · Electrical Eng. & Systems 2025-07-23 Hanwen Liu , Yifeng Gong , Zuwei Yan , Zeheng Zhuang , Jiaxuan Lu

Cancer detection and prognosis relies heavily on medical imaging, particularly CT and PET scans. Deep Neural Networks (DNNs) have shown promise in tumor segmentation by fusing information from these modalities. However, a critical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Numan Saeed , Shahad Hardan , Muhammad Ridzuan , Nada Saadi , Karthik Nandakumar , Mohammad Yaqub

Skin lesion segmentation is a critical task in computer-aided diagnosis systems for dermatological diseases. Accurate segmentation of skin lesions from medical images is essential for early detection, diagnosis, and treatment planning. In…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Viet-Thanh Nguyen , Van-Truong Pham , Thi-Thao Tran

Answering semantically-complicated questions according to an image is challenging in Visual Question Answering (VQA) task. Although the image can be well represented by deep learning, the question is always simply embedded and cannot well…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 JianJian Cao , Xiameng Qin , Sanyuan Zhao , Jianbing Shen

Point cloud segmentation is an important topic in 3D understanding that has traditionally has been tackled using either the CNN or Transformer. Recently, Mamba has emerged as a promising alternative, offering efficient long-range contextual…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yong Xien Chng , Xuchong Qiu , Yizeng Han , Yifan Pu , Jiewei Cao , Gao Huang

Recently, the Mamba architecture based on state space models has demonstrated remarkable performance in a series of natural language processing tasks and has been rapidly applied to remote sensing change detection (CD) tasks. However, most…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Haotian Zhang , Keyan Chen , Chenyang Liu , Hao Chen , Zhengxia Zou , Zhenwei Shi

Unmanned Aerial Vehicle (UAV) remote sensing, with its advantages of rapid information acquisition and low cost, has been widely applied in scenarios such as emergency response. However, due to the long imaging distance and complex imaging…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Kejun Ren , Xin Wu , Lianming Xu , Li Wang

Transformers have become foundational for visual tasks such as object detection, semantic segmentation, and video understanding, but their quadratic complexity in attention mechanisms presents scalability challenges. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Fady Ibrahim , Guangjun Liu , Guanghui Wang

\textit{Objectives}: Data scarcity and domain shifts lead to biased training sets that do not accurately represent deployment conditions. A related practical problem is cross-modal image segmentation, where the objective is to segment…

Image and Video Processing · Electrical Eng. & Systems 2024-04-01 Guillaume Sallé , Pierre-Henri Conze , Julien Bert , Nicolas Boussion , Dimitris Visvikis , Vincent Jaouen

Medicine is inherently a multimodal discipline. Medical images can reflect the pathological changes of cancer and tumors, while the expression of specific genes can influence their morphological characteristics. However, most deep learning…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Jiaying Zhou , Mingzhou Jiang , Junde Wu , Jiayuan Zhu , Ziyue Wang , Yueming Jin

Robot grasping, whether handling isolated objects, cluttered items, or stacked objects, plays a critical role in industrial and service applications. However, current visual grasp detection methods based on Convolutional Neural Networks…

Robotics · Computer Science 2025-03-11 Songsong Xiong , Hamidreza Kasaei

Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Hui Liu , Chen Jia , Fan Shi , Xu Cheng , Mengfei Shi , Xia Xie , Shengyong Chen

Imaging modalities such as Computed Tomography (CT) and Positron Emission Tomography (PET) are key in cancer detection, inspiring Deep Neural Networks (DNN) models that merge these scans for tumor segmentation. When both CT and PET scans…

Image and Video Processing · Electrical Eng. & Systems 2024-04-23 Nada Saadi , Numan Saeed , Mohammad Yaqub , Karthik Nandakumar