English
Related papers

Related papers: MMBERT: Multimodal BERT Pretraining for Improved M…

200 papers

Mental health is a critical issue in modern society, and mental disorders could sometimes turn to suicidal ideation without adequate treatment. Early detection of mental disorders and suicidal ideation from social content provides a…

Computation and Language · Computer Science 2022-07-19 Shaoxiong Ji , Tianlin Zhang , Luna Ansari , Jie Fu , Prayag Tiwari , Erik Cambria

Medical Visual Question Answering (MedVQA) aims to answer medical questions according to medical images. However, the complexity of medical data leads to confounders that are difficult to observe, so bias between images and questions is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zibo Xu , Qiang Li , Weizhi Nie , Weijie Wang , Anan Liu

Several medical Multimodal Large Languange Models (MLLMs) have been developed to address tasks involving visual images with textual instructions across various medical modalities, achieving impressive results. Most current medical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Lehan Wang , Haonan Wang , Honglong Yang , Jiaji Mao , Zehong Yang , Jun Shen , Xiaomeng Li

Multi-modal pretraining for learning high-level multi-modal representation is a further step towards deep learning and artificial intelligence. In this work, we propose a novel model, namely InterBERT (BERT for Interaction), which is the…

Computation and Language · Computer Science 2021-04-23 Junyang Lin , An Yang , Yichang Zhang , Jie Liu , Jingren Zhou , Hongxia Yang

Visual question answering (VQA) is known as an AI-complete task as it requires understanding, reasoning, and inferring about the vision and the language content. Over the past few years, numerous neural architectures have been suggested for…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Övgü Özdemir , Erdem Akagündüz

While mainstream vision-language models (VLMs) have advanced rapidly in understanding image level information, they still lack the ability to focus on specific areas designated by humans. Rather, they typically rely on large volumes of…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Kangyu Zhu , Ziyuan Qin , Huahui Yi , Zekun Jiang , Qicheng Lao , Shaoting Zhang , Kang Li

The extraction of visual features is an essential step in Visual Question Answering (VQA). Building a good visual representation of the analyzed scene is indeed one of the essential keys for the system to be able to correctly understand the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Hichem Boussaid , Lucrezia Tosato , Flora Weissgerber , Camille Kurtz , Laurent Wendling , Sylvain Lobry

Visual Question Answering (VQA) is a challenging task that requires systems to provide accurate answers to questions based on image content. Current VQA models struggle with complex questions due to limitations in capturing and integrating…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Peiyuan Chen , Zecheng Zhang , Yiping Dong , Li Zhou , Han Wang

With the new generation of satellite technologies, the archives of remote sensing (RS) images are growing very fast. To make the intrinsic information of each RS image easily accessible, visual question answering (VQA) has been introduced…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Tim Siebert , Kai Norman Clasen , Mahdyar Ravanbakhsh , Begüm Demir

Obtaining large pre-trained models that can be fine-tuned to new tasks with limited annotated samples has remained an open challenge for medical imaging data. While pre-trained deep networks on ImageNet and vision-language foundation models…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Duy M. H. Nguyen , Hoang Nguyen , Nghiem T. Diep , Tan N. Pham , Tri Cao , Binh T. Nguyen , Paul Swoboda , Nhat Ho , Shadi Albarqouni , Pengtao Xie , Daniel Sonntag , Mathias Niepert

State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach…

Visual question answering (VQA) is the task of answering questions about an image. The task assumes an understanding of both the image and the question to provide a natural language answer. VQA has gained popularity in recent years due to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-01 Deepanway Ghosal , Navonil Majumder , Roy Ka-Wei Lee , Rada Mihalcea , Soujanya Poria

While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has shown promising cross-domain performance. In this work, we present the first large-scale study…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Leon Mayer , Tim Rädsch , Dominik Michael , Lucas Luttner , Amine Yamlahi , Evangelia Christodoulou , Patrick Godau , Marcel Knopp , Annika Reinke , Fiona Kolbinger , Lena Maier-Hein

Multimodal large language models (MLLMs) have achieved rapid progress, yet their scaling behavior remains less clearly characterized and often less predictable than that of text-only LLMs. Increasing model size and task diversity often…

Computation and Language · Computer Science 2026-04-16 Hongjian Zou , Yue Ge , Qi Ding , Yixuan Liao , Xiaoxin Chen

Existing Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical…

Computation and Language · Computer Science 2025-06-03 Yucheng Zhou , Lingran Song , Jianbing Shen

Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Zhifan Jiang , Dong Yang , Vishwesh Nath , Abhijeet Parida , Nishad P. Kulkarni , Ziyue Xu , Daguang Xu , Syed Muhammad Anwar , Holger R. Roth , Marius George Linguraru

In this paper, we study how to use masked signal modeling in vision and language (V+L) representation learning. Instead of developing masked language modeling (MLM) and masked image modeling (MIM) independently, we propose to build joint…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Gukyeong Kwon , Zhaowei Cai , Avinash Ravichandran , Erhan Bas , Rahul Bhotika , Stefano Soatto

In recent years, significant progress has been made in the field of surgical scene understanding, particularly in the task of Visual Question Localized-Answering in robotic surgery (Surgical-VQLA). However, existing Surgical-VQLA models…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Pengfei Hao , Shuaibo Li , Hongqiu Wang , Zhizhuo Kou , Junhang Zhang , Guang Yang , Lei Zhu

Medical visual question answering (Med-VQA) aims to automate the prediction of correct answers for medical images and questions, thereby assisting physicians in reducing repetitive tasks and alleviating their workload. Existing approaches…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Tiancheng Gu , Kaicheng Yang , Dongnan Liu , Weidong Cai

Acquiring high-quality knowledge is a central focus in Knowledge-Based Visual Question Answering (KB-VQA). Recent methods use large language models (LLMs) as knowledge engines for answering. These methods generally employ image captions as…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Yan Zhang , Jiaqing Lin , Miao Zhang , Kui Xiao , Xiaoju Hou , Yue Zhao , Zhifei Li