English
Related papers

Related papers: Unified Supervision For Vision-Language Modeling i…

200 papers

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However,…

Computation and Language · Computer Science 2025-09-19 Pengyu Wang , Shaojun Zhou , Chenkun Tan , Xinghao Wang , Wei Huang , Zhen Ye , Zhaowei Li , Botian Jiang , Dong Zhang , Xipeng Qiu

Handwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and variability in handwriting styles. Prior methods have faced…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yu Li , Jin Jiang , Jianhua Zhu , Shuai Peng , Baole Wei , Yuxuan Zhou , Liangcai Gao

Accurate segmentation of regions of interest in biomedical images holds substantial value in image analysis. Although several foundation models for biomedical segmentation have currently achieved excellent performance on certain datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Manyu Li , Ruian He , Zixian Zhang , Chenxi Ma , Weimin Tan , Bo Yan

Innovations in digital intelligence are transforming robotic surgery with more informed decision-making. Real-time awareness of surgical instrument presence and actions (e.g., cutting tissue) is essential for such systems. Yet, despite…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Jiajun Cheng , Xianwu Zhao , Sainan Liu , Xiaofan Yu , Ravi Prakash , Patrick J. Codd , Jonathan Elliott Katz , Shan Lin

While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge when training on 3D-CT volumetric data. In this study, we propose TotalFM, a radiological…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Kohei Yamamoto , Tomohiro Kikuchi

While making a tremendous impact in various fields, deep neural networks usually require large amounts of labeled data for training which are expensive to collect in many applications, especially in the medical domain. Unlabeled data, on…

Computer Vision and Pattern Recognition · Computer Science 2020-02-25 Yingda Xia , Fengze Liu , Dong Yang , Jinzheng Cai , Lequan Yu , Zhuotun Zhu , Daguang Xu , Alan Yuille , Holger Roth

Vision-language pre-training, i.e., aligning images with paired text, is a powerful paradigm to create encoders that can be directly used for tasks such as classification, retrieval, and segmentation. In the 3D medical image domain, these…

Vision-Language Models (VLMs) have emerged as the dominant approach for zero-shot recognition, adept at handling diverse scenarios and significant distribution changes. However, their deployment in risk-sensitive areas requires a deeper…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Weijie Tu , Weijian Deng , Dylan Campbell , Stephen Gould , Tom Gedeon

Recently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. However, the prevailing approaches primarily regard the visual…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Yang Jin , Kun Xu , Kun Xu , Liwei Chen , Chao Liao , Jianchao Tan , Quzhe Huang , Bin Chen , Chenyi Lei , An Liu , Chengru Song , Xiaoqiang Lei , Di Zhang , Wenwu Ou , Kun Gai , Yadong Mu

Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Ziyang Zhang , Yang Yu , Yucheng Chen , Xulei Yang , Si Yong Yeo

Significant research efforts have been made to scale and improve vision-language model (VLM) training approaches. Yet, with an ever-growing number of benchmarks, researchers are tasked with the heavy burden of implementing each protocol,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Haider Al-Tahan , Quentin Garrido , Randall Balestriero , Diane Bouchacourt , Caner Hazirbas , Mark Ibrahim

Unified vision-language frameworks have greatly advanced in recent years, most of which adopt an encoder-decoder architecture to unify image-text tasks as sequence-to-sequence generation. However, existing video-language (VidL) models still…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Linjie Li , Zhe Gan , Kevin Lin , Chung-Ching Lin , Zicheng Liu , Ce Liu , Lijuan Wang

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

In medical image classification, supervised learning is challenging due to the scarcity of labeled medical images. To address this, we leverage the visual-textual alignment within Vision-Language Models (VLMs) to enable unsupervised…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Umaima Rahman , Raza Imam , Mohammad Yaqub , Boulbaba Ben Amor , Dwarikanath Mahapatra

Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Yiyan Xu , Qiulin Wang , Wenjie Wang , Yunyao Mao , Xintao Wang , Pengfei Wan , Kun Gai , Fuli Feng

Latent diffusion models (LDM) have revolutionized text-to-image generation, leading to the proliferation of various advanced models and diverse downstream applications. However, despite these significant advancements, current diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Jiacheng Zhang , Jie Wu , Yuxi Ren , Xin Xia , Huafeng Kuang , Pan Xie , Jiashi Li , Xuefeng Xiao , Weilin Huang , Shilei Wen , Lean Fu , Guanbin Li

Recent advancements in vision-language pre-training via contrastive learning have significantly improved performance across computer vision tasks. However, in the medical domain, obtaining multimodal data is often costly and challenging due…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Ameera Bawazir , Kebin Wu , Wenbin Li

With the rapid advancement of deep learning, particularly in the field of medical image analysis, an increasing number of Vision-Language Models (VLMs) are being widely applied to solve complex health and biomedical challenges. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Haiyang Yu , Siyang Yi , Ke Niu , Minghan Zhuo , Bin Li

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Yang Xing , Jiong Wu , Savas Ozdemir , Ying Zhang , Yang Yang , Wei Shao , Kuang Gong

Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains largely underexplored due to the lack of fine-grained, annotated datasets and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Meng-Xun Li , Wen-Hui Deng , Zhi-Xing Wu , Chun-Xiao Jin , Jia-Min Wu , Yue Han , James Kit Hon Tsoi , Gui-Song Xia , Cui Huang