English
Related papers

Related papers: Florence-2: Advancing a Unified Representation for…

200 papers

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on…

We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2, a generative vision foundation model. Unlike the widely used CLIP-style vision transformer trained…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jiuhai Chen , Jianwei Yang , Haiping Wu , Dianqi Li , Jianfeng Gao , Tianyi Zhou , Bin Xiao

Vision-Language Models (VLMs) have emerged as powerful tools in artificial intelli-gence, capable of integrating textual and visual data for a unified understanding of complex scenes. While models such as Florence2, built on transformer…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Aysegul Ucar , Soumyadeep Ro , Sanapala Satwika , Pamarthi Yasoda Gayathri , Mohmmad Ghaith Balsha

Large Language Models (LLMs) have demonstrated exceptional proficiency in text understanding and embedding tasks. However, their potential in multimodal representation, particularly for item-to-item (I2I) recommendations, remains…

Information Retrieval · Computer Science 2025-01-22 Chao Zhang , Haoxin Zhang , Shiwei Wu , Di Wu , Tong Xu , Xiangyu Zhao , Yan Gao , Yao Hu , Enhong Chen

Geometric Dimensioning and Tolerancing (GD&T) plays a critical role in manufacturing by defining acceptable variations in part features to ensure component quality and functionality. However, extracting GD&T information from 2D engineering…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Muhammad Tayyab Khan , Lequn Chen , Ye Han Ng , Wenhe Feng , Nicholas Yew Jin Tan , Seung Ki Moon

We introduce Falcon2-11B, a foundation model trained on over five trillion tokens, and its multimodal counterpart, Falcon2-11B-vlm, which is a vision-to-text model. We report our findings during the training of the Falcon2-11B which follows…

We present Fashion Florence, a Florence-2 vision-language model fine-tuned with LoRA to extract structured fashion attributes from clothing images. Given a single photograph, the model generates a JSON object containing category, color,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Anushree Berlia

Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to eliminating this inconsistency is to use generalist models for…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Hao Li , Jinguo Zhu , Xiaohu Jiang , Xizhou Zhu , Hongsheng Li , Chun Yuan , Xiaohua Wang , Yu Qiao , Xiaogang Wang , Wenhai Wang , Jifeng Dai

We present a multi-task framework for the MediaEval Medico 2025 challenge, leveraging a LoRA-tuned Florence-2 model for simultaneous visual question answering (VQA), explanation generation, and visual grounding. The proposed system…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Itbaan Safwan , Muhammad Annas Shaikh , Muhammad Haaris , Ramail Khan , Muhammad Atif Tahir

Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models' decisions has grown…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Ruoyu Chen , Siyuan Liang , Jingzhi Li , Shiming Liu , Maosen Li , Zhen Huang , Hua Zhang , Xiaochun Cao

Vision-Language Models (VLMs) lag behind Large Language Models due to the scarcity of annotated datasets, as creating paired visual-textual annotations is labor-intensive and expensive. To address this bottleneck, we introduce SAM2Auto, the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Arash Rocky , Q. M. Jonathan Wu

Fluorescent Neuronal Cells v2 is a collection of fluorescence microscopy images and the corresponding ground-truth annotations, designed to foster innovative research in the domains of Life Sciences and Deep Learning. This dataset…

Efficient and accurate extraction of key information from 2D engineering drawings is essential for advancing digital manufacturing workflows. Such information includes geometric dimensioning and tolerancing (GD&T), measures, material…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Muhammad Tayyab Khan , Lequn Chen , Zane Yong , Jun Ming Tan , Wenhe Feng , Seung Ki Moon

Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as presentation slides, scanned documents, and webpage snapshots. Tasks involving multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Mengzhao Jia , Wenhao Yu , Kaixin Ma , Tianqing Fang , Zhihan Zhang , Siru Ouyang , Hongming Zhang , Dong Yu , Meng Jiang

Vision-Language Pre-training (VLP) has achieved impressive performance on various cross-modal downstream tasks. However, most existing methods can only learn from aligned image-caption data and rely heavily on expensive regional features,…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Wei Li , Can Gao , Guocheng Niu , Xinyan Xiao , Hao Liu , Jiachen Liu , Hua Wu , Haifeng Wang

It is still challenging to build an AI system that can perform tasks that involve vision and language at human level. So far, researchers have singled out individual tasks separately, for each of which they have designed networks and…

Computer Vision and Pattern Recognition · Computer Science 2018-12-04 Duy-Kien Nguyen , Takayuki Okatani

With an increasing demand for assistive technologies that promote the independence and mobility of visually impaired people, this study suggests an innovative real-time system that gives audio descriptions of a user's surroundings to…

Human-Computer Interaction · Computer Science 2025-03-26 Kunal Chavan , Keertan Balaji , Spoorti Barigidad , Samba Raju Chiluveru

This paper introduces a holistic vision-language foundation model tailored for remote sensing, named Falcon. Falcon offers a unified, prompt-based paradigm that effectively executes comprehensive and complex remote sensing tasks. Falcon…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Kelu Yao , Nuo Xu , Rong Yang , Yingying Xu , Zhuoyan Gao , Titinunt Kitrungrotsakul , Yi Ren , Pu Zhang , Jin Wang , Ning Wei , Chao Li

While Ferret seamlessly integrates regional understanding into the Large Language Model (LLM) to facilitate its referring and grounding capability, it poses certain limitations: constrained by the pre-trained fixed visual encoder and failed…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Haotian Zhang , Haoxuan You , Philipp Dufter , Bowen Zhang , Chen Chen , Hong-You Chen , Tsu-Jui Fu , William Yang Wang , Shih-Fu Chang , Zhe Gan , Yinfei Yang

Vision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English languages, we introduce UC2, the first machine…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Mingyang Zhou , Luowei Zhou , Shuohang Wang , Yu Cheng , Linjie Li , Zhou Yu , Jingjing Liu
‹ Prev 1 2 3 10 Next ›