English
Related papers

Related papers: FUSAR-KLIP: Towards Multimodal Foundation Models f…

200 papers

This paper presents a grounded language-image pre-training (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Liunian Harold Li , Pengchuan Zhang , Haotian Zhang , Jianwei Yang , Chunyuan Li , Yiwu Zhong , Lijuan Wang , Lu Yuan , Lei Zhang , Jenq-Neng Hwang , Kai-Wei Chang , Jianfeng Gao

We present GeoGrid-Bench, a benchmark designed to evaluate the ability of foundation models to understand geo-spatial data in the grid structure. Geo-spatial datasets pose distinct challenges due to their dense numerical values, strong…

Computation and Language · Computer Science 2025-05-27 Bowen Jiang , Yangxinyu Xie , Xiaomeng Wang , Jiashu He , Joshua Bergerson , John K Hutchison , Jordan Branham , Camillo J Taylor , Tanwi Mallick

3D visual grounding allows an embodied agent to understand visual information in real-world 3D environments based on human instructions, which is crucial for embodied intelligence. Existing 3D visual grounding methods typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Fan Li , Zanyi Wang , Zeyi Huang , Guang Dai , Jingdong Wang , Mengmeng Wang

Remote Sensing Vision-Language Models (RS VLMs) have made much progress in the tasks of remote sensing (RS) image comprehension. While performing well in multi-modal reasoning and multi-turn conversations, the existing models lack…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Xu Liu , Zhouhui Lian

Large Multimodal Models (LMMs) demonstrate significant cross-modal reasoning capabilities. However, financial applications face challenges due to the lack of high-quality multimodal reasoning datasets and the inefficiency of existing…

Computation and Language · Computer Science 2025-06-17 Kai Lan , Jiayong Zhu , Jiangtong Li , Dawei Cheng , Guang Chen , Changjun Jiang

Accurately determining the geographic location where a single image was taken, visual geolocation, remains a formidable challenge due to the planet's vastness and the deceptive similarity among distant locations. We introduce GeoLocSFT, a…

Artificial Intelligence · Computer Science 2025-06-03 Qiang Yi , Lianlei Shan

Reconstructing an avatar from a portrait image has many applications in multimedia, but remains a challenging research problem. Extracting reflectance maps and geometry from one image is ill-posed: recovering geometry is a one-to-many…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Abdallah Dib , Luiz Gustavo Hafemann , Emeline Got , Trevor Anderson , Amin Fadaeinejad , Rafael M. O. Cruz , Marc-Andre Carbonneau

Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Kaiyu Li , Zepeng Xin , Li Pang , Chao Pang , Yupeng Deng , Jing Yao , Guisong Xia , Deyu Meng , Zhi Wang , Xiangyong Cao

Multimodal symbolic logical reasoning, which aims to deduce new facts from multimodal input via formal logic, is critical in high-stakes applications such as autonomous driving and medical diagnosis, as its rigorous, deterministic reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Jundong Xu , Hao Fei , Yuhui Zhang , Liangming Pan , Qijun Huang , Qian Liu , Preslav Nakov , Min-Yen Kan , William Yang Wang , Mong-Li Lee , Wynne Hsu

Remote sensing image-text retrieval plays a crucial role in remote sensing interpretation, yet remains challenging under both closed-domain and open-domain scenarios due to semantic noise and domain shifts. To address these issues, we…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Jiancheng Pan , Muyuan Ma , Qing Ma , Cong Bai , Shengyong Chen

Recent advances in vision-language models have enabled rich semantic understanding across modalities. However, these encoding methods lack the ability to interpret or reason about the moral dimensions of content-a crucial aspect of human…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Ana Carolina Condez , Diogo Tavares , João Magalhães

Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Zixuan Song , Jing Zhang , Di Wang , Zidie Zhou , Wenbin Liu , Haonan Guo , En Wang , Bo Du

Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address this, we propose the GeoSym Engine, an automated and scalable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jinhao Jing , Zheng Ma , Jinwei Liang , Qiannian Zhao , Shawn Chen , Jing Yang , Por Lip Yee , Prayag Tiwari , Jingjing Bai , Benyou Wang , Lewei Lu , Zhan Su

Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigation and emergency response. However, existing researches are constrained by narrow geographic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yutong Hu , Jinhui Chen , Chaoqiang Xu , Yuan Kou , Sili Zhou , Shaocheng Yan , Pengcheng Shi , Qingwu Hu , Jiayuan Li

Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation frameworks couple linguistic reasoning and pixel prediction through end-to-end supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xu Zhang , Junyao Ge , Yang Zheng , Kaitai Guo , Jimin Liang

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control. Yet their ability to generalize across new environments, tasks, and embodiments remains limited. We argue that a major bottleneck lies…

Face anti-spoofing (FAS) aims to construct a robust system that can withstand diverse attacks. While recent efforts have concentrated mainly on cross-domain generalization, two significant challenges persist: limited semantic understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Kun-Hsiang Lin , Yu-Wen Tseng , Kang-Yang Huang , Jhih-Ciang Wu , Wen-Huang Cheng

Fine-grained high-resolution remote sensing mapping typically relies on localized visual features, which restricts cross-domain generalizability and often leads to fragmented predictions of large-scale land covers. While global geospatial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Jienan Lyu , Miao Yang , Jinchen Cai , Yiwen Hu , Guanyi Lu , Junhao Qiu , Runmin Dong

Segmentation models can recognize a pre-defined set of objects in images. However, models that can reason over complex user queries that implicitly refer to multiple objects of interest are still in their infancy. Recent advances in…

Artificial Intelligence · Computer Science 2025-05-06 Jerome Quenum , Wen-Han Hsieh , Tsung-Han Wu , Ritwik Gupta , Trevor Darrell , David M. Chan

Cross-modal metric learning is a prominent research topic that bridges the semantic heterogeneity between vision and language. Existing methods frequently utilize simple cosine or complex distance metrics to transform the pairwise features…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Haiwen Diao , Ying Zhang , Shang Gao , Jiawen Zhu , Long Chen , Huchuan Lu
‹ Prev 1 4 5 6 7 8 10 Next ›