中文
相关论文

相关论文: LLMGeo: Benchmarking Large Language Models on Imag…

200 篇论文

Reliable image geolocation is crucial for several applications, ranging from social media geo-tagging to fake news detection. State-of-the-art geolocation methods surpass human performance on the task of geolocation estimation from images.…

计算机视觉与模式识别 · 计算机科学 2021-11-24 Apostolos Panagiotopoulos , Giorgos Kordopatis-Zilos , Symeon Papadopoulos

The choice of representation for geographic location significantly impacts the accuracy of models for a broad range of geospatial tasks, including fine-grained species classification, population density estimation, and biome classification.…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Aayush Dhakal , Srikumar Sastry , Subash Khanal , Adeel Ahmad , Eric Xing , Nathan Jacobs

Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal…

The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches leverage world knowledge, chain-of-thought reasoning, and…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Yuxiang Ji , Yong Wang , Ziyu Ma , Yiming Hu , Hailang Huang , Xuecai Hu , Guanhua Chen , Liaoni Wu , Xiangxiang Chu

Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we propose three new tasks and benchmarks to evaluate a model's…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Sindhu B Hegde , K R Prajwal , Taein Kwon , Andrew Zisserman

Worldwide geolocalization aims to locate the precise location at the coordinate level of photos taken anywhere on the Earth. It is very challenging due to 1) the difficulty of capturing subtle location-aware visual semantics, and 2) the…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Pengyue Jia , Yiding Liu , Xiaopeng Li , Yuhao Wang , Yantong Du , Xiao Han , Xuetao Wei , Shuaiqiang Wang , Dawei Yin , Xiangyu Zhao

Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions. This enrichment…

计算与语言 · 计算机科学 2024-07-08 Chang-Sheng Kao , Yun-Nung Chen

The rise of Multimodal Large Language Models (MLLMs) has become a transformative force in the field of artificial intelligence, enabling machines to process and generate content across multiple modalities, such as text, images, audio, and…

Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as presentation slides, scanned documents, and webpage snapshots. Tasks involving multiple…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Mengzhao Jia , Wenhao Yu , Kaixin Ma , Tianqing Fang , Zhihan Zhang , Siru Ouyang , Hongming Zhang , Dong Yu , Meng Jiang

Predicting the geographic location (geo-localization) from a single ground-level RGB image taken anywhere in the world is a very challenging problem. The challenges include huge diversity of images due to different environmental scenarios,…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Shraman Pramanick , Ewa M. Nowara , Joshua Gleason , Carlos D. Castillo , Rama Chellappa

Multimodal large language models (MLLMs) have achieved remarkable performance across diverse vision-and-language tasks. However, their potential in face recognition remains underexplored. In particular, the performance of open-source MLLMs…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Hatef Otroshi Shahreza , Sébastien Marcel

Image geolocalization is the challenging task of predicting the geographic coordinates of origin for a given photo. It is an unsolved problem relying on the ability to combine visual clues with general knowledge about the world to make…

计算机视觉与模式识别 · 计算机科学 2023-02-02 Lukas Haas , Silas Alberti , Michal Skreta

Visual localization is the problem of estimating the position and orientation from which a given image (or a sequence of images) is taken in a known scene. It is an important part of a wide range of computer vision and robotics…

计算机视觉与模式识别 · 计算机科学 2021-09-13 Ara Jafarzadeh , Manuel Lopez Antequera , Pau Gargallo , Yubin Kuang , Carl Toft , Fredrik Kahl , Torsten Sattler

Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-language models (VLMs) handle image-understanding tasks,…

In the past year, Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering, visual understanding and reasoning. However, the extensive model size and high training and…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Yizhang Jin , Jian Li , Yexin Liu , Tianjun Gu , Kai Wu , Zhengkai Jiang , Muyang He , Bo Zhao , Xin Tan , Zhenye Gan , Yabiao Wang , Chengjie Wang , Lizhuang Ma

This study aims to comprehensively review and empirically evaluate the application of multimodal large language models (MLLMs) and Large Vision Models (VLMs) in object detection for transportation systems. In the first fold, we provide a…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Huthaifa I. Ashqar , Ahmed Jaber , Taqwa I. Alhadidi , Mohammed Elhenawy

Vision-language models (VLMs) integrate visual and textual information, enabling a wide range of applications such as image captioning and visual question answering, making them crucial for modern AI systems. However, their high…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Gaurav Shinde , Anuradha Ravi , Emon Dey , Shadman Sakib , Milind Rampure , Nirmalya Roy

Large multimodal language models have demonstrated impressive capabilities in understanding and manipulating images. However, many of these models struggle with comprehending intensive textual contents embedded within the images, primarily…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Ruiyi Zhang , Yufan Zhou , Jian Chen , Jiuxiang Gu , Changyou Chen , Tong Sun

Existing Vision-Language models (VLMs) estimate either long-term trajectory waypoints or a set of control actions as a reactive solution for closed-loop planning based on their rich scene comprehension. However, these estimations are coarse…

机器人学 · 计算机科学 2024-04-01 Pranjal Paul , Anant Garg , Tushar Choudhary , Arun Kumar Singh , K. Madhava Krishna

Machine translation between many languages at once is highly challenging, since training with ground truth requires supervision between all language pairs, which is difficult to obtain. Our key insight is that, while languages may vary…

计算与语言 · 计算机科学 2022-04-04 Dídac Surís , Dave Epstein , Carl Vondrick