中文
相关论文

相关论文: MGeo: Multi-Modal Geographic Pre-Training Method

200 篇论文

Geo-localization aims to infer the geographic location where an image was captured using observable visual evidence. Traditional methods achieve impressive results through large-scale training on massive image corpora. With the emergence of…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Jinnao Li , Zijian Chen , Tingzhu Chen , Changbo Wang

Geospatial question answering (QA) is a fundamental task in navigation and point of interest (POI) searches. While existing geospatial QA datasets exist, they are limited in both scale and diversity, often relying solely on textual…

计算与语言 · 计算机科学 2025-03-12 Zekun Li , Malcolm Grossman , Eric , Qasemi , Mihir Kulkarni , Muhao Chen , Yao-Yi Chiang

Spatiotemporal relationships are critical in data science, as many prediction and reasoning tasks require analysis across both spatial and temporal dimensions--for instance, navigating an unfamiliar city involves planning itineraries that…

机器学习 · 计算机科学 2025-05-19 Xiao Han , Dayan Pan , Xiangyu Zhao , Xuyuan Hu , Zhaolin Deng , Xiangjie Kong , Guojiang Shen

Geolocation is now a vital aspect of modern life, offering numerous benefits but also presenting serious privacy concerns. The advent of large vision-language models (LVLMs) with advanced image-processing capabilities introduces new risks,…

密码学与安全 · 计算机科学 2024-08-20 Yi Liu , Junchen Ding , Gelei Deng , Yuekang Li , Tianwei Zhang , Weisong Sun , Yaowen Zheng , Jingquan Ge , Yang Liu

Witnessing the impressive achievements of pre-training techniques on large-scale data in the field of computer vision and natural language processing, we wonder whether this idea could be adapted in a grab-and-go spirit, and mitigate the…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Penghao Wu , Li Chen , Hongyang Li , Xiaosong Jia , Junchi Yan , Yu Qiao

Multimodal large language models (MLLMs) have shown remarkable capabilities across a broad range of tasks but their knowledge and abilities in the geographic and geospatial domains are yet to be explored, despite potential wide-ranging…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Jonathan Roberts , Timo Lüddecke , Rehan Sheikh , Kai Han , Samuel Albanie

Due to the rapid development of mobile Internet techniques, cloud computation and popularity of online social networking and location-based services, massive amount of multimedia data with geographical information is generated and uploaded…

多媒体 · 计算机科学 2018-08-21 Lei Zhu , Jun Long , Chengyuan Zhang , Ruipeng Chen , Xinpan Yuan , Zhan Yang

Spatio-temporal forecasting is crucial in transportation, logistics, and supply chain management. However, current methods struggle with large, complex datasets. We propose a dynamic, multi-modal approach that integrates the strengths of…

机器学习 · 计算机科学 2024-08-27 Sagar Srinivas Sakhinana , Geethan Sannidhi , Chidaksh Ravuru , Venkataramana Runkana

By employing large language models (LLMs) to retrieve documents and generate natural language responses, Generative Engines, such as Google AI overview and ChatGPT, provide significantly enhanced user experiences and have rapidly become the…

信息检索 · 计算机科学 2025-10-14 Yujiang Wu , Shanshan Zhong , Yubin Kim , Chenyan Xiong

Personalized recommendation of Points of Interest (POIs) plays a key role in satisfying users on Location-Based Social Networks (LBSNs). In this paper, we propose a probabilistic model to find the mapping between user-annotated tags and…

信息检索 · 计算机科学 2018-06-18 Mohammad Aliannejadi , Fabio Crestani

Multi-modal data in Earth Observation (EO) presents a huge opportunity for improving transfer learning capabilities when pre-training deep learning models. Unlike prior work that often overlooks multi-modal EO data, recent methods have…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Jose Sosa , Danila Rukhovich , Anis Kacem , Djamila Aouada

A key goal for the advancement of AI is to develop technologies that serve the needs not just of one group but of all communities regardless of their geographical region. In fact, a significant proportion of knowledge is locally shared by…

计算机视觉与模式识别 · 计算机科学 2023-01-06 Da Yin , Feng Gao , Govind Thattai , Michael Johnston , Kai-Wei Chang

Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Renqiu Xia , Mingsheng Li , Hancheng Ye , Wenjie Wu , Hongbin Zhou , Jiakang Yuan , Tianshuo Peng , Xinyu Cai , Xiangchao Yan , Bin Wang , Conghui He , Botian Shi , Tao Chen , Junchi Yan , Bo Zhang

Understanding human mobility behavior is crucial for numerous applications, including crowd management, location-based recommendations, and the estimation of pandemic spread. Machine learning models can predict the Points of Interest (POIs)…

机器学习 · 计算机科学 2024-11-26 Ziyao Li , Shang-Ling Hsu , Cyrus Shahabi

Location recommendation plays a vital role in improving users' travel experience. The timestamp of the POI to be predicted is of great significance, since a user will go to different places at different times. However, most existing methods…

信息检索 · 计算机科学 2023-04-11 Yan Luo , Haoyi Duan , Ye Liu , Fu-lai Chung

Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Ron Campos , Ashmal Vayani , Parth Parag Kulkarni , Rohit Gupta , Aizan Zafar , Aritra Dutta , Mubarak Shah

Nowadays large amounts of GPS trajectory data is being continuously collected by GPS-enabled devices such as vehicles navigation systems and mobile phones. GPS trajectory data is useful for applications such as traffic management, location…

人工智能 · 计算机科学 2016-05-18 Seyed Morteza Mousavi , Aaron Harwood , Shanika Karunasekera , Mojtaba Maghrebi

We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the significant…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Yin Wang , Ziyao Zhang , Zhiying Leng , Haitian Liu , Frederick W. B. Li , Mu Li , Xiaohui Liang

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based intelligence. While recent advances in large language…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Lingyao Li , Runlong Yu , Qikai Hu , Bowei Li , Min Deng , Yang Zhou , Xiaowei Jia

Visual Geo-localization (VG) refers to the process to identify the location described in query images, which is widely applied in robotics field and computer vision tasks, such as autonomous driving, metaverse, augmented reality, and SLAM.…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Chen Mao , Jingqi Hu