中文
相关论文

相关论文: Quantifying Geospatial in the Common Crawl Corpus

200 篇论文

As large language models (LLMs) continue to evolve, questions about their trustworthiness in delivering factual information have become increasingly important. This concern also applies to their ability to accurately represent the…

计算机与社会 · 计算机科学 2025-06-03 Omid Reza Abbasi , Franz Welscher , Georg Weinberger , Johannes Scholz

In this paper we present a preliminary analysis over the largest publicly accessible web dataset: the Common Crawl Corpus. We measure nine web characteristics from two levels of granularity using MapReduce and we comment on the initial…

信息检索 · 计算机科学 2014-09-30 Vasilis Kolias , Ioannis Anagnostopoulos , Eleftherios Kayafas

Large language models (LLMs) encode vast amounts of world knowledge. However, since these models are trained on large swaths of internet data, they are at risk of inordinately capturing information about dominant groups. This imbalance can…

计算与语言 · 计算机科学 2023-10-24 Pola Schwöbel , Jacek Golebiowski , Michele Donini , Cédric Archambeau , Danish Pruthi

Living in the era of data deluge, we have witnessed a web content explosion, largely due to the massive availability of User-Generated Content (UGC). In this work, we specifically consider the problem of geospatial information extraction…

数据库 · 计算机科学 2013-11-21 Georgios Skoumas , Dieter Pfoser , Anastasios Kyrillidis

Large Language Models (LLMs) are poised to play an increasingly important role in our lives, providing assistance across a wide array of tasks. In the geospatial domain, LLMs have demonstrated the ability to answer generic questions, such…

计算与语言 · 计算机科学 2024-11-13 Pasquale Balsebre , Weiming Huang , Gao Cong

Large language models (LLMs) have achieved huge success for their general knowledge and ability to solve a wide spectrum of tasks in natural language processing (NLP). Due to their impressive abilities, LLMs have shed light on potential…

This research focuses on assessing the ability of large language models (LLMs) in representing geometries and their spatial relations. We utilize LLMs including GPT-2 and BERT to encode the well-known text (WKT) format of geometries and…

计算与语言 · 计算机科学 2023-07-10 Yuhan Ji , Song Gao

The ability to transform location-centric geospatial data into meaningful computational representations has become fundamental to modern spatial analysis and decision-making. Geospatial Representation Learning (GRL), the process of…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Xixuan Hao , Yutian Jiang , Xingchen Zou , Jiabo Liu , Yifang Yin , Song Gao , Flora Salim , Tianrui Li , Yuxuan Liang

Large Language Models (LLMs) inherently carry the biases contained in their training corpora, which can lead to the perpetuation of societal harm. As the impact of these foundation models grows, understanding and evaluating their biases…

计算与语言 · 计算机科学 2024-10-08 Rohin Manvi , Samar Khanna , Marshall Burke , David Lobell , Stefano Ermon

Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of multimodal language…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Zhiqiang Wang , Dejia Xu , Rana Muhammad Shahroz Khan , Yanbin Lin , Zhiwen Fan , Xingquan Zhu

Georeferencing text documents has typically relied on either gazetteer-based methods to assign geographic coordinates to place names, or on language modelling approaches that associate textual terms with geographic locations. However, many…

人工智能 · 计算机科学 2026-01-26 Aneesha Fernando , Surangika Ranathunga , Kristin Stock , Raj Prasanna , Christopher B. Jones

Large language models (LLMs) are reported to be partial to certain cultures owing to the training data dominance from the English corpora. Since multilingual cultural data are often expensive to collect, existing efforts handle this by…

计算与语言 · 计算机科学 2024-12-04 Cheng Li , Mengzhou Chen , Jindong Wang , Sunayana Sitaram , Xing Xie

Web search queries concern place far more often than existing labelling schemes suggest, yet the landscape of geospatial web search queries - what people ask of place, and how often - remains poorly characterised at scale. We apply dense…

信息检索 · 计算机科学 2026-05-13 Ilya Ilyankou , Stefano Cavazzi , James Haworth

Foundation models have shown remarkable performance across diverse tasks, yet their ability to construct internal spatial world models for reasoning and planning remains unclear. We systematically evaluate the spatial understanding of large…

人工智能 · 计算机科学 2026-04-14 Weijiang Li , Yilin Zhu , Rajarshi Das , Parijat Dube

Recent work has shown that Pre-trained Language Models (PLMs) store the relational knowledge learned from data and utilize it for performing downstream tasks. However, commonsense knowledge across different regions may vary. For instance,…

计算与语言 · 计算机科学 2022-11-30 Da Yin , Hritik Bansal , Masoud Monajatipoor , Liunian Harold Li , Kai-Wei Chang

With the rise of electronic data, particularly Earth observation data, data-based geospatial modelling using machine learning (ML) has gained popularity in environmental research. Accurate geospatial predictions are vital for domain…

Interest is increasing among political scientists in leveraging the extensive information available in images. However, the challenge of interpreting these images lies in the need for specialized knowledge in computer vision and access to…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Yu Wang

Analyzing the geographic movement of humans, animals, and other phenomena is a growing field of research. This research has benefited urban planning, logistics, animal migration understanding, and much more. Typically, the movement is…

计算与语言 · 计算机科学 2022-02-01 Scott Pezanowski , Prasenjit Mitra

The integration of advanced Natural Language Processing (NLP) methodologies and Large Language Models (LLMs) has significantly enhanced the extraction and analysis of geospatial data from multilingual texts, impacting sectors such as…

计算与语言 · 计算机科学 2024-12-31 Kalin Kopanov

Large language models (LLMs) under-perform on low-resource languages due to limited training data. We present a method to efficiently collect text data for low-resource languages from the entire Common Crawl corpus. Our approach,…

计算与语言 · 计算机科学 2024-11-22 Bethel Melesse Tessema , Akhil Kedia , Tae-Sun Chung