English
Related papers

Related papers: Web scraping: a promising tool for geographic data…

200 papers

Scientists across disciplines often use data from the internet to conduct research, generating valuable insights about human behavior. However, as generative AI relying on massive text corpora becomes increasingly valuable, platforms have…

Computers and Society · Computer Science 2024-12-20 Megan A. Brown , Andrew Gruen , Gabe Maldoff , Solomon Messing , Zeve Sanderson , Michael Zimmer

Search engines are a combination of hardware and computer software supplied by a particular company through the website which has been determined. Search engines collect information from the web through bots or web crawlers that crawls the…

Information Retrieval · Computer Science 2014-10-22 Ahmad Josi , Leon Andretti Abdillah , Suryayusra

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

General Economics · Economics 2023-08-07 Jens Foerderer

As unconventional sources of geo-information, massive imagery and text messages from open platforms and social media form a temporally quasi-seamless, spatially multi-perspective stream, but with unknown and diverse quality. Due to its…

As the amount of data on the World Wide Web continues to grow exponentially, access to semantically structured information remains limited. The Semantic Web has emerged as a solution to enhance the machine-readability of data, making it…

Digital Libraries · Computer Science 2023-06-21 Muhammad Zohaib

One of the elements that have popularized and facilitated the use of geographical information on a variety of computational applications has been the use of Web maps; this has opened new research challenges on different subjects, from…

Computers and Society · Computer Science 2009-11-05 Rafael Ponce-Medellin , Gabriel Gonzalez-Serna , Rocio Vargas , Lirio Ruiz

The understanding of the immense and intricate topological structure of the World Wide Web (WWW) is a major scientific and technological challenge. This has been tackled recently by characterizing the properties of its representative graphs…

Networking and Internet Architecture · Computer Science 2008-01-23 M. Angeles Serrano , Ana Maguitman , Marian Boguna , Santo Fortunato , Alessandro Vespignani

As the exploration of digital behavioral data revolutionizes communication research, understanding the nuances of data collection methodologies becomes increasingly pertinent. This study focuses on one prominent data collection approach,…

Computers and Society · Computer Science 2024-12-03 Roberto Ulloa , Frank Mangold , Felix Schmidt , Judith Gilsbach , Sebastian Stier

The clear, social, and dark web have lately been identified as rich sources of valuable cyber-security information that -given the appropriate tools and methods-may be identified, crawled and subsequently leveraged to actionable…

Cryptography and Security · Computer Science 2021-09-16 Paris Koloveas , Thanasis Chantzios , Christos Tryfonopoulos , Spiros Skiadopoulos

Gathering enough data to create sufficiently useful training datasets for generative artificial intelligence requires scraping most public websites. The scraping is conducted using pieces of code (scraping bots) that make copies of website…

Computers and Society · Computer Science 2025-04-02 David Atkinson

Volunteered Geographic Information projects like OpenStreetMap which allow accessing and using the raw data, are a treasure trove for investigations - e.g. cultural topics, urban planning, or accessibility of services. Among the concerns…

Computers and Society · Computer Science 2023-06-09 Philipp Weigell

Web scraping is a technique that allows us to extract data from websites automatically. in the field of medicine, web scraping can be used to collect information about medical procedures, treatments, and healthcare providers. this…

Computation and Language · Computer Science 2023-06-22 Niketha Sabesan , Nivethitha , J. N Shreyah , Pranauv A J , Shyam R

Web Data Extraction is an important problem that has been studied by means of different scientific tools and in a broad range of applications. Many approaches to extracting data from the Web have been designed to solve specific problems and…

Information Retrieval · Computer Science 2017-03-07 Emilio Ferrara , Pasquale De Meo , Giacomo Fiumara , Robert Baumgartner

Web crawlers visit internet applications, collect data, and learn about new web pages from visited pages. Web crawlers have a long and interesting history. Early web crawlers collected statistics about the web. In addition to collecting…

In recent years, ethical issues in the networking field are getting moreimportant. In particular, there is a consistent debate about how Internet Service Providers (ISPs) should collect and treat network measurements. This kind of…

Computers and Society · Computer Science 2017-03-23 Leonardo Regano , Ali Safari Khatouni , Martino Trevisan , Alessio Viticchie

Geographic data plays an essential role in various Web, Semantic Web and machine learning applications. OpenStreetMap and knowledge graphs are critical complementary sources of geographic data on the Web. However, data veracity, the lack of…

Artificial Intelligence · Computer Science 2023-02-20 Elena Demidova , Alishiba Dsouza , Simon Gottschalk , Nicolas Tempelmeier , Ran Yu

Websites are regarded as domains of limitless information which anyone and everyone can access. The new trend of technology put us to change the way we are doing our business. The Internet now is fastly becoming a new place for business and…

Information Retrieval · Computer Science 2021-09-03 Ikechukwu Onyenwe , Ebele Onyedinma , Chidinma Nwafor , Obinna Agbata

Web archive analytics is the exploitation of publicly accessible web pages and their evolution for research purposes -- to the extent organizationally possible for researchers. In order to better understand the complexity of this task, the…

Digital Libraries · Computer Science 2021-07-05 Michael Völske , Janek Bevendorff , Johannes Kiesel , Benno Stein , Maik Fröbe , Matthias Hagen , Martin Potthast

Nowadays, the users' browsing activity on the Internet is not completely private due to many entities that collect and use such data, either for legitimate or illegal goals. The implications are serious, from a person who exposes…

Computers and Society · Computer Science 2021-05-05 Luca Vassio , Hassan Metwalley , Danilo Giordano

In this paper, we analyze the nature and distribution of structured data on the Web. Web-scale information extraction, or the problem of creating structured tables using extraction from the entire web, is gathering lots of research…

Databases · Computer Science 2012-03-30 Nilesh Dalvi , Ashwin Machanavajjhala , Bo Pang
‹ Prev 1 2 3 10 Next ›