English
Related papers

Related papers: Web scraping: a promising tool for geographic data…

200 papers

Web traffic is a valuable data source, typically used in the marketing space to track brand awareness and advertising effectiveness. However, web traffic is also a rich source of information for cybersecurity monitoring efforts. To better…

Information Retrieval · Computer Science 2019-04-04 Han Qin , Kit Riehle , Haozhen Zhao

This study investigates the mechanisms of Surveillance Capitalism, focusing on personal data transfer during web navigation and searching. Analyzing network traffic reveals how various entities track and harvest digital footprints. The…

Artificial Intelligence · Computer Science 2024-12-25 Antony Seabra de Medeiros , Luiz Afonso Glatzl Junior , Sergio Lifschitz

Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Vikram V. Ramaswamy , Sing Yu Lin , Dora Zhao , Aaron B. Adcock , Laurens van der Maaten , Deepti Ghadiyaram , Olga Russakovsky

Web images come in hand with valuable contextual information. Although this information has long been mined for various uses such as image annotation, clustering of images, inference of image semantic content, etc., insufficient attention…

Multimedia · Computer Science 2020-05-21 F. Fauzi , H. J. Long , M. Belkhatir

In recent years, the study of complex networks has received a lot of attention. Real systems have gained importance in scientific publications, despite of an important drawback: the difficulty of retrieving and manage such great quantity of…

Computers and Society · Computer Science 2007-10-29 Massimiliano Zanin

Getting informed of what is registered in the Web space on time, can greatly help the psychologists, marketers and political analysts to familiarize, analyse, make decision and act correctly based on the society`s different needs. The great…

Information Retrieval · Computer Science 2012-02-10 Mehdi Naghavi , Mohsen Sharifi

Web archives preserve portions of the web, but quantifying their completeness remains challenging. Prior approaches have estimated the coverage of a crawl by either comparing the outcomes of multiple crawlers, or by comparing the results of…

Physics and Society · Physics 2026-04-07 Michael Paris , Grigori Paris , Fabian Baumann

Nowadays, the huge amount of information distributed through the Web motivates studying techniques to be adopted in order to extract relevant data in an efficient and reliable way. Both academia and enterprises developed several approaches…

Artificial Intelligence · Computer Science 2013-06-06 Emilio Ferrara , Robert Baumgartner

Living in the era of data deluge, we have witnessed a web content explosion, largely due to the massive availability of User-Generated Content (UGC). In this work, we specifically consider the problem of geospatial information extraction…

Databases · Computer Science 2013-11-21 Georgios Skoumas , Dieter Pfoser , Anastasios Kyrillidis

This article evaluates the quality of data collection in individual-level desktop information tracking used in the social sciences and shows that the existing approaches face sampling issues, validity issues due to the lack of content-level…

Today's geo-location estimation approaches are able to infer the location of a target image using its visual content alone. These approaches exploit visual matching techniques, applied to a large collection of background images with known…

Computers and Society · Computer Science 2016-03-07 Jaeyoung Choi , Martha Larson , Xinchao Li , Gerald Friedland , Alan Hanjalic

Networks are structures that pervade many natural and man-made phenomena. Recent findings have characterized many networks as not random structures, but as efficent complex formations. Current research has examined complex networks as…

Disordered Systems and Neural Networks · Physics 2007-05-23 Sean P. Gorman , Rajendra Kulkarni

`Tracking' is the collection of data about an individual's activity across multiple distinct contexts and the retention, use, or sharing of data derived from that activity outside the context in which it occurred. This paper aims to…

Cryptography and Security · Computer Science 2022-01-27 Reuben Binns

Mobile devices with rich features can record videos, traffic parameters or air quality readings along user trajectories. Although such data may be valuable, users are seldom rewarded for collecting them. Emerging digital marketplaces allow…

Cryptography and Security · Computer Science 2019-10-02 Kien Nguyen , Gabriel Ghinita , Muhammad Naveed , Cyrus Shahabi

Nowadays, society has recognized that the lack of access to spatial data and tools for their analysis is the limiting factor of economic development. It came to the realization that without the single information space, which is implemented…

Software Engineering · Computer Science 2012-05-07 Evgeny V. Shulkin , Sergey M. Krasnopeyev

The growth of world-wide-web (WWW) spreads its wings from an intangible quantities of web-pages to a gigantic hub of web information which gradually increases the complexity of crawling process in a search engine. A search engine handles a…

Machine Learning · Computer Science 2012-08-15 Sudarshan Nandy , Partha Pratim Sarkar , Achintya Das

The widespread adoption of continuously connected smartphones and tablets developed the usage of mobile applications, among which many use location to provide geolocated services. These services provide new prospects for users: getting…

Cryptography and Security · Computer Science 2018-10-09 Primault Vincent , Boutet Antoine , Ben Mokhtar Sonia , Brunie Lionel

Retrieve information resources made by the machine processing may refer to multiple sources. A personal web as part of information resources in the Internet requires a feature that can be understood by computer machines. Therefore, in this…

Digital Libraries · Computer Science 2013-12-23 Istiadi , Azhari

The growth of world-wide-web (WWW) spreads its wings from an intangible quantities of web-pages to a gigantic hub of web information which gradually increases the complexity of crawling process in a search engine. A search engine handles a…

Information Retrieval · Computer Science 2012-08-14 Sudarshan Nandy , Partha Pratim Sarkar , Achintya Das

The emergence of new digital technologies has allowed the study of human behaviour at a scale and at level of granularity that were unthinkable just a decade ago. In particular, by analysing the digital traces left by people interacting in…

Computers and Society · Computer Science 2019-08-09 Mirco Musolesi