English
Related papers

Related papers: Estimating Absolute Web Crawl Coverage From Longit…

200 papers

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access…

General Economics · Economics 2023-08-07 Jens Foerderer

Online continual learning aims to get closer to a live learning experience by learning directly on a stream of data with temporally shifting distribution and by storing a minimum amount of data from that stream. In this empirical…

This paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained on large corpora…

Computation and Language · Computer Science 2020-03-16 Serge Sharoff

Existing grading systems for rock climbing routes assign a difficulty grade to a route based on the opinions of a few people. An objective approach to estimating route difficulty on an interval scale was obtained by adapting the…

Applications · Statistics 2020-07-21 Dean Scarff

Long-term Web archives comprise Web documents gathered over longer time periods and can easily reach hundreds of terabytes in size. Semantic annotations such as named entities can facilitate intelligent access to the Web archive data.…

Information Retrieval · Computer Science 2017-02-03 Tarcisio Souza , Elena Demidova , Thomas Risse , Helge Holzmann , Gerhard Gossen , Julian Szymanski

Biclustering is a two way clustering approach involving simultaneous clustering along two dimensions of the data matrix. Finding biclusters of web objects (i.e. web users and web pages) is an emerging topic in the context of web usage…

Neural and Evolutionary Computing · Computer Science 2011-06-14 R. Rathipriya , Dr. K. Thangavel , J. Bagyamani

Purpose: To provide a critical review of Bergman's 2001 study on the Deep Web. In addition, we bring a new concept into the discussion, the Academic Invisible Web (AIW). We define the Academic Invisible Web as consisting of all databases…

Digital Libraries · Computer Science 2019-01-15 Dirk Lewandowski , Philipp Mayr

Website fingerprinting attacks, which use statistical analysis on network traffic to compromise user privacy, have been shown to be effective even if the traffic is sent over anonymity-preserving networks such as Tor. The classical attack…

Cryptography and Security · Computer Science 2019-02-22 Anatoly Shusterman , Lachlan Kang , Yarden Haskal , Yosef Meltser , Prateek Mittal , Yossi Oren , Yuval Yarom

We describe our work in the collection and analysis of massive data describing the connections between participants to online social networks. Alternative approaches to social network data collection are defined and evaluated in practice,…

Social and Information Networks · Computer Science 2011-06-01 Salvatore A. Catanese , Pasquale De Meo , Emilio Ferrara , Giacomo Fiumara , Alessandro Provetti

Indexing the Web is becoming a laborious task for search engines as the Web exponentially grows in size and distribution. Presently, the most effective known approach to overcome this problem is the use of focused crawlers. A focused…

Information Retrieval · Computer Science 2015-10-02 Ali Seyfi

The web is today's primary publication medium, making web archiving an important activity for historical and analytical purposes. Web pages are increasingly interactive, resulting in pages that are increasingly difficult to archive.…

Digital Libraries · Computer Science 2016-01-21 Justin F. Brunelle , Michele C. Weigle , Michael L. Nelson

With the growing use of popular social media services like Facebook and Twitter it is challenging to collect all content from the networks without access to the core infrastructure or paying for it. Thus, if all content cannot be collected…

Social and Information Networks · Computer Science 2017-12-15 Fredrik Erlandsson , Piotr Bródka , Martin Boldt , Henric Johnson

Volunteered Geographic Information projects like OpenStreetMap which allow accessing and using the raw data, are a treasure trove for investigations - e.g. cultural topics, urban planning, or accessibility of services. Among the concerns…

Computers and Society · Computer Science 2023-06-09 Philipp Weigell

With the increasing scale and complexity of cloud systems and big data analytics platforms, it is becoming more and more challenging to understand and diagnose the processing of a service request in such distributed platforms. One way that…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-03-16 Yong Yang , Long Wang , Jing Gu , Ying Li

Curated web archive collections contain focused digital content which is collected by archiving organizations, groups, and individuals to provide a representative sample covering specific topics and events to preserve them for future…

Digital Libraries · Computer Science 2017-02-03 Zeon Trevor Fernando , Ivana Marenzi , Wolfgang Nejdl

Cache persistence analysis is an important part of worst-case execution time (WCET) analysis. It has been extensively studied in the past twenty years. Despite these efforts, all existing persistence analyses are approximative in the sense…

Programming Languages · Computer Science 2025-07-22 Gregory Stock , Sebastian Hahn , Jan Reineke

Monitoring the responses of plants to environmental changes is essential for plant biodiversity research. This, however, is currently still being done manually by botanists in the field. This work is very laborious, and the data obtained…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Matthias Körschens , Paul Bodesheim , Christine Römermann , Solveig Franziska Bucher , Mirco Migliavacca , Josephine Ulrich , Joachim Denzler

Q-learning is widely employed for optimizing various large-dimensional networks with unknown system dynamics. Recent advancements include multi-environment mixed Q-learning (MEMQ) algorithms, which utilize multiple independent Q-learning…

Machine Learning · Computer Science 2024-11-14 Talha Bozkus , Tara Javidi , Urbashi Mitra

The speed of an exhaustive search can be measured by a cover time, which is defined as the time it takes a random searcher to visit every state in some target set. Cover times have been studied in both the physics and probability…

Statistical Mechanics · Physics 2024-07-11 Hyunjoong Kim , Sean D Lawley

Most of the current methods for mining parallel texts from the web assume that web pages of web sites share same structure across languages. We believe that there still exists a non-negligible amount of parallel data spread across sources…

Computation and Language · Computer Science 2018-04-30 Jakub Kúdela , Irena Holubová , Ondřej Bojar