English
Related papers

Related papers: Longitudinal Sampling of URLs From the Wayback Mac…

200 papers

World Wide Web is a huge repository of web pages and links. It provides abundance of information for the Internet users. The growth of web is tremendous as approximately one million pages are added daily. Users' accesses are recorded in web…

Information Retrieval · Computer Science 2010-04-09 V. Chitraa , Dr. Antony Selvdoss Davamani

Nowadays, more and more people use the Web as their primary source of up-to-date information. In this context, fast crawling and indexing of newly created Web pages has become crucial for search engines, especially because user traffic to a…

Information Retrieval · Computer Science 2013-07-25 Damien Lefortier , Liudmila Ostroumova , Egor Samosvat , Pavel Serdyukov

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web…

Digital Libraries · Computer Science 2021-01-26 Shawn M. Jones , Michele C. Weigle , Michael L. Nelson

The availability of high definition video content on the web has brought about a significant change in the characteristics of Internet video, but not many studies on characterizing video have been done after this change. Video…

Multimedia · Computer Science 2014-08-26 Saba Ahsan , Varun Singh , Jörg Ott

There are several ideas being used today for Web information retrieval, and specifically in Web search engines. The PageRank algorithm is one of those that introduce a content-neutral ranking function over Web pages. This ranking is applied…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-05-23 Giorgos Kollias , Efstratios Gallopoulos , Daniel B. Szyld

ClueWeb22, the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information. Its design was influenced by the need for a high quality, large scale web corpus to support a range of academic…

Information Retrieval · Computer Science 2022-12-05 Arnold Overwijk , Chenyan Xiong , Xiao Liu , Cameron VandenBerg , Jamie Callan

In this work we propose MementoMap, a flexible and adaptive framework to efficiently summarize holdings of a web archive. We described a simple, yet extensible, file format suitable for MementoMap. We used the complete index of the…

Digital Libraries · Computer Science 2019-05-30 Sawood Alam , Michele C. Weigle , Michael L. Nelson , Fernando Melo , Daniel Bicho , Daniel Gomes

Modern browsers give access to several attributes that can be collected to form a browser fingerprint. Although browser fingerprints have primarily been studied as a web tracking tool, they can contribute to improve the current state of web…

Cryptography and Security · Computer Science 2021-10-05 Nampoina Andriamilanto , Tristan Allard , Gaëtan Le Guelvouit , Alexandre Garel

Web archiving is the process of collecting portions of the Web to ensure that the information is preserved for future exploitation. However, despite the increasing number of web archives worldwide, the absence of efficient and meaningful…

Digital Libraries · Computer Science 2018-10-25 Pavlos Fafalios , Helge Holzmann , Vaibhav Kasturia , Wolfgang Nejdl

To prevent the spread of disinformation on Instagram, we need to study the accounts and content of disinformation actors. However, due to their malicious nature, Instagram often bans accounts that are responsible for spreading…

Digital Libraries · Computer Science 2024-01-05 Rachel Zheng , Michele C. Weigle

Online data sources offer tremendous promise to demography and other social sciences, but researchers worry that the group of people who are represented in online datasets can be different from the general population. We show that by…

Applications · Statistics 2019-07-01 Dennis M. Feehan , Curtiss Cobb

User response prediction, which models the user preference w.r.t. the presented items, plays a key role in online services. With two-decade rapid development, nowadays the cumulated user behavior sequences on mature Internet service…

Information Retrieval · Computer Science 2019-05-14 Kan Ren , Jiarui Qin , Yuchen Fang , Weinan Zhang , Lei Zheng , Weijie Bian , Guorui Zhou , Jian Xu , Yong Yu , Xiaoqiang Zhu , Kun Gai

In this paper we present a preliminary analysis over the largest publicly accessible web dataset: the Common Crawl Corpus. We measure nine web characteristics from two levels of granularity using MapReduce and we comment on the initial…

Information Retrieval · Computer Science 2014-09-30 Vasilis Kolias , Ioannis Anagnostopoulos , Eleftherios Kayafas

Now no web search engine can cover more than 60 percent of all the pages on Internet. The update interval of most pages database is almost one month. This condition hasn't changed for many years. Converge and recency problems have become…

Networking and Internet Architecture · Computer Science 2007-05-23 Wang Liang , Guo YiPing , Fang Ming

Over the past few years, we have built a system that has exposed large volumes of Deep-Web content to Google.com users. The content that our system exposes contributes to more than 1000 search queries per-second and spans over 50 languages…

Databases · Computer Science 2009-09-15 Jayant Madhavan , Loredana Afanasiev , Lyublena Antova , Alon Halevy

In this paper, we focused on the problem of extracting information from web pages containing many records, a task of growing importance in the era of massive web data. Recently, the development of neural network methods has improved the…

Computation and Language · Computer Science 2025-02-21 Alexander Kustenkov , Maksim Varlamov , Alexander Yatskov

Web usage mining is a process of extracting useful information from server logs i.e. users history. Web usage mining is a process of finding out what users are looking for on the internet. Some users might be looking at only textual data,…

Information Retrieval · Computer Science 2013-10-25 P YesuRaju , P KiranSree

Summarizing web graphs is challenging due to the heterogeneity of the modeled information and its changes over time. We investigate the use of neural networks for lifelong graph summarization. Assuming we observe the web graph at a certain…

Machine Learning · Computer Science 2024-12-23 Jonatan Frank , Marcel Hoffmann , Nicolas Lell , David Richerby , Ansgar Scherp

This paper describes how born digital primary sources could be used to reconstruct the recent history of scientific institutions. The case study is an analysis of the first 25 years online of the University of Bologna. The focus of this…

Digital Libraries · Computer Science 2016-04-21 Federico Nanni

Wikipedia serves as a key infrastructure for public access to scientific knowledge, but it faces challenges in maintaining the credibility of cited sources--especially when scientific papers are retracted. This paper investigates how…

Human-Computer Interaction · Computer Science 2026-01-27 Haohan Shi , Yulin Yu , Daniel M. Romero , Emőke-Ágnes Horvát
‹ Prev 1 4 5 6 7 8 10 Next ›