English
Related papers

Related papers: Full-Text and URL Search Over Web Archives

200 papers

We present a framework for web-scale archiving of the dark web. While commonly associated with illicit and illegal activity, the dark web provides a way to privately access web information. This is a valuable and socially beneficial tool to…

Digital Libraries · Computer Science 2021-07-12 Justin F. Brunelle , Ryan Farley , Grant Atkins , Trevor Bostic , Marites Hendrix , Zak Zebrowski

Due to the increasing storage data on Web Applications, it becomes very difficult to use only keyword-based searches to provide comprehensive search results, thus increasing the difficulty for web users to search information on the web. In…

Information Retrieval · Computer Science 2021-10-12 Ikechukwu Onyenwe , Stanley Ogbonna , Ebele Onyedimma , Onyedikachukwu Ikechukwu-Onyenwe , Chidinma Nwafor

Since the era of big data, the Internet has been flooded with all kinds of information. Browsing information through the Internet has become an integral part of people's daily life. Unlike the news data and social data in the Internet, the…

Information Retrieval · Computer Science 2022-10-26 Yang Jiang , Zhe Xue , Ang Li

With the advent of the cloud computing era, the cost of creating, capturing and managing information has gradually decreased. The amount of data in the Internet is also showing explosive growth, and more and more scientific and…

Digital Libraries · Computer Science 2022-04-12 Yue Wang , Zhe Xue , Ang Li

Retrieval and content management are assumed to be mutually exclusive. In this paper we suggest that they need not be so. In the usual information retrieval scenario, some information about queries leading to a website (due to `hits' or…

Information Retrieval · Computer Science 2019-08-29 C Ravindranath Chowdary , Anil Kumar Singh , Anil Nelakanti

Rapidly growing online podcast archives contain diverse content on a wide range of topics. These archives form an important resource for entertainment and professional use, but their value can only be realized if users can rapidly and…

Information Retrieval · Computer Science 2021-08-27 Ben Carterette , Rosie Jones , Gareth F. Jones , Maria Eskevich , Sravana Reddy , Ann Clifton , Yongze Yu , Jussi Karlgren , Ian Soboroff

Search engines like Google, Yahoo or Bing are an excellent support for finding documents, but this strength also imposes a limitation. As they are optimized for document retrieval tasks, they perform less well when it comes to more complex…

Information Retrieval · Computer Science 2012-06-27 Kristiina Singer , Georg Singer , Krista Lepik , Ulrich Norbisrath , Pille Pruulmann-Vengerfeldt

Personal and private Web archives are proliferating due to the increase in the tools to create them and the realization that Internet Archive and other public Web archives are unable to capture personalized (e.g., Facebook) and private…

Digital Libraries · Computer Science 2018-06-05 Mat Kelly , Michael L. Nelson , Michele C. Weigle

YouTube (http://www.youtube.com) is an online, public-access video-sharing site that allows users to post short streaming-video submissions for open viewing. Along with Google, MySpace, Facebook, etc. it is one of the great success stories…

Physics Education · Physics 2008-08-27 A. P. Micolich

In this world, globalization has become a basic and most popular human trend. To globalize information, people are going to publish the documents in the internet. As a result, information volume of internet has become huge. To handle that…

Information Retrieval · Computer Science 2013-11-26 Sukanta Sinha , Rana Dattagupta , Debajyoti Mukhopadhyay

Although the Internet Archive's Wayback Machine is the largest and most well-known web archive, there have been a number of public web archives that have emerged in the last several years. With varying resources, audiences and collection…

Digital Libraries · Computer Science 2013-01-08 Scott G. Ainsworth , Ahmed AlSum , Hany SalahEldeen , Michele C. Weigle , Michael L. Nelson

In this paper, we present a meta-analysis of several Web content extraction algorithms, and make recommendations for the future of content extraction on the Web. First, we find that nearly all Web content extractors do not consider a very…

Information Retrieval · Computer Science 2015-08-19 Tim Weninger , Rodrigo Palacios , Valter Crescenzi , Thomas Gottron , Paolo Merialdo

Web crawlers visit internet applications, collect data, and learn about new web pages from visited pages. Web crawlers have a long and interesting history. Early web crawlers collected statistics about the web. In addition to collecting…

We are presenting a set of multilingual text analysis tools that can help analysts in any field to explore large document collections quickly in order to determine whether the documents contain information of interest, and to find the…

Computation and Language · Computer Science 2007-05-23 Camelia Ignat , Bruno Pouliquen , Ralf Steinberger , Tomaz Erjavec

The Information and Communication Technologies revolution brought a digital world with huge amounts of data available. Enterprises use mining technologies to search vast amounts of data for vital insight and knowledge. Mining tools such as…

Information Retrieval · Computer Science 2013-04-15 Abdul-Aziz Rashid Al-Azmi

In this work, we study how URL extraction results depend on input format. We compiled a pilot dataset by extracting URLs from 10 arXiv papers and used the same heuristic method to extract URLs from four formats derived from the PDF files or…

Digital Libraries · Computer Science 2025-09-08 Rochana R. Obadage , Lamia Salsabil , Sawood Alam , Bipasha Banarjee , William A. Ingram , Edward A. Fox , Jian Wu

Now no web search engine can cover more than 60 percent of all the pages on Internet. The update interval of most pages database is almost one month. This condition hasn't changed for many years. Converge and recency problems have become…

Networking and Internet Architecture · Computer Science 2007-05-23 Wang Liang , Guo YiPing , Fang Ming

The vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose…

Archives are an important source of study for various scholars. Digitization and the web have made archives more accessible and led to the development of several time-aware exploratory search systems. However these systems have been…

Information Retrieval · Computer Science 2018-10-26 Jaspreet Singh , Wolfgang Nejdl , Avishek Anand

Web archives preserve portions of the web, but quantifying their completeness remains challenging. Prior approaches have estimated the coverage of a crawl by either comparing the outcomes of multiple crawlers, or by comparing the results of…

Physics and Society · Physics 2026-04-07 Michael Paris , Grigori Paris , Fabian Baumann
‹ Prev 1 3 4 5 6 7 10 Next ›