中文
相关论文

相关论文: Reconstructing Websites for the Lazy Webmaster

200 篇论文

Web archives preserve unique and historically valuable information. They hold a record of past events and memories published by all kinds of people, such as journalists, politicians and ordinary people who have shared their testimony and…

数字图书馆 · 计算机科学 2021-08-04 Miguel Costa , Julien Masanès

The rise of Generative AI Search is fundamentally transforming how users and intelligent systems interact with the Internet. LLMs increasingly act as intermediaries between humans and web information. Yet the web remains optimized for human…

网络与互联网体系结构 · 计算机科学 2025-11-25 Muhammad Bilal , Zafar Qazi , Marco Canini

Although the Internet Archive's Wayback Machine is the largest and most well-known web archive, there have been a number of public web archives that have emerged in the last several years. With varying resources, audiences and collection…

数字图书馆 · 计算机科学 2013-01-08 Scott G. Ainsworth , Ahmed AlSum , Hany SalahEldeen , Michele C. Weigle , Michael L. Nelson

Web space is the huge repository of data. Everyday lots of new information get added to this web space. The more the information, more is demand for tools to access that information. Answering users' queries about the online information…

信息检索 · 计算机科学 2012-06-19 Sudhir Ahuja , Mr. Rinkaj Goyal

Introduction: Before embarking on the design of any computer system it is first necessary to assess the magnitude of the problem. In the case of a web search engine this assessment amounts to determining the current size of the web, the…

信息检索 · 计算机科学 2013-07-05 Andrew Trotman , Jinglan Zhang

There are several ideas being used today for Web information retrieval, and specifically in Web search engines. The PageRank algorithm is one of those that introduce a content-neutral ranking function over Web pages. This ranking is applied…

分布式、并行与集群计算 · 计算机科学 2007-05-23 Giorgos Kollias , Efstratios Gallopoulos , Daniel B. Szyld

Crawler-based search engines are the mostly used search engines among web and Internet users, involve web crawling, storing in database, ranking, indexing and displaying to the user. But it is noteworthy that because of increasing changes…

信息检索 · 计算机科学 2013-05-14 Ali Tourani , Amir Seyed Danesh

Recent advances in browser-based LLM agents have shown promise for automating tasks ranging from simple form filling to hotel booking or online shopping. Current benchmarks measure agent performance in controlled environments, such as…

人工智能 · 计算机科学 2025-10-07 Su Kara , Fazle Faisal , Suman Nath

We present a framework for web-scale archiving of the dark web. While commonly associated with illicit and illegal activity, the dark web provides a way to privately access web information. This is a valuable and socially beneficial tool to…

数字图书馆 · 计算机科学 2021-07-12 Justin F. Brunelle , Ryan Farley , Grant Atkins , Trevor Bostic , Marites Hendrix , Zak Zebrowski

The internet contains large amounts of low-quality content, yet users expect web search engines to deliver high-quality, relevant results. The abundant presence of low-quality pages can negatively impact retrieval and crawling processes by…

信息检索 · 计算机科学 2025-04-16 Francesca Pezzuti , Ariane Mueller , Sean MacAvaney , Nicola Tonellotto

E-journal preservation systems have to ingest millions of articles each year. Ingest, especially of the "long tail" of journals from small publishers, is the largest element of their cost. Cost is the major reason that archives contain less…

数字图书馆 · 计算机科学 2016-05-23 Herbert Van de Sompel , David S. H. Rosenthal , Michael L. Nelson

Now no web search engine can cover more than 60 percent of all the pages on Internet. The update interval of most pages database is almost one month. This condition hasn't changed for many years. Converge and recency problems have become…

网络与互联网体系结构 · 计算机科学 2007-05-23 Wang Liang , Guo YiPing , Fang Ming

The field of web archiving provides a unique mix of human and automated agents collaborating to achieve the preservation of the web. Centuries old theories of archival appraisal are being transplanted into the sociotechnical environment of…

数字图书馆 · 计算机科学 2016-11-09 Ed Summers , Ricardo Punzalan

Nowadays, more and more people use the Web as their primary source of up-to-date information. In this context, fast crawling and indexing of newly created Web pages has become crucial for search engines, especially because user traffic to a…

信息检索 · 计算机科学 2013-07-25 Damien Lefortier , Liudmila Ostroumova , Egor Samosvat , Pavel Serdyukov

Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex,…

计算与语言 · 计算机科学 2025-08-12 Jialong Wu , Wenbiao Yin , Yong Jiang , Zhenglin Wang , Zekun Xi , Runnan Fang , Linhai Zhang , Yulan He , Deyu Zhou , Pengjun Xie , Fei Huang

World Wide Web consists of more than 50 billion pages online. It is highly dynamic i.e. the web continuously introduces new capabilities and attracts many people. Due to this explosion in size, the effective information retrieval system or…

信息检索 · 计算机科学 2012-05-15 Sk. AbdulNabi , P. Premchand

In web search, typically a candidate generation step selects a small set of documents---from collections containing as many as billions of web pages---that are subsequently ranked and pruned before being presented to the user. In Bing, the…

信息检索 · 计算机科学 2018-08-21 Corby Rosset , Damien Jose , Gargi Ghosh , Bhaskar Mitra , Saurabh Tiwary

Common Crawl is a multi-petabyte longitudinal dataset containing over 100 billion web pages which is widely used as a source of language data for sequence model training and in web science research. Each of its constituent archives is on…

网络与互联网体系结构 · 计算机科学 2024-04-16 Henry S. Thompson

The Semantic Web initiative puts emphasis not primarily on putting data on the Web, but rather on creating links in a way that both humans and machines can explore the Web of data. When such users access the Web, they leave a trail as Web…

信息检索 · 计算机科学 2011-04-07 Markus Kirchberg , Ryan K L Ko , Bu Sung Lee

Indexing the Web is becoming a laborious task for search engines as the Web exponentially grows in size and distribution. Presently, the most effective known approach to overcome this problem is the use of focused crawlers. A focused…

信息检索 · 计算机科学 2015-10-02 Ali Seyfi