中文
相关论文

相关论文: Who and What Links to the Internet Archive

200 篇论文

The web is the prominent way information is exchanged in the 21st century. However, ensuring web-based information is accessible is complicated, particularly with web applications that rely on JavaScript and other technologies to deliver…

Web browsers are the most common tool to perform various activities over the internet. Along with normal mode, all modern browsers have private browsing mode. The name of the mode varies from browser to browser but the purpose of the…

密码学与安全 · 计算机科学 2018-03-01 Abu Awal Md Shoeb

As part of its 25th anniversary vision-setting process, the arXiv team at Cornell University Library conducted a user survey in April 2016 to seek input from the global user community about arXiv's current services and future directions. We…

数字图书馆 · 计算机科学 2016-07-28 Oya Y. Rieger , Gail Steinhart , Deborah Cooper

In the last few years, open-domain question answering (ODQA) has advanced rapidly due to the development of deep learning techniques and the availability of large-scale QA datasets. However, the current datasets are essentially designed for…

计算与语言 · 计算机科学 2022-02-23 Jiexin Wang , Adam Jatowt , Masatoshi Yoshikawa

To prevent the spread of disinformation on Instagram, we need to study the accounts and content of disinformation actors. However, due to their malicious nature, Instagram often bans accounts that are responsible for spreading…

数字图书馆 · 计算机科学 2024-01-05 Rachel Zheng , Michele C. Weigle

The core of the Web is a hyperlink navigation system collaboratively set up by webmasters to help users find desired websites. But does this system really work as expected? We show that the answer seems to be negative: there is a…

信息检索 · 计算机科学 2013-07-31 Lingfei Wu , Robert Ackland

In recent years, journalists and other researchers have used web archives as an important resource for their study of disinformation. This paper provides several examples of this use and also brings together some of the work that the Old…

数字图书馆 · 计算机科学 2023-06-19 Michele C. Weigle

The Open Archives Initiative (OAI) was created as a practical way to promote interoperability between eprint repositories. Although the scope of the OAI has been broadened, eprint repositories still represent a significant fraction of OAI…

数字图书馆 · 计算机科学 2007-05-23 Simeon Warner

Many online content portals allow users to ask questions to supplement their understanding (e.g., of lectures). While information retrieval (IR) systems may provide answers for such user queries, they do not directly assist content creators…

信息检索 · 计算机科学 2024-03-07 Rose E. Wang , Pawan Wirawarn , Omar Khattab , Noah Goodman , Dorottya Demszky

Crawler-based search engines are the mostly used search engines among web and Internet users, involve web crawling, storing in database, ranking, indexing and displaying to the user. But it is noteworthy that because of increasing changes…

信息检索 · 计算机科学 2013-05-14 Ali Tourani , Amir Seyed Danesh

Wikipedia is the largest existing knowledge repository that is growing on a genuine crowdsourcing support. While the English Wikipedia is the most extensive and the most researched one with over five million articles, comparatively little…

数字图书馆 · 计算机科学 2017-10-20 Kristina Ban , Matjaz Perc , Zoran Levnajic

E-journal preservation systems have to ingest millions of articles each year. Ingest, especially of the "long tail" of journals from small publishers, is the largest element of their cost. Cost is the major reason that archives contain less…

数字图书馆 · 计算机科学 2016-05-23 Herbert Van de Sompel , David S. H. Rosenthal , Michael L. Nelson

We analyse the darkweb and find its structure is unusual. For example, $ \sim 87 \%$ of darkweb sites \emph{never} link to another site. To call the darkweb a "web" is thus a misnomer -- it's better described as a set of largely isolated…

物理与社会 · 物理学 2020-06-04 Kevin P. O'Keeffe , Virgil Griffith , Yang Xu , Paolo Santi , Carlo Ratti

Objective: Information retrieval (IR, also known as search) systems are ubiquitous in modern times. How does the emergence of generative artificial intelligence (AI), based on large language models (LLMs), fit into the IR process? Process:…

信息检索 · 计算机科学 2025-01-20 William R. Hersh

Technology and the fruition of cultural heritage are becoming increasingly more entwined, especially with the advent of smart audio guides, virtual and augmented reality, and interactive installations. Machine learning and computer vision…

计算机视觉与模式识别 · 计算机科学 2020-12-30 Pietro Bongini , Federico Becattini , Andrew D. Bagdanov , Alberto Del Bimbo

A proposal for building an index of the Web that separates the infrastructure part of the search engine - the index - from the services part that will form the basis for myriad search engines and other services utilizing Web data on top of…

信息检索 · 计算机科学 2019-03-12 Dirk Lewandowski

Web archives are a valuable resource for researchers of various disciplines. However, to use them as a scholarly source, researchers require a tool that provides efficient access to Web archive data for extraction and derivation of smaller…

数字图书馆 · 计算机科学 2017-02-06 Helge Holzmann , Vinay Goel , Avishek Anand

Internet search engines function in a present which changes continuously. The search engines update their indices regularly, overwriting Web pages with newer ones, adding new pages to the index, and losing older ones. Some search engines…

信息检索 · 计算机科学 2009-11-19 Iina Hellsten , Loet Leydesdorff , Paul Wouters

The Internet and cyberspace are inseparable aspects of everyone's life. Cyberspace is a concept that describes widespread, interconnected, and online digital technology. Cyberspace refers to the online world that is separate from everyday…

人机交互 · 计算机科学 2023-03-27 Sarah Sharifi

Along with the rapid development of information technology, the amount of information generated at a given time far exceeds human's ability to organize, search, and manipulate without the help of automatic systems. Now a days so many tools…

信息检索 · 计算机科学 2013-05-08 Urmila Shrawankar , Anjali Mahajan