中文
相关论文

相关论文: Who and What Links to the Internet Archive

200 篇论文

This article seeks to determine the extent to which the principle of persistence is observed by repositories and the organizations that operate them. We also evaluate the impact that negative repository persistence levels may be having on…

数字图书馆 · 计算机科学 2026-01-13 George Macgregor , Joy Davidson

A variety of fan-based wikis about episodic fiction (e.g., television shows, novels, movies) exist on the World Wide Web. These wikis provide a wealth of information about complex stories, but if readers are behind in their viewing they run…

数字图书馆 · 计算机科学 2015-06-23 Shawn M. Jones , Michael L. Nelson

The World Wide Web (WWW) is the repository of large number of web pages which can be accessed via Internet by multiple users at the same time and therefore it is Ubiquitous in nature. The search engine is a key application used to search…

数据库 · 计算机科学 2012-09-25 K. C. Srikantaiah , P. L. Srikanth , V. Tejaswi , K. Shaila , K. R. Venugopal , L. M. Patnaik

The rise of intelligent assistant systems like Siri and Alexa have led to the emergence of Conversational Search, a research track of Information Retrieval (IR) that involves interactive and iterative information-seeking user-system dialog.…

信息检索 · 计算机科学 2021-02-09 Somil Gupta , Neeraj Sharma

A large amount of data on the WWW remains inaccessible to crawlers of Web search engines because it can only be exposed on demand as users fill out and submit forms. The Hidden web refers to the collection of Web data which can be accessed…

信息检索 · 计算机科学 2014-07-23 Sonali Gupta , Komal Kumar Bhatia

This article provides a quantitative analysis of privacy-compromising mechanisms on 1 million popular websites. Findings indicate that nearly 9 in 10 websites leak user data to parties of which the user is likely unaware; more than 6 in 10…

密码学与安全 · 计算机科学 2015-11-03 Timothy Libert

Web archives, a key area of digital preservation, meet the needs of journalists, social scientists, historians, and government organizations. The use cases for these groups often require that they guide the archiving process themselves,…

数字图书馆 · 计算机科学 2021-01-26 Shawn M. Jones , Alexander Nwala , Michele C. Weigle , Michael L. Nelson

Curated web archive collections contain focused digital content which is collected by archiving organizations, groups, and individuals to provide a representative sample covering specific topics and events to preserve them for future…

数字图书馆 · 计算机科学 2017-02-03 Zeon Trevor Fernando , Ivana Marenzi , Wolfgang Nejdl

Social graph construction from various sources has been of interest to researchers due to its application potential and the broad range of technical challenges involved. The World Wide Web provides a huge amount of continuously updated data…

社会与信息网络 · 计算机科学 2017-01-13 Miroslav Shaltev , Jan-Hendrik Zab , Philipp Kemkes , Stefan Siersdorfer , Sergej Zerr

This paper investigates the composition of search engine results pages. We define what elements the most popular web search engines use on their results pages (e.g., organic results, advertisements, shortcuts) and to which degree they are…

信息检索 · 计算机科学 2015-11-19 Nadine Hoechstoetter , Dirk Lewandowski

Lawrence (2001)found computer science articles that were openly accessible (OA) on the Web were cited more. We replicated this in physics. We tested 1,307,038 articles published across 12 years (1992-2003) in 10 disciplines (Biology,…

数字图书馆 · 计算机科学 2007-05-23 C. Hajjem , S. Harnad , Y. Gingras

This paper describes how born digital primary sources could be used to reconstruct the recent history of scientific institutions. The case study is an analysis of the first 25 years online of the University of Bologna. The focus of this…

数字图书馆 · 计算机科学 2016-04-21 Federico Nanni

Software is often developed using versioned controlled software, such as Git, and hosted on centralized Web hosts, such as GitHub and GitLab. These Web hosted software repositories are made available to users in the form of traditional HTML…

数字图书馆 · 计算机科学 2025-05-22 David Calano , Michele C. Weigle , Michael L. Nelson

HTTP Server is a computer programs that serves webpage content to clients. A webpage is a document or resource of information that is suitable for the World Wide Web and can be accessed through a web browser and displayed on a computer…

其他计算机科学 · 计算机科学 2010-03-09 Bala Dhandayuthapani Veerasamy

Artificial intelligence (AI) like deep learning, cloud AI computation has been advancing at a rapid pace since 2014. There is no doubt that the prosperity of AI is inseparable with the development of the Internet. However, there has been…

计算机与社会 · 计算机科学 2018-01-19 Feng Liu , Yong Shi , Peijia Lia

A goal shared by artificial intelligence and information retrieval is to create an oracle, that is, a machine that can answer our questions, no matter how difficult they are. A more limited, but still instrumental, version of this oracle is…

信息检索 · 计算机科学 2019-08-20 Rodrigo Nogueira

The Web is a tangled mass of interconnected services, where websites import a range of external resources from various third-party domains. However, the latter can further load resources hosted on other domains. For each website, this…

密码学与安全 · 计算机科学 2019-02-19 Muhammad Ikram , Rahat Masood , Gareth Tyson , Mohamed Ali Kaafar , Noha Loizon , Roya Ensafi

Web search engines have marked everyone's life by transforming how one searches and accesses information. Search engines give special attention to the user interface, especially search engine result pages (SERP). The well-known ''10 blue…

信息检索 · 计算机科学 2023-01-23 B. Oliveira , C. T. Lopes

Looking into the growth of information in the web it is a very tedious process of getting the exact information the user is looking for. Many search engines generate user profile related data listing. This paper involves one such process…

信息检索 · 计算机科学 2011-09-12 L. K. Joshila Grace , V. Maheswari , Dhinaharan Nagamalai

Common Crawl is a multi-petabyte longitudinal dataset containing over 100 billion web pages which is widely used as a source of language data for sequence model training and in web science research. Each of its constituent archives is on…

网络与互联网体系结构 · 计算机科学 2024-04-16 Henry S. Thompson