中文
相关论文

相关论文: Carbon Dating The Web: Estimating the Age of Web R…

200 篇论文

Energy is today the most critical environmental challenge. The amount of carbon emissions contributing to climate change is significantly influenced by both the production and consumption of energy. Measuring and reducing the energy…

Event collections are frequently built by crawling the live web on the basis of seed URIs nominated by human experts. Focused web crawling is a technique where the crawler is guided by reference content pertaining to the event. Given the…

数字图书馆 · 计算机科学 2018-04-06 Martin Klein , Lyudmila Balakireva , Herbert Van de Sompel

Nowadays there are algorithms, methods, and platforms that are being created to accelerate the discovery of materials that are able to absorb or adsorb $CO_2$ molecules that are in the atmosphere or during the combustion in power plants,…

人工智能 · 计算机科学 2022-05-20 Maira Gatti de Bayser

Web crawling is the problem of keeping a cache of webpages fresh, i.e., having the most recent copy available when a page is requested. This problem is usually coupled with the natural restriction that the bandwidth available to the web…

机器学习 · 计算机科学 2019-11-26 Utkarsh Upadhyay , Robert Busa-Fekete , Wojciech Kotlowski , David Pal , Balazs Szorenyi

Most archived HTML pages embed other web resources, such as images and stylesheets. Playback of the archived web pages typically provides only the capture date (or Memento-Datetime) of the root resource and not the Memento-Datetime of the…

数字图书馆 · 计算机科学 2014-10-07 Scott G. Ainsworth , Michael L. Nelson , Herbert Van de Sompel

In the past decade, global warming made several headlines and turned the attention of the whole world to it. Carbon footprint is the main factor that drives greenhouse emissions up and results in the temperature increase of the planet with…

计算机与社会 · 计算机科学 2023-04-04 Michalis Pachilakis , Savino Dambra , Iskander Sanchez-Rola , Leyla Bilge

While the environmental impact of digitalization is becoming more and more evident, the climate crisis has become a major issue for society. For instance, data centers alone account for 2.7% of Europe's energy consumption today. A…

分布式、并行与集群计算 · 计算机科学 2023-10-31 Henrik Claßen , Jonas Thierfeldt , Julian Tochman-Szewc , Philipp Wiesner , Odej Kao

Information integration applications, such as mediators or mashups, that require access to information resources currently rely on users manually discovering and integrating them in the application. Manual resource discovery is a slow…

人工智能 · 计算机科学 2016-09-08 Anon Plangprasopchok , Kristina Lerman

The vastness of the web imposes a prohibitive cost on building large-scale search engines with limited resources. Crawl frontiers thus need to be optimized to improve the coverage and freshness of crawled content. In this paper, we propose…

Web archives preserve portions of the web, but quantifying their completeness remains challenging. Prior approaches have estimated the coverage of a crawl by either comparing the outcomes of multiple crawlers, or by comparing the results of…

物理与社会 · 物理学 2026-04-07 Michael Paris , Grigori Paris , Fabian Baumann

The carbon footprint share of the information and communication technology (ICT) sector has steadily increased in the past decade and is predicted to make up as much as 23 \% of global emissions in 2030. This shows a pressing need for…

信息检索 · 计算机科学 2024-06-07 Noah Gießing , Madhurima Deb , Ankit Satpute , Moritz Schubotz , Olaf Teschke

The use of citation counts to assess the impact of research articles is well established. However, the citation impact of an article can only be measured several years after it has been published. As research articles are increasingly…

信息检索 · 计算机科学 2007-05-23 Tim Brody , Stevan Harnad

There is unison is the scientific community about human induced climate change. Despite this, we see the web awash with claims around climate change scepticism, thus driving the need for fact checking them but at the same time providing an…

计算与语言 · 计算机科学 2021-08-02 Shraey Bhatia , Jey Han Lau , Timothy Baldwin

When a user requests a web page from a web archive, the user will typically either get an HTTP 200 if the page is available, or an HTTP 404 if the web page has not been archived. This is because web archives are typically accessed by URI…

数字图书馆 · 计算机科学 2019-08-09 Lulwah M. Alkwai , Michael L. Nelson , Michele C. Weigle

Looking into the growth of information in the web it is a very tedious process of getting the exact information the user is looking for. Many search engines generate user profile related data listing. This paper involves one such process…

信息检索 · 计算机科学 2011-09-12 L. K. Joshila Grace , V. Maheswari , Dhinaharan Nagamalai

Nowadays, more and more people use the Web as their primary source of up-to-date information. In this context, fast crawling and indexing of newly created Web pages has become crucial for search engines, especially because user traffic to a…

信息检索 · 计算机科学 2013-07-25 Damien Lefortier , Liudmila Ostroumova , Egor Samosvat , Pavel Serdyukov

As the exploration of digital behavioral data revolutionizes communication research, understanding the nuances of data collection methodologies becomes increasingly pertinent. This study focuses on one prominent data collection approach,…

计算机与社会 · 计算机科学 2024-12-03 Roberto Ulloa , Frank Mangold , Felix Schmidt , Judith Gilsbach , Sebastian Stier

Looking into the growth of information in the web it is a very tedious process of getting the exact information the user is looking for. Many search engines generate user profile related data listing. This paper involves one such process…

信息检索 · 计算机科学 2011-09-12 L. K. Joshila Grace , V. Maheswari , Dhinaharan Nagamalai

As Digital Libraries (DL) become more aligned with the web architecture, their functional components need to be fundamentally rethought in terms of URIs and HTTP. Annotation, a core scholarly activity enabled by many DL solutions, exhibits…

数字图书馆 · 计算机科学 2010-03-22 Robert Sanderson , Herbert Van de Sompel

For providing quick and accurate results, a search engine maintains a local snapshot of the entire web. And, to keep this local cache fresh, it employs a crawler for tracking changes across various web pages. However, finite bandwidth…

信息检索 · 计算机科学 2020-04-07 Konstantin Avrachenkov , Kishor Patil , Gugan Thoppe
‹ 上一页 1 2 3 10 下一页 ›