中文
相关论文

相关论文: Reconstructing Websites for the Lazy Webmaster

200 篇论文

A large amount of data on the WWW remains inaccessible to crawlers of Web search engines because it can only be exposed on demand as users fill out and submit forms. The Hidden web refers to the collection of Web data which can be accessed…

信息检索 · 计算机科学 2014-07-23 Sonali Gupta , Komal Kumar Bhatia

Large language models are trained on massive scrapes of the web, which are often unstructured, noisy, and poorly phrased. Current scaling laws show that learning from such data requires an abundance of both compute and data, which grows…

计算与语言 · 计算机科学 2024-01-30 Pratyush Maini , Skyler Seto , He Bai , David Grangier , Yizhe Zhang , Navdeep Jaitly

The emergence of Linked Data on the WWW has spawned research interest in an online execution of declarative queries over this data. A particularly interesting approach is traversal-based query execution which fetches data by traversing data…

数据库 · 计算机科学 2016-07-06 Olaf Hartig , M. Tamer Özsu

The WWW is the most important source of information. But, there is no guarantee for information correctness and lots of conflicting information is retrieved by the search engines and the quality of provided information also varies from low…

信息检索 · 计算机科学 2009-11-18 Sumalatha Ramachandran , Sujaya Paulraj , Sharon Joseph , Vetriselvi Ramaraj

Archiving Web pages into themed collections is a method for ensuring these resources are available for posterity. Services such as Archive-It exists to allow institutions to develop, curate, and preserve collections of Web resources.…

数字图书馆 · 计算机科学 2017-05-18 Yasmin AlNoamany , Michele C. Weigle , Michael L. Nelson

Increasingly more data is becoming available on the Web, estimates speaking of 1 billion documents in 2002. Most of the documents are Web pages whose data is considered to be in XML format, expecting it to eventually replace HTML. A common…

数据库 · 计算机科学 2007-05-23 Martin Bernauer

Popular web pages are archived frequently, which makes it difficult to visualize the progression of the site through the years at web archives. The What Did It Look Like (WDILL) Twitter bot shows web page transitions by creating a timelapse…

数字图书馆 · 计算机科学 2021-04-30 Dhruv Patel , Alexander C. Nwala , Michael L. Nelson , Michele C. Weigle

Small organizations, start ups, and self-hosted servers face increasing strain from automated web crawlers and AI bots, whose online presence has increased dramatically in the past few years. Modern bots evade traditional throttling and can…

密码学与安全 · 计算机科学 2025-08-06 Rama Carl Hoetzlein

Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model weaknesses -- and even expert-curated challenge sets quickly…

计算与语言 · 计算机科学 2026-05-27 Wenda Xu , Vilém Zouhar , Parker Riley , Mara Finkelstein , Markus Freitag , Daniel Deutsch

The technology of automatic document summarization is maturing and may provide a solution to the information overload problem. Nowadays, document summarization plays an important role in information retrieval. With a large volume of…

信息检索 · 计算机科学 2012-04-10 Mohsen Pourvali , Mohammad Saniee Abadeh

Efficiently discovering relevant Web services with respect to a specific user query has become a growing challenge owing to the incredible growth in the field of web technologies. In previous works, different clustering models have been…

机器学习 · 计算机科学 2022-10-05 Anirudha Rayasam , Siddhartha R Thota , Avinash N Bukkittu , Sowmya Kamath

When a network is attacked, cyber defenders need to precisely identify which systems (i.e., computers or devices) were compromised and what damage may have been inflicted. This process is sometimes referred to as cyber triage and is an…

密码学与安全 · 计算机科学 2024-09-18 Eric Ficke , Raymond M. Bateman , Shouhuai Xu

Web archiving frameworks are commonly assessed by the quality of their archival records and by their ability to operate at scale. The ubiquity of dynamic web content poses a significant challenge for crawler-based solutions such as the…

数字图书馆 · 计算机科学 2019-09-11 Martin Klein , Harihar Shankar , Lyudmila Balakireva , Herbert Van de Sompel

Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in…

数字图书馆 · 计算机科学 2017-07-31 Gerhard Gossen , Elena Demidova , Thomas Risse

Automated process discovery is a class of process mining methods that allow analysts to extract business process models from event logs. Traditional process discovery methods extract process models from a snapshot of an event log stored in…

机器学习 · 计算机科学 2018-04-10 Volodymyr Leno , Abel Armas-Cervantes , Marlon Dumas , Marcello La Rosa , Fabrizio M. Maggi

Recovering from crises, such as hurricanes or wildfires, is a complex process that can take weeks, months, or even decades to overcome. Crises have both acute (immediate) and chronic (long-term) effects on communities. Crisis informatics…

社会与信息网络 · 计算机科学 2025-03-12 Casey Randazzo , Minkyung Kim , Melanie Kwestel , Marya L Doerfel , Tawfiq Ammari

Web cache deception (WCD) is an attack proposed in 2017, where an attacker tricks a caching proxy into erroneously storing private information transmitted over the Internet and subsequently gains unauthorized access to that cached data. Due…

密码学与安全 · 计算机科学 2020-02-17 Seyed Ali Mirheidari , Sajjad Arshad , Kaan Onarlioglu , Bruno Crispo , Engin Kirda , William Robertson

Password updates are a critical account security measure and an essential part of the password lifecycle. Service providers and common security recommendations advise users to update their passwords in response to incidents or as a critical…

密码学与安全 · 计算机科学 2025-11-17 Alexander Krause , Jacques Suray , Lea Schmüser , Marten Oltrogge , Oliver Wiese , Maximilian Golla , Sascha Fahl

Linguistic theories formulated in the architecture of {\sc hpsg} can be very precise and explicit since {\sc hpsg} provides a formally well-defined setup. However, when querying a faithful implementation of such an explicit theory, the…

cmp-lg · 计算机科学 2008-02-03 Thilo Götz , Walt Detmar Meurers

Internet search engines function in a present which changes continuously. The search engines update their indices regularly, overwriting Web pages with newer ones, adding new pages to the index, and losing older ones. Some search engines…

信息检索 · 计算机科学 2009-11-19 Iina Hellsten , Loet Leydesdorff , Paul Wouters