中文
相关论文

相关论文: The Many Shapes of Archive-It

200 篇论文

In theory, a major advantage to the big data approach in studying online communities is that it should be possible to collect a representative random sample from a broadly defined population. However, in practice, data collection processes…

社会与信息网络 · 计算机科学 2021-02-02 Muhammad Umer Gurchani

From popular uprisings to pandemics, the Web is an essential source consulted by scientists and historians for reconstructing and studying past events. Unfortunately, the Web is plagued by reference rot which causes important Web resources…

数字图书馆 · 计算机科学 2021-07-07 Alexander C. Nwala , Michele C. Weigle , Michael L. Nelson

Used by a variety of researchers, web archive collections have become invaluable sources of evidence. If a researcher is presented with a web archive collection that they did not create, how do they know what is inside so that they can use…

数字图书馆 · 计算机科学 2021-01-26 Shawn M. Jones , Michele C. Weigle , Michael L. Nelson

For traditional library collections, archivists can select a representative sample from a collection and display it in a featured physical or digital library space. Web archive collections may consist of thousands of archived pages, or…

数字图书馆 · 计算机科学 2021-03-23 Shawn M. Jones , Martin Klein , Michele C. Weigle , Michael L. Nelson

Web archives preserve unique and historically valuable information. They hold a record of past events and memories published by all kinds of people, such as journalists, politicians and ordinary people who have shared their testimony and…

数字图书馆 · 计算机科学 2021-08-04 Miguel Costa , Julien Masanès

Web archiving frameworks are commonly assessed by the quality of their archival records and by their ability to operate at scale. The ubiquity of dynamic web content poses a significant challenge for crawler-based solutions such as the…

数字图书馆 · 计算机科学 2019-09-11 Martin Klein , Harihar Shankar , Lyudmila Balakireva , Herbert Van de Sompel

Memento aggregators enable users to query multiple web archives for captures of a URI in time through a single HTTP endpoint. While this one-to-many access point is useful for researchers and end-users, aggregators are in a position to…

数字图书馆 · 计算机科学 2023-01-10 Mat Kelly

The humanities, like many other areas of society, are currently undergoing major changes in the wake of digital transformation. However, in order to make collection of digitised material in this area easily accessible, we often still lack…

信息检索 · 计算机科学 2021-03-23 Vuong M. Ngo , Sven Helmer , Nhien-An Le-Khac , M-Tahar Kechadi

Many social Web sites allow users to annotate the content with descriptive metadata, such as tags, and more recently to organize content hierarchically. These types of structured metadata provide valuable evidence for learning how a…

人工智能 · 计算机科学 2010-05-28 Anon Plangprasopchok , Kristina Lerman , Lise Getoor

Text categorization is an essential task in Web content analysis. Considering the ever-evolving Web data and new emerging categories, instead of the laborious supervised setting, in this paper, we focus on the minimally-supervised setting…

计算与语言 · 计算机科学 2021-02-24 Xinyang Zhang , Chenwei Zhang , Luna Xin Dong , Jingbo Shang , Jiawei Han

Personal and private Web archives are proliferating due to the increase in the tools to create them and the realization that Internet Archive and other public Web archives are unable to capture personalized (e.g., Facebook) and private…

数字图书馆 · 计算机科学 2018-06-05 Mat Kelly , Michael L. Nelson , Michele C. Weigle

The new era of the Web is known as the semantic Web or the Web of data. The semantic Web depends on ontologies that are seen as one of its pillars. The bigger these ontologies, the greater their exploitation. However, when these ontologies…

人工智能 · 计算机科学 2017-09-26 Noreddine Gherabi , Redouane Nejjahi , Abderrahim Marzouk

In Social Science research, multimedia documents are often collected to answer particular research questions like: "Which of the aesthetic properties of a photo are considered important on the web" or "How has Street Art developed over the…

信息检索 · 计算机科学 2015-03-13 Zeon Trevor Fernando

Modern machine learning research relies on relatively few carefully curated datasets. Even in these datasets, and typically in `untidy' or raw data, practitioners are faced with significant issues of data quality and diversity which can be…

机器学习 · 计算机科学 2022-09-22 Shoaib Ahmed Siddiqui , Nitarshan Rajkumar , Tegan Maharaj , David Krueger , Sara Hooker

With the huge amount of information available online, the World Wide Web is a fertile area for data mining research. The Web mining research is at the cross road of research from several research communities, such as database, information…

机器学习 · 计算机科学 2007-05-23 Raymond Kosala , Hendrik Blockeel

Research data are often released upon journal publication to enable result verification and reproducibility. For that reason, research dissemination infrastructures typically support diverse datasets coming from numerous disciplines, from…

数字图书馆 · 计算机科学 2023-05-29 Ana Trisovic

The Memento aggregator currently polls every known public web archive when serving a request for an archived web page, even though some web archives focus on only specific domains and ignore the others. Similar to query routing in…

数字图书馆 · 计算机科学 2013-09-17 Ahmed AlSum , Michele C. Weigle , Michael L. Nelson , Herbert Van de Sompel

The Web is ephemeral. Many resources have representations that change over time, and many of those representations are lost forever. A lucky few manage to reappear as archived resources that carry their own URIs. For example, some content…

The web is today's primary publication medium, making web archiving an important activity for historical and analytical purposes. Web pages are increasingly interactive, resulting in pages that are increasingly difficult to archive.…

数字图书馆 · 计算机科学 2016-01-21 Justin F. Brunelle , Michele C. Weigle , Michael L. Nelson

As web technologies evolve, web archivists work to keep up so that our digital history is preserved. Recent advances in web technologies have introduced client-side executed scripts that load data without a referential identifier or that…

数字图书馆 · 计算机科学 2019-05-17 Mat Kelly , Justin F. Brunelle , Michele C. Weigle , Michael L. Nelson