中文
相关论文

相关论文: The Many Shapes of Archive-It

200 篇论文

Understanding and extracting of information from large documents, such as business opportunities, academic articles, medical documents and technical reports, poses challenges not present in short documents. Such large documents may be…

计算与语言 · 计算机科学 2019-10-10 Muhammad Mahbubur Rahman , Tim Finin

In this paper we present the results of a study into the persistence and availability of web resources referenced from papers in scholarly repositories. Two repositories with different characteristics, arXiv and the UNT digital library, are…

数字图书馆 · 计算机科学 2011-05-18 Robert Sanderson , Mark Phillips , Herbert Van de Sompel

In our daily lives, organizing resources into a set of categories is a common task. Categorization becomes more useful as the collection of resources increases. Large collections of books, movies, and web pages, for instance, are cataloged…

数字图书馆 · 计算机科学 2012-05-01 Arkaitz Zubiaga

The application of unsupervised learning approaches, and in particular of clustering techniques, represents a powerful exploration means for the analysis of network measurements. Discovering underlying data characteristics, grouping similar…

人工智能 · 计算机科学 2020-03-11 Andrea Morichetta , Pedro Casas , Marco Mellia

Instead of mining coherent topics from a given text corpus in a completely unsupervised manner, seed-guided topic discovery methods leverage user-provided seed words to extract distinctive and coherent topics so that the mined topics can…

计算与语言 · 计算机科学 2023-01-12 Yu Zhang , Yunyi Zhang , Martin Michalski , Yucheng Jiang , Yu Meng , Jiawei Han

Structured prediction problems are one of the fundamental tools in machine learning. In order to facilitate algorithm development for their numerical solution, we collect in one place a large number of datasets in easy to read formats for a…

The information retrieval (IR) community has a strong tradition of making the computational artifacts and resources available for future reuse, allowing the validation of experimental results. Besides the actual test collections, the…

信息检索 · 计算机科学 2022-07-20 Timo Breuer , Jüri Keller , Philipp Schaer

Argument mining automatically identifies and extracts the structure of inference and reasoning conveyed in natural language arguments. To the best of our knowledge, most of the state-of-the-art works in this field have focused on using…

计算与语言 · 计算机科学 2023-02-28 Pranjal Srivastava , Pranav Bhatnagar , Anurag Goel

In the scientific digital libraries, some papers from different research communities can be described by community-dependent keywords even if they share a semantically similar topic. Articles that are not tagged with enough keyword…

数字图书馆 · 计算机科学 2018-06-22 Hussein T. Al-Natsheh , Lucie Martinet , Fabrice Muhlenbach , Fabien Rico , Djamel A. Zighed

Retrieval-Augmented Generation systems depend on retrieving semantically relevant document chunks to support accurate, grounded outputs from large language models. In structured and repetitive corpora such as regulatory filings, chunk…

信息检索 · 计算机科学 2026-01-21 Raquib Bin Yousuf , Shengzhe Xu , Mandar Sharma , Andrew Neeser , Chris Latimer , Naren Ramakrishnan

Text extraction from web pages has many applications, including web crawling optimization and document clustering. Though much has been written about the acquisition of content from live web pages, content acquisition of archived web pages,…

数字图书馆 · 计算机科学 2016-02-24 Shawn M. Jones , Harihar Shankar

Web archives have grown to petabytes. In addition to providing invaluable background knowledge on many social and cultural developments over the last 30 years, they also provide vast amounts of training data for machine learning. To benefit…

数字图书馆 · 计算机科学 2022-09-27 Niklas Deckers , Martin Potthast

The application of semantic technologies to content on the web is, in many regards, important and urgent. Search engines, chatbots, intelligent personal assistants and other technologies increasingly rely on content published as semantic…

信息检索 · 计算机科学 2017-10-03 Elias Kärle , Umutcan Şimşek , Dieter Fensel

We present a method for mapping Reddit communities that accounts for temporal shifts, using quantitative and qualitative analyses of clustering techniques to produce high-quality, stable, and meaningful maps for researchers, journalists and…

社会与信息网络 · 计算机科学 2024-10-15 Virginia Partridge , Jasmine Mangat , Rebecca Curran , Ryan McGrady , Ethan Zuckerman

The Web is a typical example of a social network. One of the most intriguing features of the Web is its self-organization behavior, which is usually faced through the existence of communities. The discovery of the communities in a Web-graph…

信息检索 · 计算机科学 2015-03-20 Antonis Sidiropoulos

In this work we propose MementoMap, a flexible and adaptive framework to efficiently summarize holdings of a web archive. We described a simple, yet extensible, file format suitable for MementoMap. We used the complete index of the…

数字图书馆 · 计算机科学 2019-05-30 Sawood Alam , Michele C. Weigle , Michael L. Nelson , Fernando Melo , Daniel Bicho , Daniel Gomes

Understanding and analyzing big data is firmly recognized as a powerful and strategic priority. For deeper interpretation of and better intelligence with big data, it is important to transform raw data (unstructured, semi-structured and…

信息检索 · 计算机科学 2016-12-13 Seyed-Mehdi-Reza Beheshti , Alireza Tabebordbar , Boualem Benatallah , Reza Nouri

The creation of open archives i.e. archives where access is regulated by open licensing models (content, source, data), should be seen as part of a broader socio-economic phenomenon that finds legal expression in specific organizational and…

数字图书馆 · 计算机科学 2011-09-06 Prodromos Tsiavos , Petros Stefaneas

Web archives are typically very broad in scope and extremely large in scale. This makes data analysis appear daunting, especially for non-computer scientists. These collections constitute an increasingly important source for researchers in…

数字图书馆 · 计算机科学 2017-07-31 Gerhard Gossen , Elena Demidova , Thomas Risse

The World Wide Web is the most wide known information source that is easily available and searchable. It consists of billions of interconnected documents Web pages are authored by millions of people. Accesses made by various users to pages…

数据库 · 计算机科学 2014-08-26 Priyanka Verma , Nishtha Kesswani