English
Related papers

Related papers: First in the Web, but Where are the Pieces?

200 papers

In this paper we study the prevalence of unique entity identifiers on the Web. These are, e.g., ISBNs (for books), GTINs (for commercial products), DOIs (for documents), email addresses, and others. We show how these identifiers can be…

Databases · Computer Science 2016-07-19 Aliaksandr Talaika , Joanna Biega , Antoine Amarilli , Fabian M. Suchanek

One of the first pre-processing steps for constructing web-scale LLM pretraining datasets involves extracting text from HTML. Despite the immense diversity of web content, existing open-source datasets predominantly apply a single fixed…

Annotation graphs and annotation servers offer infrastructure to support the analysis of human language resources in the form of time-series data such as text, audio and video. This paper outlines areas of common need among empirical…

Computation and Language · Computer Science 2007-05-23 Christopher Cieri , Steven Bird

Web archive collections are created with a particular purpose in mind. A curator selects seeds, or original resources, which are then captured by an archiving system and stored as archived web pages, or mementos. The systems that build web…

Digital Libraries · Computer Science 2021-01-26 Shawn M. Jones , Michele C. Weigle , Michael L. Nelson

Many web-search queries serve as the beginning of an exploration of an unknown space of information, rather than looking for a specific web page. To answer such queries effec- tively, the search engine should attempt to organize the space…

Information Retrieval · Computer Science 2014-01-17 Fei Wu , Jayant Madhavan , Alon Halevy

The first step in a science project is the acquisition and understanding of the relevant data. This paper outlines the results of a project to design and test network tools specifically oriented at retrieving astronomical data. The tools…

Instrumentation and Methods for Astrophysics · Physics 2009-10-07 James Schombert

This short paper gives an introduction to a research project to analyze how digital documents are structured and described. Using a phenomenological approach, this research will reveal common patterns that are used in data, independent from…

Digital Libraries · Computer Science 2014-08-12 Jakob Voß

This paper presents a link analysis approach for identifying privileged documents by constructing a network of human entities derived from email header metadata. Entities are classified as either counsel or non-counsel based on a predefined…

Information Retrieval · Computer Science 2025-12-10 Jianping Zhang , Han Qin , Nathaniel Huber-Fliflet

Many recent discoveries in astrophysics involve phenomena that are highly complex. Carefully designed experiments, together with sophisticated computer simulations, are required to gain insights into the underlying physics. We show that…

Astrophysics · Physics 2017-08-23 Johnny S. T. Ng

Knowledge discovery is defined as non-trivial extraction of implicit, previously unknown and potentially useful information from given data. Knowledge extraction from web documents deals with unstructured, free-format documents whose number…

Neural and Evolutionary Computing · Computer Science 2007-05-23 Vitaly Schetinin

The large set of technical documentation of legacy accelerator systems, coupled with the retirement of experienced personnel, underscores the urgent need for efficient methods to preserve and transfer specialized knowledge. This paper…

Information Retrieval · Computer Science 2025-09-03 Qing Dai , Rasmus Ischebeck , Maruisz Sapinski , Adam Grycner

This study provides a conceptual overview of the literature dealing with the process of citing documents (focusing on the literature from the recent decade). It presents theories, which have been proposed for explaining the citation…

Digital Libraries · Computer Science 2018-05-07 Iman Tahamtan , Lutz Bornmann

Line separators are used to segregate text-lines from one another in document image analysis. Finding the separator points at every line terminal in a document image would enable text-line segmentation. In particular, identifying the…

Computer Vision and Pattern Recognition · Computer Science 2017-08-21 Amarnath R , P. Nagabhushan

Web archives are large longitudinal collections that store webpages from the past, which might be missing on the current live Web. Consequently, temporal search over such collections is essential for finding prominent missing webpages and…

Information Retrieval · Computer Science 2017-02-07 Helge Holzmann , Wolfgang Nejdl , Avishek Anand

We present 1-Pager the first system that answers a question and retrieves evidence using a single Transformer-based model and decoding process. 1-Pager incrementally partitions the retrieval corpus using constrained decoding to select a…

Computation and Language · Computer Science 2023-10-26 Palak Jain , Livio Baldini Soares , Tom Kwiatkowski

Machine Learning software documentation is different from most of the documentations that were studied in software engineering research. Often, the users of these documentations are not software experts. The increasing interest in using…

Software Engineering · Computer Science 2020-02-03 Yalda Hashemi , Maleknaz Nayebi , Giuliano Antoniol

Active Internet measurement studies rely on a list of targets to be scanned. While probing the entire IPv4 address space is feasible for scans of limited complexity, more complex scans do not scale to measuring the full Internet. Thus, a…

Networking and Internet Architecture · Computer Science 2018-02-09 Quirin Scheitle , Jonas Jelten , Oliver Hohlfeld , Luca Ciprian , Georg Carle

Over the past decade, astronomers have been using an increasingly larger number of web-based applications and archives to conduct their research. However, despite the early success in creating links across projects and data centers, the…

Instrumentation and Methods for Astrophysics · Physics 2015-03-17 Alberto Accomazzi , Michael J. Kurtz , Stephen S. Murray

The ever increasing prevalence of publicly available structured data on the World Wide Web enables new applications in a variety of domains. In this paper, we provide a conceptual approach that leverages such data in order to explain the…

Artificial Intelligence · Computer Science 2017-10-13 Md Kamruzzaman Sarker , Ning Xie , Derek Doran , Michael Raymer , Pascal Hitzler

Wikipedia is a rich and invaluable source of information. Its central place on the Web makes it a particularly interesting object of study for scientists. Researchers from different domains used various complex datasets related to Wikipedia…

Information Retrieval · Computer Science 2019-03-21 Nicolas Aspert , Volodymyr Miz , Benjamin Ricaud , Pierre Vandergheynst
‹ Prev 1 8 9 10 Next ›