English
Related papers

Related papers: PDF/A-3u as an archival format for Accessible math…

200 papers

Text documents are complex high dimensional objects. To effectively visualize such data it is important to reduce its dimensionality and visualize the low dimensional embedding as a 2-D or 3-D scatter plot. In this paper we explore…

Computation and Language · Computer Science 2010-03-03 Yi Mao , Krishnakumar Balasubramanian , Guy Lebanon

Arithmetic Coding is an efficient lossless compression scheme applied for many multimedia standards such as JPEG, JPEG2000, H.263, H.264 and H.265. Due to nonlinearity, high error propagation and high error sensitivity of arithmetic coders,…

Cryptography and Security · Computer Science 2018-04-13 Alaa Eldin Rohiem Shehata , Hassan Yakout El-Arsh

PDFs can be studied directly using lattice QCD by evaluating matrix elements of non-local operators. A number of groups are pursuing numerical calculations and investigating possible systematic uncertainties. One systematic that has…

High Energy Physics - Lattice · Physics 2018-11-08 Raúl A. Briceño , Juan V. Guerrero , Maxwell T. Hansen , Christopher J. Monahan

This survey analyzed the quality of the PDF documents on online repositories in Switzerland, examining their accessibility for people with visual impairments. Two minimal accessibility features were analyzed: the PDFs had to have tags and a…

Digital Libraries · Computer Science 2023-06-16 Alireza Darvishy , Rolf Sethe , Ines Engler , Oriane Pierres , Juliet Manning

Product sustainability reports provide valuable insights into the environmental impacts of a product and are often distributed in PDF format. These reports often include a combination of tables and text, which complicates their analysis.…

Computation and Language · Computer Science 2025-08-06 Kaiwen Zhao , Bharathan Balaji , Stephen Lee

Document subject classification is essential for structuring (digital) libraries and allowing readers to search within a specific field. Currently, the classification is typically made by human domain experts. Semi-supervised Machine…

Information Retrieval · Computer Science 2021-09-03 Philipp Scharpf , Moritz Schubotz , Bela Gipp

This thesis investigates in the use of access log data as a source of information for identifying related scientific papers. This is done for arXiv.org, the authority for publication of e-prints in several fields of physics. Compared to…

Digital Libraries · Computer Science 2007-05-23 Stefan Pohl

This paper describes the KnowledgeHub tool, a scientific literature Information Extraction (IE) and Question Answering (QA) pipeline. This is achieved by supporting the ingestion of PDF documents that are converted to text and structured…

Identifying academic plagiarism is a pressing problem, among others, for research institutions, publishers, and funding organizations. Detection approaches proposed so far analyze lexical, syntactical, and semantic text similarity. These…

Information Retrieval · Computer Science 2021-06-11 Norman Meuschke

Text documents, including programs, typically have human-readable semantic structure. Historically, programmatic access to these semantics has required explicit in-document tagging. Especially in systems where the text has an execution…

Computation and Language · Computer Science 2024-03-07 Edward Misback , Zachary Tatlock , Steven L. Tanimoto

The number of published PDF documents has increased exponentially in recent decades. There is a growing need to make their rich content discoverable to information retrieval tools. In this paper, we present a novel approach to document…

In this work, we explored data acquisition protocols and improved data reduction protocols using PDFgetX3 to obtain reliable data for atomic pair distribution function (PDF) analysis from a laboratory-based Mo x-ray source. A variable…

In this paper, we investigate mathematical content representations suitable for the automated classification of and the similarity search in STEM documents using standard machine learning algorithms: the Latent Dirichlet Allocation (LDA)…

Information Retrieval · Computer Science 2021-10-11 Michal Růžička , Petr Sojka

In the paper, we present the ADD-Lib, our efficient and easy to use framework for Algebraic Decision Diagrams (ADDs). The focus of the ADD-Lib is not so much on its efficient implementation of individual operations, which are taken by other…

Machine Learning · Computer Science 2020-02-18 Frederik Gossen , Alnis Murtovi , Philip Zweihoff , Bernhard Steffen

Mathematical formulae represent complex semantic information in a concise form. Especially in Science, Technology, Engineering, and Mathematics, mathematical formulae are crucial to communicate information, e.g., in scientific papers, and…

Digital Libraries · Computer Science 2018-04-16 Moritz Schubotz , Andre Greiner-Petter , Philipp Scharpf , Norman Meuschke , Howard Cohl , Bela Gipp

This paper introduces the concept of accessibility from the field of transportation planning and adopts it within the context of Information Retrieval (IR). An analogy is drawn between the fields, which motivates the development of document…

Information Retrieval · Computer Science 2024-04-15 Leif Azzopardi , Vishwa Vinay

This demo paper presents the first tool to annotate the reuse of text, images, and mathematical formulae in a document pair -- TEIMMA. Annotating content reuse is particularly useful to develop plagiarism detection algorithms. Real-world…

Information Retrieval · Computer Science 2023-06-14 Ankit Satpute , André Greiner-Petter , Moritz Schubotz , Norman Meuschke , Akiko Aizawa , Olaf Teschke , Bela Gipp

This paper reports the use of a document distance-based approach to automatically expand the number of available relevance judgements when these are limited and reduced to only positive judgements. This may happen, for example, when the…

Information Retrieval · Computer Science 2015-01-27 Diego Mollá , Iman Amini , David Martinez

We present a framework consisting of three approaches that can enhance meta-analyses: 1) scoping reviews (evidence map), 2) bibliometrics, and 3) alternative impact metrics. These three "enrichment" approaches facilitate the research…

Digital Libraries · Computer Science 2023-12-05 Yefeng Yang , Malgorzata Lagisz , Shinichi Nakagawa

Retrieved documents containing noise will hinder RAG from detecting answer clues and make the inference process slow and expensive. Therefore, context compression is necessary to enhance its accuracy and efficiency. Existing context…

Computation and Language · Computer Science 2026-04-28 Qianchi Zhang , Hainan Zhang , Liang Pang , Hongwei Zheng , Zhiming Zheng
‹ Prev 1 3 4 5 6 7 10 Next ›