中文
相关论文

相关论文: Toward Robust URL Extraction for Open Science: A S…

200 篇论文

In recent years, funding agencies and journals increasingly advocate for open science practices (e.g. data and method sharing) to improve the transparency, access, and reproducibility of science. However, quantifying these practices at…

数字图书馆 · 计算机科学 2023-10-06 Hancheng Cao , Jesse Dodge , Kyle Lo , Daniel A. McFarland , Lucy Lu Wang

In the evolving landscape of clinical informatics, the integration and utilization of software tools developed through governmental funding represent a pivotal advancement in research and application. However, the dispersion of these tools…

数字图书馆 · 计算机科学 2024-03-28 Jeremy R. Harper

In October 2023, arXiv made HTML formatted papers available to readers. This was the exciting outcome of over a year of accessibility research and development with the scientific community. Currently, only 2.4% of research outputs meet…

数字图书馆 · 计算机科学 2024-02-15 Charles Frankston , Jonathan Godfrey , Shamsi Brinn , Alison Hofer , Mark Nazzaro

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

Metadata of scientific articles such as title, abstract, keywords or index terms, body text, conclusion, reference and others play a decisive role in collecting, managing and storing academic data in scientific databases, academic journals…

信息检索 · 计算机科学 2018-07-25 Jahongir Azimjonov , Jumabek Alikhanov

The research content hosted by arXiv is not fully accessible to everyone due to disabilities and other barriers. This matters because a significant proportion of people have reading and visual disabilities, it is important to our community…

One of the first pre-processing steps for constructing web-scale LLM pretraining datasets involves extracting text from HTML. Despite the immense diversity of web content, existing open-source datasets predominantly apply a single fixed…

Extracting information from academic PDF documents is crucial for numerous indexing, retrieval, and analysis use cases. Choosing the best tool to extract specific content elements is difficult because many, technically diverse tools are…

信息检索 · 计算机科学 2023-03-20 Norman Meuschke , Apurva Jagdale , Timo Spinde , Jelena Mitrović , Bela Gipp

Non-textual components such as charts, diagrams and tables provide key information in many scientific documents, but the lack of large labeled datasets has impeded the development of data-driven methods for scientific figure extraction. In…

数字图书馆 · 计算机科学 2018-06-01 Noah Siegel , Nicholas Lourie , Russell Power , Waleed Ammar

arXiv is the largest open-access repository for scientific literature. When submitting a paper, authors upload the manuscript's source files, from which the final PDF is compiled. These source files are also publicly downloadable,…

网络与互联网体系结构 · 计算机科学 2026-01-19 Giovanni Apruzzese , Aurore Fass

Astrophysics papers often rely on software which may or may not be available, and URLs are often used as proxy citations for software and data. We extracted all URLs from two journals' 2015 research articles, removed those from certain…

天体物理仪器与方法 · 物理学 2019-03-22 P. Wesley Ryan , Alice Allen , Peter Teuben

This study explores three approaches to processing table data in scientific papers to enhance extractive question answering and develop a software tool for the systematic review process. The methods evaluated include: (1) Optical Character…

信息检索 · 计算机科学 2025-08-27 Dongyoun Kim , Hyung-do Choi , Youngsun Jang , John Kim

Many scientific papers such as those in arXiv and PubMed data collections have abstracts with varying lengths of 50-1000 words and average length of approximately 200 words, where longer abstracts typically convey more information about the…

计算与语言 · 计算机科学 2022-06-03 Sajad Sotudeh , Nazli Goharian

In this work, we aim at developing an extractive summarizer in the multi-document setting. We implement a rank based sentence selection using continuous vector representations along with key-phrases. Furthermore, we propose a model to…

计算与语言 · 计算机科学 2020-06-26 Mir Tafseer Nayeem , Yllias Chali

This thesis investigates in the use of access log data as a source of information for identifying related scientific papers. This is done for arXiv.org, the authority for publication of e-prints in several fields of physics. Compared to…

数字图书馆 · 计算机科学 2007-05-23 Stefan Pohl

In this paper we present the results of a study into the persistence and availability of web resources referenced from papers in scholarly repositories. Two repositories with different characteristics, arXiv and the UNT digital library, are…

数字图书馆 · 计算机科学 2011-05-18 Robert Sanderson , Mark Phillips , Herbert Van de Sompel

We describe a strategy for identifying the universe of research publications relevant to the application and development of artificial intelligence. The approach leverages the arXiv corpus of scientific preprints, in which authors choose…

数字图书馆 · 计算机科学 2020-05-29 James Dunham , Jennifer Melot , Dewey Murdick

In this work, we compare two simple methods of tagging scientific publications with labels reflecting their content. As a first source of labels Wikipedia is employed, second label set is constructed from the noun phrases occurring in the…

计算与语言 · 计算机科学 2014-11-04 Michał Łopuszyński , Łukasz Bolikowski

Long-term Web archives comprise Web documents gathered over longer time periods and can easily reach hundreds of terabytes in size. Semantic annotations such as named entities can facilitate intelligent access to the Web archive data.…

信息检索 · 计算机科学 2017-02-03 Tarcisio Souza , Elena Demidova , Thomas Risse , Helge Holzmann , Gerhard Gossen , Julian Szymanski

The availability of metadata for scientific documents is pivotal in propelling scientific knowledge forward and for adhering to the FAIR principles (i.e. Findability, Accessibility, Interoperability, and Reusability) of research findings.…

信息检索 · 计算机科学 2025-01-10 Zeyd Boukhers , Cong Yang
‹ 上一页 1 2 3 10 下一页 ›