Related papers: Updates of PDFs in the MSTW framework
Research data and software are widely accepted as an outcome of scientific work. However, in comparison to text-based publications, there is not yet an established process to assess and evaluate quality of research data and research…
We discuss the main features of the recent NNPDF2.1 NLO set, a determination of parton distributions from a global set of hard scattering data using the NNPDF methodology including heavy quark mass effects. We present the implications for…
The xFitter project (former HERAFitter project) is an open-source package that provides a framework for the determination of the parton distribution functions (PDFs) of the proton for many different kinds of analyses in Quantum…
The majority of scientific papers are distributed in PDF, which pose challenges for accessibility, especially for blind and low vision (BLV) readers. We characterize the scope of this problem by assessing the accessibility of 11,397 PDFs…
We discuss implementation of the LHC experimental data sets in the new CT18 global analysis of quantum chromodynamics (QCD) at the next-to-next-leading order of the QCD coupling strength. New methodological developments in the fitting…
Understanding and extracting of information from large documents, such as business opportunities, academic articles, medical documents and technical reports, poses challenges not present in short documents. Such large documents may be…
We survey recent developments in the theory of achievement sets and present a substantial collection of open problems.
PDFs remain the dominant format for scholarly communication, despite significant accessibility challenges for blind and low-vision users. While various tools attempt to evaluate PDF accessibility, there is no standardized methodology to…
In this project, we semantically enriched and enhanced the metadata of long text documents, theses and dissertations, retrieved from the HathiTrust Digital Library in English published from 1920 to 2020 through a combination of manual…
Recently, there has been a growing interest among large language model (LLM) developers in LLM-based document reading systems, which enable users to upload their own documents and pose questions related to the document contents, going…
This technical report introduces Docling, an easy to use, self-contained, MIT-licensed open-source package for PDF document conversion. It is powered by state-of-the-art specialized AI models for layout analysis (DocLayNet) and table…
Improving data quality in unstructured documents is a long-standing challenge. Unstructured data, especially in textual form, inherently lacks defined semantics, which poses significant challenges for effective processing and for ensuring…
The LHC has recently released precise measurements of the transverse momentum distribution of the Z-boson that provide a unique constraint on the structure of the proton. Theoretical developments now allow the prediction of these…
The present note relies on the recently published conceptual design report of the LHeC and extends the first contribution to the European strategy debate in emphasising the role of the LHeC to complement and complete the high luminosity LHC…
We compare predictions of nCTEQ15 nuclear parton distribution functions with proton-lead vector boson production data from the LHC. We select data sets that are most sensitive to nuclear PDFs and have potential to constrain them. We…
Parton Distribution Functions (PDFs) contribute significantly to the uncertainty on the determination of the top-quark pole mass and other precision measurements at the Large Hadron Collider (LHC). It is crucial to understand these…
Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works…
As software systems become more complex, modern software development requires more attention to human perspectives, and active participation of development teams in requirements elicitation tasks. In this context, incomplete or ambiguous…
PDF fits in the HERAPDF1.0 formalism have been made to the combined HERA-I inclusive data and the newly combined $F_2$(charm) data from the H1 and ZEUS experiments. The charm data are found to be sensitive to the value of the charm mass and…
Even for a conservative estimate, 80% of enterprise data reside in unstructured files, stored in data lakes that accommodate heterogeneous formats. Classical search engines can no longer meet information seeking needs, especially when the…