中文
相关论文

相关论文: Scaling Accessible Mathematics on arXiv: HTML Conv…

200 篇论文

Standing at the forefront of knowledge dissemination, digital libraries curate vast collections of scientific literature. However, these scholarly writings are often laden with jargon and tailored for domain experts rather than the general…

计算与语言 · 计算机科学 2024-08-08 Haining Wang , Jason Clark

Recent advances in large language models (LLMs) have enabled AI agents to autonomously generate scientific proposals, conduct experiments, author papers, and perform peer reviews. Yet this flood of AI-generated research content collides…

An analysis of 2,765 articles published in four math journals from 1997 to 2005 indicate that articles deposited in the arXiv received 35% more citations on average than non-deposited articles (an advantage of about 1.1 citations per…

数字图书馆 · 计算机科学 2007-05-23 Philip M. Davis , Michael J. Fromerth

We propose to demonstrate LiquidXML, a platform for managing large corpora of XML documents in large-scale P2P networks. All LiquidXML peers may publish XML documents to be shared with all the network peers. The challenge then is to…

The large-scale training of multi-modal models on data scraped from the web has shown outstanding utility in infusing these models with the required world knowledge to perform effectively on multiple downstream tasks. However, one downside…

For blind and low-vision (BLV) individuals, digital math communication is uniquely difficult due to the lack of accessible tools. Currently, the state of the art is either code-based, like LaTeX, or WYSIWYG, like visual editors. However,…

人机交互 · 计算机科学 2026-03-17 Kenneth Ge , JooYoung Seo

The NLLG (Natural Language Learning & Generation) arXiv reports assist in navigating the rapidly evolving landscape of NLP and AI research across cs.CL, cs.CV, cs.AI, and cs.LG categories. This fourth installment captures a transformative…

数字图书馆 · 计算机科学 2024-12-18 Christoph Leiter , Jonas Belouadi , Yanran Chen , Ran Zhang , Daniil Larionov , Aida Kostikova , Steffen Eger

The majority of scientific papers are distributed in PDF, which pose challenges for accessibility, especially for blind and low vision (BLV) readers. We characterize the scope of this problem by assessing the accessibility of 11,397 PDFs…

In the evolving landscape of clinical informatics, the integration and utilization of software tools developed through governmental funding represent a pivotal advancement in research and application. However, the dispersion of these tools…

数字图书馆 · 计算机科学 2024-03-28 Jeremy R. Harper

Scientific publishing lays the foundation of science by disseminating research findings, fostering collaboration, encouraging reproducibility, and ensuring that scientific knowledge is accessible, verifiable, and built upon over time.…

Large language models (LLMs) are dramatically influencing AI research, spurring discussions on what has changed so far and how to shape the field's future. To clarify such questions, we analyze a new dataset of 16,979 LLM-related arXiv…

数字图书馆 · 计算机科学 2024-04-30 Rajiv Movva , Sidhika Balachandar , Kenny Peng , Gabriel Agostini , Nikhil Garg , Emma Pierson

Most scholarly works are distributed online in PDF format, which can present significant accessibility challenges for blind and low-vision readers. To characterize the scope of this issue, we perform a large-scale analysis of 20K open- and…

数字图书馆 · 计算机科学 2024-10-07 Anukriti Kumar , Lucy Lu Wang

Verifiable formal languages like Lean have profoundly impacted mathematical reasoning, particularly through the use of large language models (LLMs) for automated reasoning. A significant challenge in training LLMs for these formal languages…

计算与语言 · 计算机科学 2025-02-28 Guoxiong Gao , Yutong Wang , Jiedong Jiang , Qi Gao , Zihan Qin , Tianyi Xu , Bin Dong

The use of large language models (LLMs) in scholarly publications has grown dramatically since the launch of ChatGPT in late 2022. This usage is often undisclosed, and it can be challenging for readers and reviewers to identify human…

数字图书馆 · 计算机科学 2025-12-02 Andrew Gray

The rapid growth of information in the field of Generative Artificial Intelligence (AI), particularly in the subfields of Natural Language Processing (NLP) and Machine Learning (ML), presents a significant challenge for researchers and…

计算机与社会 · 计算机科学 2023-08-15 Steffen Eger , Christoph Leiter , Jonas Belouadi , Ran Zhang , Aida Kostikova , Daniil Larionov , Yanran Chen , Vivian Fresen

This paper explores articles hosted on the arXiv preprint server with the aim to uncover valuable insights hidden in this vast collection of research. Employing text mining techniques and through the application of natural language…

数字图书馆 · 计算机科学 2024-04-08 Michele Leonardo Bianchi

This paper describes the intense software filtering that has allowed the arXiv eprint repository to sort and process large numbers of submissions with minimal human intervention, making it one of the most important and influential cases of…

物理与社会 · 物理学 2016-09-27 Luis Reyes-Galindo

XSLT is an increasingly popular language for processing XML data. It is widely supported by application platform software. However, little optimization effort has been made inside the current XSLT processing engines. Evaluating a very…

数据库 · 计算机科学 2007-05-23 Zhimao Guo , Min Li , Xiaoling Wang , Aoying Zhou

Nowadays, Machine Learning (ML) is seen as the universal solution to improve the effectiveness of information retrieval (IR) methods. However, while mathematics is a precise and accurate science, it is usually expressed by less accurate and…

数字图书馆 · 计算机科学 2020-12-22 André Greiner-Petter , Terry Ruas , Moritz Schubotz , Akiko Aizawa , William Grosky , Bela Gipp

The rapid growth of preprint servers has accelerated scientific dissemination but has also shifted the technical burden of manuscript preparation to authors. This challenge is particularly acute in computational research, where manuscripts…

数字图书馆 · 计算机科学 2025-12-19 Bruno M. Saraiva , António D. Brito , Guillaume Jaquemet , Ricardo Henriques