中文
相关论文

相关论文: Self-Driving Datasets: From 20 Million Papers to N…

200 篇论文

The increasing volume and complexity of scientific literature demand robust methods for organizing and understanding research documents. In this study, we investigate whether structured knowledge, specifically, subject-predicate-object…

计算与语言 · 计算机科学 2026-04-21 Mihael Arcan

Biological data and knowledge bases increasingly rely on Semantic Web technologies and the use of knowledge graphs for data integration, retrieval and federated queries. We propose a solution for automatically semantifying biological…

人工智能 · 计算机科学 2021-12-03 Marco Anteghini , Jennifer D'Souza , Vitor A. P. Martins dos Santos , Sören Auer

Computational reproducibility is fundamental to trustworthy science, yet remains difficult to achieve in practice across various research workflows, including Jupyter notebooks published alongside scholarly articles. Environment drift,…

软件工程 · 计算机科学 2026-04-02 Sheeba Samuel , Daniel Mietchen , Hemanta Lo , Martin Gaedke

Electronic health records contain inconsistently structured or free-text data, requiring efficient preprocessing to enable predictive health care models. Although artificial intelligence-driven natural language processing tools show promise…

Corpus linguistics has traditionally relied on human researchers to formulate hypotheses, construct queries, and interpret results - a process demanding specialized technical skills and considerable time. We propose Agent-Driven Corpus…

计算与语言 · 计算机科学 2026-04-09 Jia Yu , Weiwei Yu , Pengfei Xiao , Fukun Xing

The development of deep segmentation models for computational pathology (CPath) can help foster the investigation of interpretable morphological biomarkers. Yet, there is a major bottleneck in the success of such approaches because…

We introduce Nomad, a system for autonomous data exploration and insight discovery. Given a corpus of documents, databases, or other data sources, users rarely know the full set of questions, hypotheses, or connections that could be…

人工智能 · 计算机科学 2026-04-03 Bokang Jia , Samta Kamboj , Satheesh Katipomu , Seung Hun Han , Neha Sengupta , Andrew Jackson

Most existing large-scale academic search engines are built to retrieve text-based information. However, there are no large-scale retrieval services for scientific figures and tables. One challenge for such services is understanding…

人工智能 · 计算机科学 2023-01-31 Zeba Karishma , Shaurya Rohatgi , Kavya Shrinivas Puranik , Jian Wu , C. Lee Giles

Large Language Models (LLMs) have demonstrated remarkable performance across various domains, including healthcare. However, their ability to effectively represent structured non-textual data, such as the alphanumeric medical codes used in…

Clinical narratives encode temporal dynamics essential for modeling patient trajectories, yet large-scale temporally annotated resources are scarce. We introduce PMOA-TTS, a corpus of 124,699 single-patient PubMed Open Access case reports…

计算与语言 · 计算机科学 2026-01-16 Shahriar Noroozizadeh , Sayantan Kumar , George H. Chen , Jeremy C. Weiss

Citation graphs can be helpful in generating high-quality summaries of scientific papers, where references of a scientific paper and their correlations can provide additional knowledge for contextualising its background and main…

信息检索 · 计算机科学 2023-02-24 Zheheng Luo , Qianqian Xie , Sophia Ananiadou

Research profiles highlight scientists' research focus, enabling talent discovery and collaborations, but are often outdated. Automated, scalable methods are urgently needed to keep profiles current. We design and evaluate two Large…

计算与语言 · 计算机科学 2026-01-07 Yilun Liang , Gongbo Zhang , Edward Sun , Betina Idnay , Yilu Fang , Fangyi Chen , Casey Ta , Yifan Peng , Chunhua Weng

Pre-trained language models induce dense entity representations that offer strong performance on entity-centric NLP tasks, but such representations are not immediately interpretable. This can be a barrier to model uptake in important…

计算与语言 · 计算机科学 2021-06-18 Diego Garcia-Olano , Yasumasa Onoe , Ioana Baldini , Joydeep Ghosh , Byron C. Wallace , Kush R. Varshney

While coreference resolution is traditionally used as a component in individual document understanding, in this work we take a more global view and explore what can we learn about a domain from the set of all document-level coreference…

计算与语言 · 计算机科学 2024-10-23 Shir Ashury-Tahan , Amir David Nissan Cohen , Nadav Cohen , Yoram Louzoun , Yoav Goldberg

Given the dominance of dense retrievers that do not generalize well beyond their training dataset distributions, domain-specific test sets are essential in evaluating retrieval. There are few test datasets for retrieval systems intended for…

Recent research advances achieve human-level accuracy for de-identifying free-text clinical notes on research datasets, but gaps remain in reproducing this in large real-world settings. This paper summarizes lessons learned from building a…

计算与语言 · 计算机科学 2023-12-15 Veysel Kocaman , Hasham Ul Haq , David Talby

There is an ongoing need for scalable tools to aid researchers in both retrospective and prospective standardization of discrete entity types -- such as disease names, cell types or chemicals -- that are used in metadata associated with…

数据库 · 计算机科学 2024-07-04 Rafael S. Gonçalves , Jason Payne , Amelia Tan , Carmen Benitez , Jamie Haddock , Robert Gentleman

As modern science becomes increasingly data-intensive, the ability to analyze and visualize large-scale, complex datasets is critical to accelerating discovery. However, many domain scientists lack the programming expertise required to…

软件工程 · 计算机科学 2025-12-01 Apu Kumar Chakroborti , Yi Ding , Lipeng Wan

The paper introduces a framework for the evaluation of the encoding of factual scientific knowledge, designed to streamline the manual evaluation process typically conducted by domain experts. Inferring over and extracting information from…

计算与语言 · 计算机科学 2024-10-21 Magdalena Wysocka , Oskar Wysocki , Maxime Delmas , Vincent Mutel , Andre Freitas

Peer reviewing is a central component in the scientific publishing process. We present the first public dataset of scientific peer reviews available for research purposes (PeerRead v1) providing an opportunity to study this important…