English
Related papers

Related papers: DMDD: A Large-Scale Dataset for Dataset Mentions D…

200 papers

We present DMCD (DataMap Causal Discovery), a two-phase causal discovery framework that integrates LLM-based semantic drafting from variable metadata with statistical validation on observational data. In Phase I, a large language model…

Artificial Intelligence · Computer Science 2026-02-25 Samarth KaPatel , Sofia Nikiforova , Giacinto Paolo Saggese , Paul Smith

Acronym extraction is the task of identifying acronyms and their expanded forms in texts that is necessary for various NLP applications. Despite major progress for this task in recent years, one limitation of existing AE research is that…

Computation and Language · Computer Science 2022-02-22 Amir Pouran Ben Veyseh , Nicole Meister , Seunghyun Yoon , Rajiv Jain , Franck Dernoncourt , Thien Huu Nguyen

The progress of event extraction research has been hindered by the absence of wide-coverage, large-scale datasets. To make event extraction systems more accessible, we build a general-purpose event detection dataset GLEN, which covers 205K…

Computation and Language · Computer Science 2023-11-01 Qiusi Zhan , Sha Li , Kathryn Conger , Martha Palmer , Heng Ji , Jiawei Han

We present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as well as annotation guidelines containing straightforward--but…

Computation and Language · Computer Science 2022-04-08 Rustem Yeshpanov , Yerbolat Khassanov , Huseyin Atakan Varol

A citation is a well-established mechanism for connecting scientific artifacts. Citation networks are used by citation analysis for a variety of reasons, prominently to give credit to scientists' work. However, because of current citation…

Information Retrieval · Computer Science 2020-01-17 Tong Zeng , Longfeng Wu , Sarah Bratt , Daniel E. Acuna

The importance of databases of reliable and accurate data in chemistry has substantially increased in the past two decades. Their main usage is to parametrize electronic structure theory methods, and to assess their capabilities and…

Chemical Physics · Physics 2023-11-10 Pierpaolo Morgante , Roberto Peverati

Large text data sets, such as publications, websites, and other text-based media, inherit two distinct types of features: (1) the text itself, its information conveyed through semantics, and (2) its relationship to other texts through…

Computation and Language · Computer Science 2026-02-05 Tim Kunt , Annika Buchholz , Imene Khebouri , Thorsten Koch , Ida Litzel , Thi Huong Vu

Cancer diseases constitute one of the most significant societal challenges. In this paper, we introduce a novel histopathological dataset for prostate cancer detection. The proposed dataset, consisting of over 2.6 million tissue patches…

The Environmental Microorganism Image Dataset Seventh Version (EMDS-7) is a microscopic image data set, including the original Environmental Microorganism images (EMs) and the corresponding object labeling files in ".XML" format file. The…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Hechen Yang , Chen Li , Xin Zhao , Bencheng Cai , Jiawei Zhang , Pingli Ma , Peng Zhao , Ao Chen , Tao Jiang , Hongzan Sun , Yueyang Teng , Shouliang Qi , Tao Jiang , Marcin Grzegorzek

Dataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density,…

Background: Machine learning methods for clinical named entity recognition and entity normalization systems can utilize both labeled corpora and Knowledge Graphs (KGs) for learning. However, infrequently occurring concepts may have few…

Computation and Language · Computer Science 2024-10-11 Kuleen Sasse , Shinjitha Vadlakonda , Richard E. Kennedy , John D. Osborne

Researchers apply machine-learning techniques for code smell detection to counter the subjectivity of many code smells. Such approaches need a large, manually annotated dataset for training and benchmarking. Existing literature offers a few…

Software Engineering · Computer Science 2023-03-16 Himesh Nandani , Mootez Saad , Tushar Sharma

Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GPT-4 excel in medical question answering but may face…

Computation and Language · Computer Science 2025-07-02 Bowen Wang , Jiuyang Chang , Yiming Qian , Guoxin Chen , Junhao Chen , Zhouqiang Jiang , Jiahao Zhang , Yuta Nakashima , Hajime Nagahara

Climate models have been key for assessing the impact of climate change and simulating future climate scenarios. The machine learning (ML) community has taken an increased interest in supporting climate scientists' efforts on various tasks…

The extensive amounts of data required for training deep neural networks pose significant challenges on storage and transmission fronts. Dataset distillation has emerged as a promising technique to condense the information of massive…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Ali Abbasi , Ashkan Shahbazi , Hamed Pirsiavash , Soheil Kolouri

Objective: Retrieval-based Clinical Decision Support (ReCDS) can aid clinical workflow by providing relevant literature and similar patients for a given patient. However, the development of ReCDS systems has been severely obstructed by the…

Computation and Language · Computer Science 2023-12-21 Zhengyun Zhao , Qiao Jin , Fangyuan Chen , Tuorui Peng , Sheng Yu

Current citation practices observed in articles are very noisy, confusing, and not standardised, making identifying the cited works problematic for hu-mans and any reference extraction software. In this work, we want to investigate such…

Digital Libraries · Computer Science 2022-07-22 Erika Alves dos Santos , Silvio Peroni , Marcos Luiz Mucheroni

Structured information extraction from scientific literature is crucial for capturing core concepts and emerging trends in specialized fields. While existing datasets aid model development, most focus on specific publication sections due to…

Computation and Language · Computer Science 2026-04-06 Decheng Duan , Yingyi Zhang , Jitong Peng , Chengzhi Zhang

At the core of many important machine learning problems faced by online streaming services is a need to model how users interact with the content they are served. Unfortunately, there are no public datasets currently available that enable…

Information Retrieval · Computer Science 2020-10-16 Brian Brost , Rishabh Mehrotra , Tristan Jehan

Anomaly detection (AD) aims to identify defects using normal-only training data. Existing anomaly detection benchmarks (e.g., MVTec-AD with 15 categories) cover only a narrow range of categories, limiting the evaluation of cross-context…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Hai Ling , Jia Guo , Zhulin Tao , Yunkang Cao , Donglin Di , Hongyan Xu , Xiu Su , Yang Song , Lei Fan
‹ Prev 1 8 9 10 Next ›