中文
相关论文

相关论文: DMDD: A Large-Scale Dataset for Dataset Mentions D…

200 篇论文

Scientific full text papers are usually stored in separate places than their underlying research datasets. Authors typically make references to datasets by mentioning them for example by using their titles and the year of publication.…

数字图书馆 · 计算机科学 2016-03-30 Behnam Ghavimi , Philipp Mayr , Sahar Vahdati , Christoph Lange

Knowledge Graph has been proven effective in modeling structured information and conceptual knowledge, especially in the medical domain. However, the lack of high-quality annotated corpora remains a crucial problem for advancing the…

计算与语言 · 计算机科学 2021-09-22 Dejie Chang , Mosha Chen , Chaozhen Liu , Liping Liu , Dongdong Li , Wei Li , Fei Kong , Bangchang Liu , Xiaobin Luo , Ji Qi , Qiao Jin , Bin Xu

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural language descriptions…

Scientific information extraction (SciIE) is critical for converting unstructured knowledge from scholarly articles into structured data (entities and relations). Several datasets have been proposed for training and validating SciIE models.…

计算与语言 · 计算机科学 2024-10-29 Qi Zhang , Zhijia Chen , Huitong Pan , Cornelia Caragea , Longin Jan Latecki , Eduard Dragut

Dataset Condensation is a newly emerging technique aiming at learning a tiny dataset that captures the rich information encoded in the original dataset. As the size of datasets contemporary machine learning models rely on becomes…

机器学习 · 计算机科学 2022-10-18 Justin Cui , Ruochen Wang , Si Si , Cho-Jui Hsieh

Medical imaging papers often focus on methodology, but the quality of the algorithms and the validity of the conclusions are highly dependent on the datasets used. As creating datasets requires a lot of effort, researchers often use…

Automatically locating named entities in natural language text - named entity recognition - is an important task in the biomedical domain. Many named entity mentions are ambiguous between several bioconcept types, however, causing text…

计算与语言 · 计算机科学 2019-09-24 Chih-Hsuan Wei , Kyubum Lee , Robert Leaman , Zhiyong Lu

Scientific document understanding is challenging as the data is highly domain specific and diverse. However, datasets for tasks with scientific text require expensive manual annotation and tend to be small and limited to only one or a few…

计算与语言 · 计算机科学 2021-05-26 Dustin Wright , Isabelle Augenstein

We introduce ChemDisGene, a new dataset for training and evaluating multi-class multi-label document-level biomedical relation extraction models. Our dataset contains 80k biomedical research abstracts labeled with mentions of chemicals,…

计算与语言 · 计算机科学 2022-04-15 Dongxu Zhang , Sunil Mohan , Michaela Torkar , Andrew McCallum

Data intensive research requires the support of appropriate datasets. However, it is often time-consuming to discover usable datasets matching a specific research topic. We formulate the dataset discovery problem on an attributed…

信息检索 · 计算机科学 2021-06-08 Basmah Altaf , Shichao Pei , Xiangliang Zhang

Scientific talks are a growing medium for disseminating research, and automatically identifying relevant literature that grounds or enriches a talk would be highly valuable for researchers and students alike. We introduce Reference…

计算与语言 · 计算机科学 2025-10-29 Frederik Broy , Maike Züfle , Jan Niehues

Most existing large-scale academic search engines are built to retrieve text-based information. However, there are no large-scale retrieval services for scientific figures and tables. One challenge for such services is understanding…

人工智能 · 计算机科学 2023-01-31 Zeba Karishma , Shaurya Rohatgi , Kavya Shrinivas Puranik , Jian Wu , C. Lee Giles

Mentorship in science is crucial for topic choice, career decisions, and the success of mentees and mentors. Typically, researchers who study mentorship use article co-authorship and doctoral dissertation datasets. However, available…

数字图书馆 · 计算机科学 2021-06-14 Qing Ke , Lizhen Liang , Ying Ding , Stephen V. David , Daniel E. Acuna

Recent regulatory initiatives like the European AI Act and relevant voices in the Machine Learning (ML) community stress the need to describe datasets along several key dimensions for trustworthy AI, such as the provenance processes and…

数字图书馆 · 计算机科学 2024-05-27 Joan Giner-Miguelez , Abel Gómez , Jordi Cabot

Autonomous driving is among the largest domains in which deep learning has been fundamental for progress within the last years. The rise of datasets went hand in hand with this development. All the more striking is the fact that researchers…

机器学习 · 计算机科学 2022-05-04 Daniel Bogdoll , Felix Schreyer , J. Marius Zöllner

Literature analysis facilitates researchers to acquire a good understanding of the development of science and technology. The traditional literature analysis focuses largely on the literature metadata such as topics, authors, abstracts,…

人工智能 · 计算机科学 2021-01-29 Linlin Hou , Ji Zhang , Ou Wu , Ting Yu , Zhen Wang , Zhao Li , Jianliang Gao , Yingchun Ye , Rujing Yao

Open datasets play a crucial role in three research domains that intersect data science and education: learning analytics, educational data mining, and artificial intelligence in education. Researchers in these domains apply computational…

计算机与社会 · 计算机科学 2026-04-14 Valdemar Švábenský , Brendan Flanagan , Erwin Daniel López Zapata , Atsushi Shimada

The proliferation of datasets across open data portals and enterprise data lakes presents an opportunity for deriving data-driven insights. Widely-used dataset search systems rely on keyword search over dataset metadata, including…

数据库 · 计算机科学 2025-12-19 Haoxiang Zhang , Yurong Liu , Aécio Santos , Wei-Lun Hung , Juliana Freire

Training Data Detection (TDD) is a task aimed at determining whether a specific data instance is used to train a machine learning model. In the computer security literature, TDD is also referred to as Membership Inference Attack (MIA).…

密码学与安全 · 计算机科学 2025-08-12 Zhihao Zhu , Yi Yang , Defu Lian

Data plays a vital role in machine learning studies. In the research of recommendation, both user behaviors and side information are helpful to model users. So, large-scale real scenario datasets with abundant user behaviors will contribute…

信息检索 · 计算机科学 2021-06-14 Bin Hao , Min Zhang , Weizhi Ma , Shaoyun Shi , Xinxing Yu , Houzhi Shan , Yiqun Liu , Shaoping Ma