中文
相关论文

相关论文: DMDD: A Large-Scale Dataset for Dataset Mentions D…

200 篇论文

We describe a large, high-quality benchmark for the evaluation of Mention Detection tools. The benchmark contains annotations of both named entities as well as other types of entities, annotated on different types of text, ranging from…

计算与语言 · 计算机科学 2018-01-26 Yosi Mass , Lili Kotlerman , Shachar Mirkin , Elad Venezian , Gera Witzling , Noam Slonim

Identity documents recognition is an important sub-field of document analysis, which deals with tasks of robust document detection, type identification, text fields recognition, as well as identity fraud prevention and document authenticity…

In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of…

信息检索 · 计算机科学 2017-03-06 Nemanja Spasojevic , Preeti Bhargava , Guoning Hu

Cloud computing has become a powerful and indispensable technology for complex, high performance and scalable computation. The exponential expansion in the deployment of cloud technology has produced a massive amount of data from a variety…

密码学与安全 · 计算机科学 2020-05-27 Fadi Salo , MohammadNoor Injadat , Ali Bou Nassif , Aleksander Essex

The detection and extraction of abbreviations from unstructured texts can help to improve the performance of Natural Language Processing tasks, such as machine translation and information retrieval. However, in terms of publicly available…

计算与语言 · 计算机科学 2022-05-02 Leonardo Zilio , Hadeel Saadany , Prashant Sharma , Diptesh Kanojia , Constantin Orăsan

The Web today has millions of datasets, and the number of datasets continues to grow at a rapid pace. These datasets are not standalone entities; rather, they are intricately connected through complex relationships. Semantic relationships…

信息检索 · 计算机科学 2024-08-28 Kate Lin , Tarfah Alrashed , Natasha Noy

Context: Data Mining (DM) method has been evolving year by year and as of today there is also the enhancement of DM technique that can be run several times faster than the traditional one, called Distributed Data Mining (DDM). It is not a…

分布式、并行与集群计算 · 计算机科学 2020-09-23 Fauzi Adi Rafrastara , Qi Deyu

A growing number of papers are published in the area of superconducting materials science. However, novel text and data mining (TDM) processes are still needed to efficiently access and exploit this accumulated knowledge, paving the way…

Citation recommendation describes the task of recommending citations for a given text. Due to the overload of published scientific works in recent years on the one hand, and the need to cite the most appropriate publications when writing…

信息检索 · 计算机科学 2020-09-09 Michael Färber , Adam Jatowt

With the recent progress in machine learning, boosted by techniques such as deep learning, many tasks can be successfully solved once a large enough dataset is available for training. Nonetheless, human-annotated datasets are often…

计算与语言 · 计算机科学 2019-08-19 Daniel Specht Menezes , Pedro Savarese , Ruy Luiz Milidiú

Proper citation is of great importance in academic writing for it enables knowledge accumulation and maintains academic integrity. However, citing properly is not an easy task. For published scientific entities, the ever-growing academic…

数字图书馆 · 计算机科学 2022-10-20 Jialiang Lin , Yao Yu , Jiaxin Song , Xiaodong Shi

This paper explores methods for building a comprehensive citation graph using big data techniques to evaluate scientific impact more accurately. Traditional citation metrics have limitations, and this work investigates merging large…

数字图书馆 · 计算机科学 2025-05-08 Inci Yueksel-Erguen , Ida Litzel , Hanqiu Peng

Recent progress in face detection (including keypoint detection), and recognition is mainly being driven by (i) deeper convolutional neural network architectures, and (ii) larger datasets. However, most of the large datasets are maintained…

计算机视觉与模式识别 · 计算机科学 2017-05-23 Ankan Bansal , Anirudh Nanduri , Carlos Castillo , Rajeev Ranjan , Rama Chellappa

This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the foundational infrastructure analogous to a root system that…

计算与语言 · 计算机科学 2024-02-29 Yang Liu , Jiahuan Cao , Chongyu Liu , Kai Ding , Lianwen Jin

Plant disease recognition has witnessed a significant improvement with deep learning in recent years. Although plant disease datasets are essential and many relevant datasets are public available, two fundamental questions exist. First, how…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Mingle Xu , Ji Eun Park , Jaehwan Lee , Jucheng Yang , Sook Yoon

Hypothetical induction is recognized as the main reasoning type when scientists make observations about the world and try to propose hypotheses to explain those observations. Past research on hypothetical induction is under a constrained…

计算与语言 · 计算机科学 2024-06-13 Zonglin Yang , Xinya Du , Junxian Li , Jie Zheng , Soujanya Poria , Erik Cambria

Dataset Distillation (DD) is a promising technique to synthesize a smaller dataset that preserves essential information from the original dataset. This synthetic dataset can serve as a substitute for the original large-scale one, and help…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Yao Lu , Jianyang Gu , Xuguang Chen , Saeed Vahidian , Qi Xuan

Dataset distillation (DD) is a newly emerging research area aiming at alleviating the heavy computational load in training models on large datasets. It tries to distill a large dataset into a small and condensed one so that models trained…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Yuxuan Duan , Jianfu Zhang , Liqing Zhang

This document gives a set of recommendations to build and manipulate the datasets used to develop and/or validate machine learning models such as deep neural networks. This document is one of the 3 documents defined in [1] to ensure the…

Program code as a data source is gaining popularity in the data science community. Possible applications for models trained on such assets range from classification for data dimensionality reduction to automatic code generation. However,…

软件工程 · 计算机科学 2022-10-31 Anastasia Drozdova , Polina Guseva , Ekaterina Trofimova , Anna Scherbakova , Andrey Ustyuzhanin