English
Related papers

Related papers: CC-GPX: Extracting High-Quality Annotated Geospati…

200 papers

Pre-training Large Language Models (LLMs) on high-quality, meticulously curated datasets is widely recognized as critical for enhancing their performance and generalization capabilities. This study explores the untapped potential of Common…

Hosting over 10 million of software projects, GitHub is one of the most important data sources to study behavior of developers and software projects. However, with the increase of the size of open source datasets, the potential threats to…

Software Engineering · Computer Science 2018-05-09 Can Cheng , Bing Li , Zengyang Li , Peng Liang

Geoscientists, as well as researchers in many fields, need to read a huge amount of literature to locate, extract, and aggregate relevant results and data to enable future research or to build a scientific database, but there is no existing…

Human-Computer Interaction · Computer Science 2022-02-25 Shao Zhang , Yuting Jia , Hui Xu , Ying Wen , Dakuo Wang , Xinbing Wang

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

Computation and Language · Computer Science 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

The rapid advancements in Large Language Models (LLMs) have revolutionized natural language processing, with GPTs, customized versions of ChatGPT available on the GPT Store, emerging as a prominent technology for specific domains and tasks.…

Software Engineering · Computer Science 2024-05-27 Xinyi Hou , Yanjie Zhao , Shenao Wang , Haoyu Wang

Analyzing and predicting the traffic scene around the ego vehicle has been one of the key challenges in autonomous driving. Datasets including the trajectories of all road users present in a scene, as well as the underlying road topology…

Computer Vision and Pattern Recognition · Computer Science 2020-07-17 Antonia Breuer , Jan-Aike Termöhlen , Silviu Homoceanu , Tim Fingscheidt

Although Natural Language Processing (NLP) research on argument mining has advanced considerably in recent years, most studies draw on corpora of asynchronous and written texts, often produced by individuals. Few published corpora of…

Computation and Language · Computer Science 2020-05-26 Christopher Olshefski , Luca Lugini , Ravneet Singh , Diane Litman , Amanda Godley

Recent research on pattern discovery has progressed from mining frequent patterns and sequences to mining structured patterns, such as trees and graphs. Graphs as general data structure can model complex relations among data with wide…

Databases · Computer Science 2013-12-17 Ghazi Al-Naymat

This paper contributes to cross-lingual image annotation and retrieval in terms of data and baseline methods. We propose COCO-CN, a novel dataset enriching MS-COCO with manually written Chinese sentences and tags. For more effective…

Computation and Language · Computer Science 2019-01-16 Xirong Li , Chaoxi Xu , Xiaoxu Wang , Weiyu Lan , Zhengxiong Jia , Gang Yang , Jieping Xu

Recently we create so much data (2.5 quintillion bytes every day) that 90% of the data in the world today has been created in the last two years alone [1]. This data comes from sensors used to gather traffic or climate information, posts to…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-12-20 Afsin Akdogan , Hien To

Large scale comparative research into municipal governance is often prohibitively difficult due to a lack of high-quality data. But, recent advances in speech-to-text algorithms and natural language processing has made it possible to more…

Digital Libraries · Computer Science 2022-09-01 Eva Maxfield Brown , Nicholas Weber

Community Search (CS) is one of the fundamental graph analysis tasks, which is a building block of various real applications. Given any query nodes, CS aims to find cohesive subgraphs that query nodes belong to. Recently, a large number of…

Databases · Computer Science 2023-10-10 Shuheng Fang , Kangfei Zhao , Guanghua Li , Jeffery Xu Yu

This paper evaluates large georeferenced corpora, taken from both web-crawled and social media sources, against ground-truth population and language-census datasets. The goal is to determine (i) which dataset best represents population…

Computation and Language · Computer Science 2020-04-03 Jonathan Dunn , Ben Adams

Clustering trajectory data attracted considerable attention in the last few years. Most of prior work assumed that moving objects can move freely in an euclidean space and did not consider the eventual presence of an underlying road network…

Machine Learning · Computer Science 2013-10-22 Mohamed Khalil El Mahrsi , Fabrice Rossi

Building high-quality datasets for specialized tasks is a time-consuming and resource-intensive process that often requires specialized domain knowledge. We propose Corpus Retrieval and Augmentation for Fine-Tuning (CRAFT), a method for…

Computation and Language · Computer Science 2025-12-08 Ingo Ziegler , Abdullatif Köksal , Desmond Elliott , Hinrich Schütze

Public omics databases like the Gene Expression Omnibus and the Sequence Read Archive offer substantial opportunities for data reuse to address novel biomedical questions. However, it is still difficult to find samples and studies of…

We propose a new large-scale (nearly a million questions) ultra-long-context (more than 50,000 words average document length) reading comprehension dataset. Using GPT 3.5, we summarized each scene in 1,500 hand-curated fiction books from…

Computation and Language · Computer Science 2023-12-11 Arseny Moskvichev , Ky-Vinh Mai

The advancements of neural dialogue generation models show promising results on modeling short-text conversations. However, training such models usually needs a large-scale high-quality dialogue corpus, which is hard to access. In this…

Computation and Language · Computer Science 2022-04-27 Yida Wang , Pei Ke , Yinhe Zheng , Kaili Huang , Yong Jiang , Xiaoyan Zhu , Minlie Huang

Graph condensation (GC) has recently garnered considerable attention due to its ability to reduce large-scale graph datasets while preserving their essential properties. The core concept of GC is to create a smaller, more manageable graph…

Machine Learning · Computer Science 2024-11-25 Qingyun Sun , Ziying Chen , Beining Yang , Cheng Ji , Xingcheng Fu , Sheng Zhou , Hao Peng , Jianxin Li , Philip S. Yu

This paper presents SwissCrawl, the largest Swiss German text corpus to date. Composed of more than half a million sentences, it was generated using a customized web scraping tool that could be applied to other low-resource languages as…

Computation and Language · Computer Science 2020-06-17 Lucy Linder , Michael Jungo , Jean Hennebert , Claudiu Musat , Andreas Fischer
‹ Prev 1 4 5 6 7 8 10 Next ›