English
Related papers

Related papers: Datasheet for the Pile

200 papers

Code large language models mark a pivotal breakthrough in artificial intelligence. They are specifically crafted to understand and generate programming languages, significantly boosting the efficiency of coding development workflows. In…

Software Engineering · Computer Science 2024-03-26 Rui Xie , Zhengran Zeng , Zhuohao Yu , Chang Gao , Shikun Zhang , Wei Ye

Peer reviewing is a central component in the scientific publishing process. We present the first public dataset of scientific peer reviews available for research purposes (PeerRead v1) providing an opportunity to study this important…

Computation and Language · Computer Science 2018-04-26 Dongyeop Kang , Waleed Ammar , Bhavana Dalvi , Madeleine van Zuylen , Sebastian Kohlmeier , Eduard Hovy , Roy Schwartz

Dehumanization is a mental process that enables the exclusion and ill treatment of a group of people. In this paper, we present two data sets of dehumanizing text, a large, automatically collected corpus and a smaller, manually annotated…

Computation and Language · Computer Science 2024-02-15 Paul Engelmann , Peter Brunsgaard Trolle , Christian Hardmeier

Relevant information in documents is often summarized in tables, helping the reader to identify useful facts. Most benchmark datasets support either document layout analysis or table understanding, but lack in providing data to apply both…

Computation and Language · Computer Science 2023-02-14 Andrea Gemelli , Emanuele Vivoli , Simone Marinai

Data plays a vital role in machine learning studies. In the research of recommendation, both user behaviors and side information are helpful to model users. So, large-scale real scenario datasets with abundant user behaviors will contribute…

Information Retrieval · Computer Science 2021-06-14 Bin Hao , Min Zhang , Weizhi Ma , Shaoyun Shi , Xinxing Yu , Houzhi Shan , Yiqun Liu , Shaoping Ma

ML/AI is the field of computer science and computer engineering that arguably received the most attention and funding over the last decade. Data is the key element of ML/AI, so it is becoming increasingly important to ensure that users are…

Digital Libraries · Computer Science 2025-03-19 Marco Rondina , Antonio Vetrò , Juan Carlos De Martin

The current scientific and technological landscape is characterised by the increasing availability of data resources and processing tools and services. In this setting, metadata have emerged as a key factor facilitating management, sharing…

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of…

This paper explores the characteristics of DataCite to determine its possibilities and potential as a new bibliometric data source to analyze the scholarly production of open data. Open science and the increasing data sharing requirements…

Digital Libraries · Computer Science 2017-10-13 Nicolas Robinson-Garcia , Philippe Mongeon , Wei Jeng , Rodrigo Costas

This contribution argues that Reddit, as a massive, categorized, open-access dataset, is a useful data source, for "almost any topic". Hence, it can be used in data science, e.g. for knowledge exploration. This statement is backed-up with…

Information Retrieval · Computer Science 2024-10-15 Jan Sawicki , Maria Ganzha , Marcin Paprzycki , Amelia Bădică

Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is compounded by the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Chao Ding , Mouxiao Bian , Pengcheng Chen , Hongliang Zhang , Tianbin Li , Lihao Liu , Jiayuan Chen , Zhuoran Li , Yabei Zhong , Yongqi Liu , Haiqing Huang , Dongming Shan , Junjun He , Jie Xu

The paper discusses the creation of a multimodal dataset of Russian-language scientific papers and testing of existing language models for the task of automatic text summarization. A feature of the dataset is its multimodal data, which…

Computation and Language · Computer Science 2024-05-14 Alena Tsanda , Elena Bruches

How should text dataset sizes be compared across languages? Even for content-matched (parallel) corpora, UTF-8 encoded text can require a dramatically different number of bytes for different languages. In our work, we define the byte…

Computation and Language · Computer Science 2024-03-04 Catherine Arnett , Tyler A. Chang , Benjamin K. Bergen

In this report, we present ChuXin, an entirely open-source language model with a size of 1.6 billion parameters. Unlike the majority of works that only open-sourced the model weights and architecture, we have made everything needed to train…

Computation and Language · Computer Science 2024-05-09 Xiaomin Zhuang , Yufan Jiang , Qiaozhi He , Zhihua Wu

Emotion-Cause analysis has attracted the attention of researchers in recent years. However, most existing datasets are limited in size and number of emotion categories. They often focus on extracting parts of the document that contain the…

Computation and Language · Computer Science 2024-08-09 Mia Huong Nguyen , Yasith Samaradivakara , Prasanth Sasikumar , Chitralekha Gupta , Suranga Nanayakkara

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

Computation and Language · Computer Science 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

Large language models (LLMs) serve as powerful tools for design, providing capabilities for both task automation and design assistance. Recent advancements have shown tremendous potential for facilitating LLM integration into the chip…

Building document-grounded dialogue systems have received growing interest as documents convey a wealth of human knowledge and commonly exist in enterprises. Wherein, how to comprehend and retrieve information from documents is a…

Computation and Language · Computer Science 2022-07-15 Zhenyu Zhang , Bowen Yu , Haiyang Yu , Tingwen Liu , Cheng Fu , Jingyang Li , Chengguang Tang , Jian Sun , Yongbin Li

Current conversational recommendation systems focus predominantly on text. However, real-world recommendation settings are generally multimodal, causing a significant gap between existing research and practical applications. To address this…

Multimedia · Computer Science 2025-04-16 Zihan Wang , Xiaocui Yang , Yongkang Liu , Shi Feng , Daling Wang , Yifei Zhang

Organisations disclose their privacy practices by posting privacy policies on their website. Even though users often care about their digital privacy, they often don't read privacy policies since they require a significant investment in…

Information Retrieval · Computer Science 2024-04-02 Mukund Srinath , Shomir Wilson , C. Lee Giles
‹ Prev 1 8 9 10 Next ›