中文
相关论文

相关论文: Datasheet for the Pile

200 篇论文

We introduce doc2dial, a new dataset of goal-oriented dialogues that are grounded in the associated documents. Inspired by how the authors compose documents for guiding end users, we first construct dialogue flows based on the content…

计算与语言 · 计算机科学 2020-11-20 Song Feng , Hui Wan , Chulaka Gunasekara , Siva Sankalp Patel , Sachindra Joshi , Luis A. Lastras

ClueWeb22, the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information. Its design was influenced by the need for a high quality, large scale web corpus to support a range of academic…

信息检索 · 计算机科学 2022-12-05 Arnold Overwijk , Chenyan Xiong , Xiao Liu , Cameron VandenBerg , Jamie Callan

A well-designed interactive human-like dialogue system is expected to take actions (e.g. smiling) and respond in a pattern similar to humans. However, due to the limitation of single-modality (only speech) or small volume of currently…

人机交互 · 计算机科学 2022-12-13 Zhiling Luo , Qiankun Shi , Sha Zhao , Wei Zhou , Haiqing Chen , Yuankai Ma , Haitao Leng

Large, diachronic datasets of political discourse are hard to come across, especially for resource-lean languages such as Greek. In this paper, we introduce a curated dataset of the Greek Parliament Proceedings that extends chronologically…

计算与语言 · 计算机科学 2022-10-25 Konstantina Dritsa , Kaiti Thoma , John Pavlopoulos , Panos Louridas

Talk2AI is a large-scale longitudinal dataset of 3,080 conversations (totaling 30,800 turns) between human participants and Large Language Models (LLMs), designed to support research on persuasion, opinion change, and human-AI interaction.…

Spreadsheets are widely used in various fields to do large numerical analysis. While several companies have relied on spreadsheets for decades, data scientists are going in the direction of using scientific programming languages such as…

软件工程 · 计算机科学 2022-11-14 Amir Nassereldine , Patrick Chen , Jinjun Xiong

The problem addressed here is that of simultaneous treatment of several gene expression datasets, possibly collected under different experimental conditions and/or platforms. Using robust statistics, a large scale statistical analysis has…

统计方法学 · 统计学 2014-10-10 Bernard Ycart , Konstantina Charmpi , Sophie Rousseaux , Jean-Jacques Fournié

How can an end-user provide feedback if a deployed structured prediction model generates inconsistent output, ignoring the structural complexity of human language? This is an emerging topic with recent progress in synthetic or constrained…

人工智能 · 计算机科学 2021-12-17 Niket Tandon , Aman Madaan , Peter Clark , Keisuke Sakaguchi , Yiming Yang

Systematic reviews are time-consuming endeavors. Historically speaking, knowledgeable humans have had to screen and extract data from studies before it can be analyzed. However, large language models (LLMs) hold promise to greatly…

人机交互 · 计算机科学 2025-01-22 Noah L. Schroeder , Chris Davis Jaldi , Shan Zhang

The arXiv has collected 1.5 million pre-print articles over 28 years, hosting literature from scientific fields including Physics, Mathematics, and Computer Science. Each pre-print features text, figures, authors, citations, categories, and…

信息检索 · 计算机科学 2019-05-02 Colin B. Clement , Matthew Bierbaum , Kevin P. O'Keeffe , Alexander A. Alemi

The Medical Algorithms Project, a web-based resource located at www.medal.org, is the world's largest collection of medical-related spreadsheets, consisting of over 13,500 Excel spreadsheets each encoding a medical algorithm from 45…

人机交互 · 计算机科学 2009-08-07 M. Sriram Iyengar , John R. Svirbely

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

计算与语言 · 计算机科学 2019-12-02 Masato Hagiwara , Masato Mita

Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on large-scale real-world data such as electronic health records…

Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-parallel datasets at scale. This work introduces HEALTHDIAL, a large-scale, multilingual,…

计算与语言 · 计算机科学 2026-05-29 Songbo Hu , Yinhong Liu , Ej Zhou , Evgeniia Razumovskaia , Xiaobin Wang , Alexander Fraser , Ivan Vulić , Anna Korhonen

Sequence-to-sequence models have recently gained the state of the art performance in summarization. However, not too many large-scale high-quality datasets are available and almost all the available ones are mainly news articles with…

计算与语言 · 计算机科学 2018-10-23 Mahnaz Koupaee , William Yang Wang

Stylistic variation in text needs to be studied with different aspects including the writer's personal traits, interpersonal relations, rhetoric, and more. Despite recent attempts on computational modeling of the variation, the lack of…

计算与语言 · 计算机科学 2019-09-04 Dongyeop Kang , Varun Gangal , Eduard Hovy

This paper introduces IGGA, a dataset of 160 industry guidelines and policy statements for the use of Generative AIs (GAIs) and Large Language Models (LLMs) in industry and workplace settings, collected from official company websites, and…

计算机与社会 · 计算机科学 2025-03-19 Junfeng Jiao , Saleh Afroogh , Kevin Chen , David Atkinson , Amit Dhurandhar

Most particle induced X-ray emission (PIXE) data analysis codes are not focused on handling multilayered samples. We have developed an open-source library called "LibCPIXE", for PIXE data analysis. It is written in standard C and implements…

材料科学 · 物理学 2007-07-18 C. Pascual-Izarra , N. P. Barradas , M. A. Reis

Intelligence analysts perform sensemaking over collections of documents using various visual and analytic techniques to gain insights from large amounts of text. As data scales grow, our work explores how to leverage two AI technologies,…

人机交互 · 计算机科学 2025-10-13 Adam Coscia , Alex Endert

This paper presents a high-quality multilingual dataset for the documentation domain to advance research on localization of structured text. Unlike widely-used datasets for translation of plain text, we collect XML-structured parallel text…

计算与语言 · 计算机科学 2020-06-25 Kazuma Hashimoto , Raffaella Buschiazzo , James Bradbury , Teresa Marshall , Richard Socher , Caiming Xiong