English
Related papers

Related papers: The Harrington Yowlumne Narrative Corpus

200 papers

Neural natural language generation (NNLG) from structured meaning representations has become increasingly popular in recent years. While we have seen progress with generating syntactically correct utterances that preserve semantics, various…

Computation and Language · Computer Science 2019-06-18 Shereen Oraby , Vrindavan Harrison , Abteen Ebrahimi , Marilyn Walker

We introduce the Emergent Language Corpus Collection (ELCC): a collection of corpora generated from open source implementations of emergent communication systems across the literature. These systems include a variety of signalling game…

Computation and Language · Computer Science 2024-12-05 Brendon Boldt , David Mortensen

Maximizing the likelihood of the next token is an established, statistically sound objective for pre-training language models. In this paper we show that we can train better models faster by pre-aggregating the corpus with a collapsed…

Computation and Language · Computer Science 2024-07-04 Ashutosh Sathe , Sunita Sarawagi

Existing curriculum learning approaches to Neural Machine Translation (NMT) require sampling sufficient amounts of "easy" samples from training data at the early training stage. This is not always achievable for low-resource languages where…

Computation and Language · Computer Science 2021-03-23 Chen Liang , Haoming Jiang , Xiaodong Liu , Pengcheng He , Weizhu Chen , Jianfeng Gao , Tuo Zhao

We introduce the Speak & Improve Corpus 2025, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The aim of the…

Computation and Language · Computer Science 2024-12-18 Kate Knill , Diane Nicholls , Mark J. F. Gales , Mengjie Qian , Pawel Stroinski

We present a crowdsourced dataset for Piedmontese, an endangered Romance language of northwestern Italy. The dataset comprises 145 Italian-Piedmontese parallel sentences derived from Flores+, with translations produced by speakers writing…

Computation and Language · Computer Science 2026-02-17 Gianluca Vico , Jindřich Libovický

We present a technique which complements Hidden Markov Models by incorporating some lexicalized states representing syntactically uncommon words. Our approach examines the distribution of transitions, selects the uncommon words, and makes…

Computation and Language · Computer Science 2007-05-23 Jin-Dong Kim , Sang-Zoo Lee , Hae-Chang Rim

We propose a novel transcription workflow which combines spoken term detection and human-in-the-loop, together with a pilot experiment. This work is grounded in an almost zero-resource scenario where only a few terms have so far been…

Computation and Language · Computer Science 2020-11-13 Éric Le Ferrand , Steven Bird , Laurent Besacier

Social media offer an abundant source of valuable raw data, however informal writing can quickly become a bottleneck for many natural language processing (NLP) tasks. Off-the-shelf tools are usually trained on formal text and cannot…

Computation and Language · Computer Science 2019-04-15 Ismini Lourentzou , Kabir Manghnani , ChengXiang Zhai

As natural language processing for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques, such as pre-trained language models, suffer from biased corpus. This case becomes more obvious regarding…

Computation and Language · Computer Science 2025-06-17 Yizhi Li , Ge Zhang , Hanhua Hong , Yiwen Wang , Chenghua Lin

Automatic abstractive text summarization is an important and challenging research topic of natural language processing. Among many widely used languages, the Chinese language has a special property that a Chinese character contains rich…

Computation and Language · Computer Science 2018-09-11 Chieh-Teng Chang , Chi-Chia Huang , Chih-Yuan Yang , Jane Yung-Jen Hsu

Large language models (LLMs) still produce plausible-sounding but ungrounded factual claims, a problem that worsens in multi-turn dialogue as context grows and early errors cascade. We introduce $\textbf{HalluHard}$, a challenging…

Artificial Intelligence · Computer Science 2026-02-03 Dongyang Fan , Sebastien Delsad , Nicolas Flammarion , Maksym Andriushchenko

We introduce HarperValleyBank, a free, public domain spoken dialog corpus. The data simulate simple consumer banking interactions, containing about 23 hours of audio from 1,446 human-human conversations between 59 unique speakers. We…

Machine Learning · Computer Science 2021-03-22 Mike Wu , Jonathan Nafziger , Anthony Scodary , Andrew Maas

Training high-performing Small Language Models (SLMs) remains costly, even with knowledge distillation and pruning from larger teacher models. Existing work often faces three key challenges: (1) information loss from hard pruning, (2)…

Computation and Language · Computer Science 2025-12-19 Jitai Hao , Qiang Huang , Hao Liu , Xinyan Xiao , Zhaochun Ren , Jun Yu

We describe an effort to annotate a corpus of natural language instructions consisting of 622 wet lab protocols to facilitate automatic or semi-automatic conversion of protocols into a machine-readable format and benefit biological…

Computation and Language · Computer Science 2018-05-02 Chaitanya Kulkarni , Wei Xu , Alan Ritter , Raghu Machiraju

Automatic phylogenetic inference plays an increasingly important role in computational historical linguistics. Most pertinent work is currently based on expert cognate judgments. This limits the scope of this approach to a small number of…

Computation and Language · Computer Science 2018-10-18 Gerhard Jäger

Prevalent retrieval-based tool-use pipelines struggle with a dual semantic challenge: their retrievers often employ encoders that fail to capture complex semantics, while the Large Language Model (LLM) itself lacks intrinsic tool knowledge…

Artificial Intelligence · Computer Science 2026-01-30 Bowen Fang , Wen Ye , Yunyue Su , Jinghao Zhang , Qiang Liu , Yesheng Liu , Xin Sun , Shu Wu , Jiabing Yang , Baole Wei , Liang Wang

We present a new, unique and freely available parallel corpus containing European Union (EU) documents of mostly legal nature. It is available in all 20 official EUanguages, with additional documents being available in the languages of the…

Computation and Language · Computer Science 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Anna Widiger , Camelia Ignat , Tomaz Erjavec , Dan Tufis , Daniel Varga

Traditionally, Text Simplification is treated as a monolingual translation task where sentences between source texts and their simplified counterparts are aligned for training. However, especially for longer input documents, summarizing the…

Computation and Language · Computer Science 2022-07-29 Dennis Aumiller , Michael Gertz