English
Related papers

Related papers: MOROCO: The Moldavian and Romanian Dialectal Corpu…

200 papers

The five idioms (i.e., varieties) of the Romansh language are largely standardized and are taught in the schools of the respective communities in Switzerland. In this paper, we present the first parallel corpus of Romansh idioms. The corpus…

Computation and Language · Computer Science 2026-02-16 Zachary Hopton , Jannis Vamvas , Andrin Büchler , Anna Rutkiewicz , Rico Cathomas , Rico Sennrich

Multilingual language models have been a crucial breakthrough as they considerably reduce the need of data for under-resourced languages. Nevertheless, the superiority of language-specific models has already been proven for languages having…

Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities. This information can further enhance many downstream tasks in the field of information retrieval…

Computation and Language · Computer Science 2023-02-14 Yash Verma , Anubhav Jangra , Raghvendra Kumar , Sriparna Saha

Dialogue is at the core of human behaviour and being able to identify the topic at hand is crucial to take part in conversation. Yet, there are few accounts of the topical organisation in casual dialogue and of how people recognise the…

Computation and Language · Computer Science 2025-01-15 Amandine Decker , Vincent Tourneur , Maxime Amblard , Ellen Breitholtz

Natural language inference (NLI), the task of recognizing the entailment relationship in sentence pairs, is an actively studied topic serving as a proxy for natural language understanding. Despite the relevance of the task in building…

Computation and Language · Computer Science 2024-10-21 Eduard Poesina , Cornelia Caragea , Radu Tudor Ionescu

In this work, we present a novel perspective on cognitive impairment classification from speech by integrating speech foundation models that explicitly recognize speech dialects. Our motivation is based on the observation that individuals…

Sound · Computer Science 2026-01-14 Tiantian Feng , Anfeng Xu , Jinkook Lee , Shrikanth Narayanan

There has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performance. The MALACH corpus (LDC catalog LDC2012S05), a 375-Hour…

Computation and Language · Computer Science 2019-08-12 Michael Picheny , Zóltan Tüske , Brian Kingsbury , Kartik Audhkhasi , Xiaodong Cui , George Saon

Norway has a large amount of dialectal variation, as well as a general tolerance to its use in the public sphere. There are, however, few available resources to study this variation and its change over time and in more informal areas, \eg…

Computation and Language · Computer Science 2021-04-13 Jeremy Barnes , Petter Mæhlum , Samia Touileb

We introduce a multilabel probing task to assess the morphosyntactic representations of word embeddings from multilingual language models. We demonstrate this task with multilingual BERT (Devlin et al., 2018), training probes for seven…

Computation and Language · Computer Science 2021-04-20 Naomi Tachikawa Shapiro , Amandalynne Paullada , Shane Steinert-Threlkeld

We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error detection and…

We present a probabilistic model that uses both prosodic and lexical cues for the automatic segmentation of speech into topically coherent units. We propose two methods for combining lexical and prosodic information using hidden Markov…

Computation and Language · Computer Science 2022-02-28 G. Tur , D. Hakkani-Tur , A. Stolcke , E. Shriberg

Our main contribution in this work is novel results of multilingual models that go beyond typical applications of rumor or misinformation detection in English social news content to identify fine-grained classes of digital deception across…

Social and Information Networks · Computer Science 2019-09-13 Maria Glenski , Ellyn Ayton , Josh Mendoza , Svitlana Volkova

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

Computation and Language · Computer Science 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

In this paper, we present a Bayesian multilingual document model for learning language-independent document embeddings. The model is an extension of BaySMM [Kesiraju et al 2020] to the multilingual scenario. It learns to represent the…

Computation and Language · Computer Science 2024-03-26 Santosh Kesiraju , Sangeet Sagar , Ondřej Glembek , Lukáš Burget , Ján Černocký , Suryakanth V Gangashetty

The importance of clear and correct text in legal documents cannot be understated, and, consequently, a grammatical error correction tool meant to assist a professional in the law must have the ability to understand the possible errors in…

Computation and Language · Computer Science 2026-04-23 Mircea Timpuriu , Mihaela-Claudia Cercel , Dumitru-Clementin Cercel

Discourse parsing is an important task useful for NLU applications such as summarization, machine comprehension, and emotion recognition. The current discourse parsing datasets based on conversations consists of written English dialogues…

Computation and Language · Computer Science 2025-06-11 Divyaksh Shukla , Ritesh Baviskar , Dwijesh Gohil , Aniket Tiwari , Atul Shree , Ashutosh Modi

To support machine learning of cross-language prosodic mappings and other ways to improve speech-to-speech translation, we present a protocol for collecting closely matched pairs of utterances across languages, a description of the…

Computation and Language · Computer Science 2023-07-17 Nigel G. Ward , Jonathan E. Avila , Emilia Rivas , Divette Marco

Researchers working in areas such as lexicography, translation studies, and computational linguistics, use a combination of automated and semi-automated tools to analyze the content of text corpora. Keywords, named entities, and events are…

Human-Computer Interaction · Computer Science 2022-03-24 Shane Sheehan , Saturnino Luz , Masood Masoodian

Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy corpus based on the original English Word2vec word analogy corpus and…

Computation and Language · Computer Science 2017-11-09 Lukas Svoboda , Slobodan Beliga