English
Related papers

Related papers: FreCDo: A Large Corpus for French Cross-Domain Dia…

200 papers

This paper explores morpho-syntactic ambiguities for French to develop a strategy for part-of-speech disambiguation that a) reflects the complexity of French as an inflected language, b) optimizes the estimation of probabilities, c) allows…

cmp-lg · Computer Science 2007-05-23 Evelyne Tzoukermann , Dragomir R. Radev , William A. Gale

Discourse segmentation is a crucial step in building end-to-end discourse parsers. However, discourse segmenters only exist for a few languages and domains. Typically they only detect intra-sentential segment boundaries, assuming gold…

Computation and Language · Computer Science 2017-04-25 Chloé Braud , Ophélie Lacroix , Anders Søgaard

Quotation extraction is a widely useful task both from a sociological and from a Natural Language Processing perspective. However, very little data is available to study this task in languages other than English. In this paper, we present a…

Computation and Language · Computer Science 2023-09-20 Ange Richard , Laura Alonzo-Canul , François Portet

In this paper we describe our efforts to make a bidirectional Congolese Swahili (SWC) to French (FRA) neural machine translation system with the motivation of improving humanitarian translation workflows. For training, we created a…

Computation and Language · Computer Science 2021-03-22 Alp Öktem , Eric DeLuca , Rodrigue Bashizi , Eric Paquin , Grace Tang

We analyze the process of creating word embedding feature representations designed for a learning task when annotated data is scarce, for example, in depressive language detection from Tweets. We start with a rich word embedding pre-trained…

Computation and Language · Computer Science 2021-06-25 Nawshad Farruque , Randy Goebel , Osmar Zaiane

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

In this paper we present a new method to learn a model robust to typos for a Named Entity Recognition task. Our improvement over existing methods helps the model to take into account the context of the sentence inside a court decision in…

Computation and Language · Computer Science 2019-09-10 Valentin Barriere , Amaury Fouret

Pronoun resolution is a major area of natural language understanding. However, large-scale training sets are still scarce, since manually labelling data is costly. In this work, we introduce WikiCREM (Wikipedia CoREferences Masked) a…

Computation and Language · Computer Science 2019-10-15 Vid Kocijan , Oana-Maria Camburu , Ana-Maria Cretu , Yordan Yordanov , Phil Blunsom , Thomas Lukasiewicz

We present the first French partition of the OLDI Seed Corpus, our submission to the WMT 2025 Open Language Data Initiative (OLDI) shared task. We detail its creation process, which involved using multiple machine translation systems and a…

Computation and Language · Computer Science 2025-08-05 Malik Marmonier , Benoît Sagot , Rachel Bawden

Thanks to the rise of self-supervised learning, automatic speech recognition (ASR) systems now achieve near-human performance on a wide variety of datasets. However, they still lack generalization capability and are not robust to domain…

Machine Learning · Computer Science 2023-03-15 Lucas Maison , Yannick Estève

Production of news content is growing at an astonishing rate. To help manage and monitor the sheer amount of text, there is an increasing need to develop efficient methods that can provide insights into emerging content areas, and stratify…

Computation and Language · Computer Science 2020-10-29 M. Tarik Altuncu , Sophia N. Yaliraki , Mauricio Barahona

The development of robust, multilingual speaker recognition systems is hindered by a lack of large-scale, publicly available and multilingual datasets, particularly for the read-speech style crucial for applications like anti-spoofing. To…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Aref Farhadipour , Jan Marquenie , Srikanth Madikeri , Eleanor Chodroff

Numerous studies have demonstrated the ability of neural language models to learn various linguistic properties without direct supervision. This work takes an initial step towards exploring the less researched topic of how neural models…

Computation and Language · Computer Science 2023-10-25 Lina Conti , Guillaume Wisniewski

We propose a new method that leverages contextual embeddings for the task of diachronic semantic shift detection by generating time specific word representations from BERT embeddings. The results of our experiments in the domain specific…

Computation and Language · Computer Science 2020-03-06 Matej Martinc , Petra Kralj Novak , Senja Pollak

Thanks to the Eslo1 ("Enqu\^ete sociolinguistique d'Orl\'eans", i.e. "Sociolinguistic Inquiery of Orl\'eans") campain, a large oral corpus has been gathered and transcribed in a textual format. The purpose of the work presented here is to…

Machine Learning · Computer Science 2010-03-31 Iris Eshkol , Isabelle Tellier , Taalab Samer , Sylvie Billot

Camouflaged object detection (COD) aims to segment objects that blend into their surroundings. However, most existing studies overlook the semantic differences among textual prompts of different targets as well as fine-grained frequency…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Dezhen Wang , Haixiang Zhao , Xiang Shen , Sheng Miao

The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models face challenges with historical content due to OCR noise and outdated spellings. This study…

Computation and Language · Computer Science 2025-03-14 Andrianos Michail , Corina Julia Raclé , Juri Opitz , Simon Clematide

Large Language Model (LLM) pre-training exhausts an ever growing compute budget, yet recent research has demonstrated that careful document selection enables comparable model quality with only a fraction of the FLOPs. Inspired by efforts…

Computation and Language · Computer Science 2024-06-10 Xiang Kong , Tom Gunter , Ruoming Pang

With the rapid evolution of social media, fake news has become a significant social problem, which cannot be addressed in a timely manner using manual investigation. This has motivated numerous studies on automating fake news detection.…

Computation and Language · Computer Science 2021-02-25 Amila Silva , Ling Luo , Shanika Karunasekera , Christopher Leckie

We present M2D2, a fine-grained, massively multi-domain corpus for studying domain adaptation in language models (LMs). M2D2 consists of 8.5B tokens and spans 145 domains extracted from Wikipedia and Semantic Scholar. Using ontologies…

Computation and Language · Computer Science 2022-10-17 Machel Reid , Victor Zhong , Suchin Gururangan , Luke Zettlemoyer
‹ Prev 1 3 4 5 6 7 10 Next ›