中文
相关论文

相关论文: FreCDo: A Large Corpus for French Cross-Domain Dia…

200 篇论文

In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on general domain data, specialized ones have emerged to more…

This paper experiments with frequency-based corpus similarity measures across 39 languages using a register prediction task. The goal is to quantify (i) the distance between different corpora from the same language and (ii) the homogeneity…

计算与语言 · 计算机科学 2022-06-10 Haipeng Li , Jonathan Dunn

Despite impressive advancements in multilingual corpora collection and model training, developing large-scale deployments of multilingual models still presents a significant challenge. This is particularly true for language tasks that are…

Encoder models trained for the embedding of sentences or short documents have proven useful for tasks such as semantic search and topic modeling. In this paper, we present a version of the SwissBERT encoder model that we specifically…

计算与语言 · 计算机科学 2024-05-14 Juri Grosjean , Jannis Vamvas

In this paper, we introduce the Fongbe to French Speech Translation Corpus (FFSTC) for the first time. This corpus encompasses approximately 31 hours of collected Fongbe language content, featuring both French transcriptions and…

计算与语言 · 计算机科学 2024-03-11 D. Fortune Kponou , Frejus A. A. Laleye , Eugene C. Ezin

We introduce word vectors for the construction domain. Our vectors were obtained by running word2vec on an 11M-word corpus that we created from scratch by leveraging freely-accessible online sources of construction-related text. We first…

计算与语言 · 计算机科学 2016-10-31 Antoine J. -P. Tixier , Michalis Vazirgiannis , Matthew R. Hallowell

In this article, we describe some discursive segmentation methods as well as a preliminary evaluation of the segmentation quality. Although our experiment were carried for documents in French, we have developed three discursive segmentation…

In this work we tackle the problem of sentence boundary detection applied to French as a binary classification task ("sentence boundary" or "not sentence boundary"). We combine convolutional neural networks with subword-level information…

计算与语言 · 计算机科学 2018-02-14 Carlos-Emiliano González-Gallardo , Juan-Manuel Torres-Moreno

Phrase detection requires methods to identify if a phrase is relevant to an image and localize it, if applicable. A key challenge for training more discriminative detection models is sampling negatives. Sampling techniques from prior work…

计算机视觉与模式识别 · 计算机科学 2022-11-16 Maan Qraitem , Bryan A. Plummer

We introduce Konooz, a novel multi-dimensional corpus covering 16 Arabic dialects across 10 domains, resulting in 160 distinct corpora. The corpus comprises about 777k tokens, carefully collected and manually annotated with 21 entity types…

计算与语言 · 计算机科学 2025-06-17 Nagham Hamad , Mohammed Khalilia , Mustafa Jarrar

We present DisMo, a multi-level annotator for spoken language corpora that integrates part-of-speech tagging with basic disfluency detection and annotation, and multi-word unit recognition. DisMo is a hybrid system that uses a combination…

计算与语言 · 计算机科学 2018-02-09 George Christodoulides , Mathieu Avanzi , Jean-Philippe Goldman

We introduce BERTweetFR, the first large-scale pre-trained language model for French tweets. Our model is initialized using the general-domain French language model CamemBERT which follows the base architecture of RoBERTa. Experiments show…

计算与语言 · 计算机科学 2021-09-22 Yanzhu Guo , Virgile Rennard , Christos Xypolopoulos , Michalis Vazirgiannis

We develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective sentences extracted…

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

计算与语言 · 计算机科学 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

Background: Transformer-based language models have shown strong performance on many Natural LanguageProcessing (NLP) tasks. Masked Language Models (MLMs) attract sustained interest because they can be adaptedto different languages and…

计算与语言 · 计算机科学 2024-04-01 Nesrine Bannour , Christophe Servan , Aurélie Névéol , Xavier Tannier

Claims are the central component of an argument. Detecting claims across different domains or data sets can often be challenging due to their varying conceptualization. We propose to alleviate this problem by fine tuning a language model…

计算与语言 · 计算机科学 2019-05-20 Tuhin Chakrabarty , Christopher Hidey , Kathleen McKeown

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

计算与语言 · 计算机科学 2025-10-28 Eric Jeangirard

This paper presents an embedding-based approach to detecting variation without relying on prior normalisation or predefined variant lists. The method trains subword embeddings on raw text and groups related forms through combined cosine and…

计算与语言 · 计算机科学 2026-02-13 Anne-Marie Lutgen , Alistair Plum , Christoph Purschke

Detecting toxic content using language models is crucial yet challenging. While substantial progress has been made in English, toxicity detection in French remains underdeveloped, primarily due to the lack of culturally relevant,…

计算与语言 · 计算机科学 2026-04-21 Axel Delaval , Shujian Yang , Haicheng Wang , Han Qiu , Jialiang Lu

Language models have proven to be very useful when adapted to specific domains. Nonetheless, little research has been done on the adaptation of domain-specific BERT models in the French language. In this paper, we focus on creating a…

计算与语言 · 计算机科学 2022-02-15 Stella Douka , Hadi Abdine , Michalis Vazirgiannis , Rajaa El Hamdani , David Restrepo Amariles