中文
相关论文

相关论文: WEKA-Based: Key Features and Classifier for French…

200 篇论文

Benchmarking plays a pivotal role in assessing and enhancing the performance of compact deep learning models designed for execution on resource-constrained devices, such as microcontrollers. Our study introduces a novel, entirely…

声音 · 计算机科学 2024-03-18 René Groh , Nina Goes , Andreas M. Kist

We present a set of probabilistic models applied to binary classification as defined in the DEFT'05 challenge. The challenge consisted a mixture of two differents problems in Natural Language Processing : identification of author (a…

计算与语言 · 计算机科学 2019-03-19 Marc El-Bèze , Juan-Manuel Torres-Moreno , Frédéric Béchet

We explore one aspect of the structure of a codified legal system at the national level using a new type of representation to understand the strong or weak dependencies between the various fields of law. In Part I of this study, we analyze…

社会与信息网络 · 计算机科学 2012-01-06 Romain Boulet

The paper reports on a series of experiments aiming at probing LeBenchmark, a pretrained acoustic model trained on 7k hours of spoken French, for syntactic information. Pretrained acoustic models are increasingly used for downstream speech…

计算与语言 · 计算机科学 2024-03-05 Zdravko Dugonjić , Adrien Pupier , Benjamin Lecouteux , Maximin Coavoux

In this paper we describe our efforts to make a bidirectional Congolese Swahili (SWC) to French (FRA) neural machine translation system with the motivation of improving humanitarian translation workflows. For training, we created a…

计算与语言 · 计算机科学 2021-03-22 Alp Öktem , Eric DeLuca , Rodrigue Bashizi , Eric Paquin , Grace Tang

We earlier described two taggers for French, a statistical one and a constraint-based one. The two taggers have the same tokeniser and morphological analyser. In this paper, we describe aspects of this work concerned with the definition of…

cmp-lg · 计算机科学 2008-02-03 Jean-Pierre Chanod , Pasi Tapanainen

In this paper, we introduce the French-YMCA corpus, a new linguistic resource specifically tailored for children and adolescents. The motivation for building this corpus is clear: children have unique language requirements, as their…

计算与语言 · 计算机科学 2026-04-08 Cherifa Ben Khelil , Jean-Yves Antoine , Anaïs Halftermeyer , Frédéric Rayar , Mathieu Thebaud

Progress in cross-lingual modeling depends on challenging, realistic, and diverse evaluation sets. We introduce Multilingual Knowledge Questions and Answers (MKQA), an open-domain question answering evaluation set comprising 10k…

计算与语言 · 计算机科学 2021-08-18 Shayne Longpre , Yi Lu , Joachim Daiber

In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed to discriminate…

计算与语言 · 计算机科学 2018-08-15 Liviu P. Dinu , Alina Maria Ciobanu , Marcos Zampieri , Shervin Malmasi

We extract a large-scale stance detection dataset from comments written by candidates of elections in Switzerland. The dataset consists of German, French and Italian text, allowing for a cross-lingual evaluation of stance detection. It…

计算与语言 · 计算机科学 2020-06-11 Jannis Vamvas , Rico Sennrich

In this work we tackle the problem of sentence boundary detection applied to French as a binary classification task ("sentence boundary" or "not sentence boundary"). We combine convolutional neural networks with subword-level information…

计算与语言 · 计算机科学 2018-02-14 Carlos-Emiliano González-Gallardo , Juan-Manuel Torres-Moreno

Context-aware Machine Translation aims to improve translations of sentences by incorporating surrounding sentences as context. Towards this task, two main architectures have been applied, namely single-encoder (based on concatenation) and…

计算与语言 · 计算机科学 2024-02-05 Paweł Mąka , Yusuf Can Semerci , Jan Scholtes , Gerasimos Spanakis

We propose a theoretical framework within which information on the vocabulary of a given corpus can be inferred on the basis of statistical information gathered on that corpus. Inferences can be made on the categories of the words in the…

计算与语言 · 计算机科学 2008-10-08 Pascal Vaillant , Richard Nock , Claudia Henry

This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological…

计算与语言 · 计算机科学 2025-12-09 Revekka Kyriakoglou , Anna Pappa

This paper evaluates global-scale dialect identification for 14 national varieties of English as a means for studying syntactic variation. The paper makes three main contributions: (i) introducing data-driven language mapping as a method…

计算与语言 · 计算机科学 2019-04-12 Jonathan Dunn

We release our synthetic parallel paraphrase corpus across 17 languages: Arabic, Catalan, Czech, German, English, Spanish, Estonian, French, Hindi, Indonesian, Italian, Dutch, Romanian, Russian, Swedish, Vietnamese, and Chinese. Our method…

Etiquettes are an essential ingredient of day-to-day interactions among people. Moreover, etiquettes are region-specific, and etiquettes in one region might contradict those in other regions. In this paper, we propose EtiCor, an Etiquettes…

计算与语言 · 计算机科学 2023-10-31 Ashutosh Dwivedi , Pradhyumna Lavania , Ashutosh Modi

Measuring a document's complexity level is an open challenge, particularly when one is working on a diverse corpus of documents rather than comparing several documents on a similar topic or working on a language other than English. In this…

计算与语言 · 计算机科学 2022-09-01 Vincent Primpied , David Beauchemin , Richard Khoury

This paper proposes the first multilingual (French, English and Arabic) and multicultural (Indo-European languages vs. less culturally close languages) irony detection system. We employ both feature-based models and neural architectures…

计算与语言 · 计算机科学 2020-02-07 Bilal Ghanem , Jihen Karoui , Farah Benamara , Paolo Rosso , Véronique Moriceau

We present DisMo, a multi-level annotator for spoken language corpora that integrates part-of-speech tagging with basic disfluency detection and annotation, and multi-word unit recognition. DisMo is a hybrid system that uses a combination…

计算与语言 · 计算机科学 2018-02-09 George Christodoulides , Mathieu Avanzi , Jean-Philippe Goldman