中文
相关论文

相关论文: FreCDo: A Large Corpus for French Cross-Domain Dia…

200 篇论文

Multilingual acoustic models have been successfully applied to low-resource speech recognition. Most existing works have combined many small corpora together and pretrained a multilingual model by sampling from each corpus uniformly. The…

计算与语言 · 计算机科学 2019-08-06 Xinjian Li , Siddharth Dalmia , Alan W. Black , Florian Metze

This study explores the integration of Building Information Modeling (BIM) with Natural Language Processing (NLP) to automate the extraction of requirements from unstructured French Building Technical Specification (BTS) documents within…

计算与语言 · 计算机科学 2025-08-20 Insaf Nahri , Romain Pinquié , Philippe Véron , Nicolas Bus , Mathieu Thorel

Automated terminology extraction refers to the task of extracting meaningful terms from domain-specific texts. This paper proposes a novel machine learning approach to terminology extraction, which combines features from traditional term…

计算与语言 · 计算机科学 2025-02-25 Andraž Repar , Nada Lavrač , Senja Pollak

This paper introduces a computational framework designed to delineate gender distribution biases in topics covered by French TV and radio news. We transcribe a dataset of 11.7k hours, broadcasted in 2023 on 21 French channels. A Large…

计算与语言 · 计算机科学 2024-07-22 Valentin Pelloin , Lena Dodson , Émile Chapuis , Nicolas Hervé , David Doukhan

Achieving consistent high-quality machine translation (MT) across diverse domains remains a significant challenge, primarily due to the limited and imbalanced parallel training data available in various domains. While large language models…

计算与语言 · 计算机科学 2024-10-04 Tianxiang Hu , Pei Zhang , Baosong Yang , Jun Xie , Derek F. Wong , Rui Wang

Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models. In an effort towards sustainable practices, we study the impact of pre-training data volume on compact language…

计算与语言 · 计算机科学 2020-10-12 Vincent Micheli , Martin d'Hoffschmidt , François Fleuret

In this paper, we introduce SaudiBERT, a monodialect Arabic language model pretrained exclusively on Saudi dialectal text. To demonstrate the model's effectiveness, we compared SaudiBERT with six different multidialect Arabic language…

计算与语言 · 计算机科学 2024-05-13 Faisal Qarah

The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. Text-audio retrieval…

音频与语音处理 · 电气工程与系统科学 2022-02-11 A. Sophia Koepke , Andreea-Maria Oncescu , João F. Henriques , Zeynep Akata , Samuel Albanie

This paper measures similarity both within and between 84 language varieties across nine languages. These corpora are drawn from digital sources (the web and tweets), allowing us to evaluate whether such geo-referenced corpora are reliable…

计算与语言 · 计算机科学 2021-04-06 Jonathan Dunn

A radio speech corpus of 9mn has been prosodically marked by a phonetician expert, and non expert listeners. this corpus is large enough to train and test an automatic boundary spotting system, namely a time delay neural network fed with F0…

cmp-lg · 计算机科学 2007-05-23 V. Pagel , N. Carbonell , Y. Laprie , J. Vaissiere

We present a dictionary-based approach to racism detection in Dutch social media comments, which were retrieved from two public Belgian social media sites likely to attract racist reactions. These comments were labeled as racist or…

计算与语言 · 计算机科学 2016-09-01 Stéphan Tulkens , Lisa Hilte , Elise Lodewyckx , Ben Verhoeven , Walter Daelemans

This paper presents our system for SemEval 2025 Task 11: Bridging the Gap in Text-Based Emotion Detection (Track A), which focuses on multi-label emotion detection in short texts. We propose a feature-centric framework that dynamically…

计算与语言 · 计算机科学 2026-02-05 Ziyi Huang , Xia Cui

To build a satisfying chatbot that has the ability of managing a goal-oriented multi-turn dialogue, accurate modeling of human conversation is crucial. In this paper we concentrate on the task of response selection for multi-turn…

计算与语言 · 计算机科学 2018-02-19 Guozhen An , Mehrnoosh Shafiee , Davood Shamsi

Existing research on fairness evaluation of document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. In this work, we assemble and publish a multilingual Twitter corpus…

计算与语言 · 计算机科学 2020-03-04 Xiaolei Huang , Linzi Xing , Franck Dernoncourt , Michael J. Paul

Large Language Models (LLMs) are increasingly leveraged for translation tasks but often fall short when translating inclusive language -- such as texts containing the singular 'they' pronoun or otherwise reflecting fair linguistic…

计算与语言 · 计算机科学 2025-05-06 Fanny Jourdan , Yannick Chevalier , Cécile Favre

Norway has a large amount of dialectal variation, as well as a general tolerance to its use in the public sphere. There are, however, few available resources to study this variation and its change over time and in more informal areas, \eg…

计算与语言 · 计算机科学 2021-04-13 Jeremy Barnes , Petter Mæhlum , Samia Touileb

Subjective bias detection is critical for applications like propaganda detection, content recommendation, sentiment analysis, and bias neutralization. This bias is introduced in natural language via inflammatory words and phrases, casting…

计算与语言 · 计算机科学 2020-06-16 Tanvi Dadu , Kartikey Pant , Radhika Mamidi

We introduce FaBERT, a Persian BERT-base model pre-trained on the HmBlogs corpus, encompassing both informal and formal Persian texts. FaBERT is designed to excel in traditional Natural Language Understanding (NLU) tasks, addressing the…

计算与语言 · 计算机科学 2024-02-12 Mostafa Masumi , Seyed Soroush Majd , Mehrnoush Shamsfard , Hamid Beigy

Verbal deception has been studied in psychology, forensics, and computational linguistics for a variety of reasons, like understanding behaviour patterns, identifying false testimonies, and detecting deception in online communication.…

计算与语言 · 计算机科学 2023-06-09 Aswathy Velutharambath , Roman Klinger

Spread of fake news using out-of-context images and captions has become widespread in this era of information overload. Since fake news can belong to different domains like politics, sports, etc. with their unique characteristics, inference…

机器学习 · 计算机科学 2025-01-08 Amartya Bhattacharya , Debarshi Brahma , Suraj Nagaje Mahadev , Anmol Asati , Vikas Verma , Soma Biswas