中文
相关论文

相关论文: MOROCO: The Moldavian and Romanian Dialectal Corpu…

200 篇论文

We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous…

计算与语言 · 计算机科学 2018-03-26 Mikhail Khodak , Nikunj Saunshi , Kiran Vodrahalli

AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce for many languages and educational systems. We introduce RoMathExam, a longitudinal…

计算机与社会 · 计算机科学 2026-04-21 Luca-Ncolae Cuclea , Sabin-Codrut Badea , Adrian-Marius Dumitran

This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.…

计算与语言 · 计算机科学 2020-11-24 William Havard , Laurent Besacier , Olivier Rosec

Running large-scale pre-trained language models in computationally constrained environments remains a challenging problem yet to be addressed, while transfer learning from these models has become prevalent in Natural Language Processing…

Current research into spoken language translation (SLT),or speech-to-text translation, is often hampered by the lack of specific data resources for this task, as currently available SLT datasets are restricted to a limited set of language…

This paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South Slavic language…

计算与语言 · 计算机科学 2024-05-28 Nikola Ljubešić , Taja Kuzman

The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. Text-audio retrieval…

音频与语音处理 · 电气工程与系统科学 2022-02-11 A. Sophia Koepke , Andreea-Maria Oncescu , João F. Henriques , Zeynep Akata , Samuel Albanie

This paper describes the Spot the Difference Corpus which contains 54 interactions between pairs of subjects interacting to find differences in two very similar scenes. The setup used, the participants' metadata and details about collection…

计算与语言 · 计算机科学 2018-05-15 José Lopes , Nils Hemmingsson , Oliver Åstrand

This paper measures similarity both within and between 84 language varieties across nine languages. These corpora are drawn from digital sources (the web and tweets), allowing us to evaluate whether such geo-referenced corpora are reliable…

计算与语言 · 计算机科学 2021-04-06 Jonathan Dunn

We present a new, unique and freely available parallel corpus containing European Union (EU) documents of mostly legal nature. It is available in all 20 official EUanguages, with additional documents being available in the languages of the…

计算与语言 · 计算机科学 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Anna Widiger , Camelia Ignat , Tomaz Erjavec , Dan Tufis , Daniel Varga

Thanks to the Eslo1 ("Enqu\^ete sociolinguistique d'Orl\'eans", i.e. "Sociolinguistic Inquiery of Orl\'eans") campain, a large oral corpus has been gathered and transcribed in a textual format. The purpose of the work presented here is to…

机器学习 · 计算机科学 2010-03-31 Iris Eshkol , Isabelle Tellier , Taalab Samer , Sylvie Billot

This thesis presents a constraint-based morphological disambiguation approach that is applicable to languages with complex morphology--specifically agglutinative languages with productive inflectional and derivational morphological…

cmp-lg · 计算机科学 2008-02-03 Gokhan Tur

Memes are becoming increasingly more popular in online media, especially in social networks. They usually combine graphical representations (images, drawings, animations or video) with text to convey powerful messages. In order to extract,…

计算与语言 · 计算机科学 2024-10-22 Vasile Păiş , Sara Niţă , Alexandru-Iulius Jerpelea , Luca Pană , Eric Curea

While there has been a surge of large language models for Norwegian in recent years, we lack any tool to evaluate their understanding of grammaticality. We present two new Norwegian datasets for this task. NoCoLA_class is a supervised…

计算与语言 · 计算机科学 2023-06-14 Matias Jentoft , David Samuel

Audio-text retrieval is a challenging task, requiring the search for an audio clip or a text caption within a database. The predominant focus of existing research on English descriptions poses a limitation on the applicability of such…

声音 · 计算机科学 2024-06-18 Zhiyong Yan , Heinrich Dinkel , Yongqing Wang , Jizhong Liu , Junbo Zhang , Yujun Wang , Bin Wang

Data scarcity has been a long standing issue in the field of open-domain social dialogue. To quench this thirst, we present SODA: the first publicly available, million-scale high-quality social dialogue dataset. By contextualizing social…

We present the second ever evaluated Arabic dialect-to-dialect machine translation effort, and the first to leverage external resources beyond a small parallel corpus. The subject has not previously received serious attention due to lack of…

计算与语言 · 计算机科学 2017-12-19 Alexander Erdmann , Nizar Habash , Dima Taji , Houda Bouamor

Pre-trained Language Models (PLMs) have shown remarkable performances in recent years, setting a new paradigm for NLP research and industry. The legal domain has received some attention from the NLP community partly due to its textual…

We present M2D2, a fine-grained, massively multi-domain corpus for studying domain adaptation in language models (LMs). M2D2 consists of 8.5B tokens and spans 145 domains extracted from Wikipedia and Semantic Scholar. Using ontologies…

计算与语言 · 计算机科学 2022-10-17 Machel Reid , Victor Zhong , Suchin Gururangan , Luke Zettlemoyer

The explosion in the amount of news and journalistic content being generated across the globe, coupled with extended and instantaneous access to information through online media, makes it difficult and time-consuming to monitor news…

计算与语言 · 计算机科学 2018-08-06 M. Tarik Altuncu , Sophia N. Yaliraki , Mauricio Barahona