English
Related papers

Related papers: Tagging the Teleman Corpus

200 papers

We present a dataset for evaluating the grammaticality of the predictions of a language model. We automatically construct a large number of minimally different pairs of English sentences, each consisting of a grammatical and an…

Computation and Language · Computer Science 2018-08-29 Rebecca Marvin , Tal Linzen

Old French is a typical example of an under-resourced historic languages, that furtherly displays animportant amount of linguistic variation. In this paper, we present the current results of a long going project (2015-...) and describe how…

Computation and Language · Computer Science 2021-09-24 Jean-Baptiste Camps , Thibault Clérice , Frédéric Duval , Lucence Ing , Naomi Kanaoka , Ariane Pinche

While modern masked language models (LMs) are trained on ever larger corpora, we here explore the effects of down-scaling training to a modestly-sized but representative, well-balanced, and publicly available English text source -- the…

Computation and Language · Computer Science 2023-05-09 David Samuel , Andrey Kutuzov , Lilja Øvrelid , Erik Velldal

Recent work raises concerns about the use of standard splits to compare natural language processing models. We propose a Bayesian statistical model comparison technique which uses k-fold cross-validation across multiple data sets to…

Computation and Language · Computer Science 2020-10-08 Piotr Szymański , Kyle Gorman

This project explores methods to enhance sign language translation of German sign language, specifically focusing on disambiguation of homonyms. Sign language is ambiguous and understudied which is the basis for our experiments. We approach…

Computation and Language · Computer Science 2024-09-16 Jana Grimm , Miriam Winkler , Oliver Kraus , Tanalp Agustoslu

We present new supertaggers trained on English grammar-based treebanks and test the effects of the best tagger on parsing speed and accuracy. The treebanks are produced automatically by large manually built grammars and feature high-quality…

Computation and Language · Computer Science 2024-10-10 Olga Zamaraeva , Carlos Gómez-Rodríguez

We introduce a memory-based approach to part of speech tagging. Memory-based learning is a form of supervised learning based on similarity-based reasoning. The part of speech tag of a word in a particular context is extrapolated from the…

cmp-lg · Computer Science 2008-02-03 Walter Daelemans , Jakub Zavrel , Peter Berck , Steven Gillis

Probabilistic approaches to part-of-speech tagging rely primarily on whole-word statistics about word/tag combinations as well as contextual information. But experience shows about 4 per cent of tokens encountered in test sets are unknown…

Computation and Language · Computer Science 2013-02-28 Greg Adams , Beth Millar , Eric Neufeld , Tim Philip

This paper describes our submission to CoNLL 2018 UD Shared Task. We have extended an LSTM-based neural network designed for sequence tagging to additionally generate character-level sequences. The network was jointly trained to produce…

Computation and Language · Computer Science 2018-09-11 Gor Arakelyan , Karen Hambardzumyan , Hrant Khachatrian

This paper investigates neural character-based morphological tagging for languages with complex morphology and large tag sets. We systematically explore a variety of neural architectures (DNN, CNN, CNNHighway, LSTM, BLSTM) to obtain…

Computation and Language · Computer Science 2016-06-22 Georg Heigold , Guenter Neumann , Josef van Genabith

We evaluate a battery of recent large language models on two benchmarks for word sense disambiguation in Swedish. At present, all current models are less accurate than the best supervised disambiguators in cases where a training set is…

Computation and Language · Computer Science 2024-10-31 Richard Johansson

Improvement in machine learning-based NLP performance are often presented with bigger models and more complex code. This presents a trade-off: better scores come at the cost of larger tools; bigger models tend to require more during…

Computation and Language · Computer Science 2021-04-19 Magnus Jacobsen , Mikkel H. Sørensen , Leon Derczynski

We present a bootstrapping method to develop an annotated corpus, which is specially useful for languages with few available resources. The method is being applied to develop a corpus of Spanish of over 5Mw. The method consists on taking…

Computation and Language · Computer Science 2007-05-23 L. Marquez , L. Padro , H. Rodriguez

We present an extended comparison of contextualized language models for Hungarian. We compare huBERT, a Hungarian model against 4 multilingual models including the multilingual BERT model. We evaluate these models through three tasks,…

Computation and Language · Computer Science 2021-02-23 Judit Ács , Dániel Lévai , Dávid Márk Nemeskey , András Kornai

Machine translation is evolving quite rapidly in terms of quality. Nowadays, we have several machine translation systems available in the web, which provide reasonable translations. However, these systems are not perfect, and their quality…

Computation and Language · Computer Science 2015-10-16 Krzysztof Wołk , Krzysztof Marasek , Wojciech Glinkowski

While LLMs have shown great success in understanding and generating text in traditional conversational settings, their potential for performing ill-defined complex tasks is largely under-studied. Indeed, we are yet to conduct comprehensive…

Artificial Intelligence · Computer Science 2023-10-26 Shubhra Kanti Karmaker Santu , Dongji Feng

Subword tokenization has become the de-facto standard for tokenization, although comparative evaluations of subword vocabulary quality across languages are scarce. Existing evaluation studies focus on the effect of a tokenization algorithm…

Computation and Language · Computer Science 2023-10-23 Lisa Beinborn , Yuval Pinter

This paper presents a new approach for measuring semantic similarity/distance between words and concepts. It combines a lexical taxonomy structure with corpus statistical information so that the semantic distance between nodes in the…

cmp-lg · Computer Science 2008-02-03 Jay J. Jiang , David W. Conrath

Text-dependent speaker verification is becoming popular in the speaker recognition society. However, the conventional i-vector framework which has been successful for speaker identification and other similar tasks works relatively poorly in…

Sound · Computer Science 2017-09-12 Yi Liu , Liang He , Yao Tian , Zhuzi Chen , Jia Liu , Michael T. Johnson

This paper provides a comparative analysis of the performance of four state-of-the-art distributional semantic models (DSMs) over 11 languages, contrasting the native language-specific models with the use of machine translation over…

Computation and Language · Computer Science 2018-05-18 Andre Freitas , Siamak Barzegar , Juliano Efson Sales , Siegfried Handschuh , Brian Davis