English
Related papers

Related papers: Tagging the Teleman Corpus

200 papers

Dividing oral histories into topically coherent segments can make them more accessible online. People regularly make judgments about where coherent segments can be extracted from oral histories. But making these judgments can be taxing, so…

Computation and Language · Computer Science 2015-09-30 Ryan Shaw

We present an evaluation of text simplification (TS) in Spanish for a production system, by means of two corpora focused in both complex-sentence and complex-word identification. We compare the most prevalent Spanish-specific readability…

Computation and Language · Computer Science 2023-08-16 Adrian de Wynter , Anthony Hevia , Si-Qing Chen

Prior studies in multilingual language modeling (e.g., Cotterell et al., 2018; Mielke et al., 2019) disagree on whether or not inflectional morphology makes languages harder to model. We attempt to resolve the disagreement and extend those…

Computation and Language · Computer Science 2021-03-29 Hyunji Hayley Park , Katherine J. Zhang , Coleman Haley , Kenneth Steimel , Han Liu , Lane Schwartz

Neural HMMs are a type of neural transducer recently proposed for sequence-to-sequence modelling in text-to-speech. They combine the best features of classic statistical speech synthesis and modern neural TTS, requiring less data and fewer…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Shivam Mehta , Ambika Kirkland , Harm Lameris , Jonas Beskow , Éva Székely , Gustav Eje Henter

We present the first parsing results on the Penn-Helsinki Parsed Corpus of Early Modern English (PPCEME), a 1.9 million word treebank that is an important resource for research in syntactic change. We describe key features of PPCEME that…

Computation and Language · Computer Science 2021-12-17 Seth Kulick , Neville Ryant , Beatrice Santorini

We study the adaptation of Link Grammar Parser to the biomedical sublanguage with a focus on domain terms not found in a general parser lexicon. Using two biomedical corpora, we implement and evaluate three approaches to addressing unknown…

Computation and Language · Computer Science 2007-05-23 Sampo Pyysalo , Tapio Salakoski , Sophie Aubin , Adeline Nazarenko

Chinese word segmentation and part-of-speech tagging are necessary tasks in terms of computational linguistics and application of natural language processing. Many re-searchers still debate the demand for Chinese word segmentation and…

Computation and Language · Computer Science 2021-12-20 Duc-Vu Nguyen , Linh-Bao Vo , Ngoc-Linh Tran , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

Using phonological speech vocoding, we propose a platform for exploring relations between phonology and speech processing, and in broader terms, for exploring relations between the abstract and physical structures of a speech signal. Our…

Computation and Language · Computer Science 2016-11-16 Milos Cernak , Stefan Benus , Alexandros Lazaridis

We describe a stochastic approach to partial parsing, i.e., the recognition of syntactic structures of limited depth. The technique utilises Markov Models, but goes beyond usual bracketing approaches, since it is capable of recognising not…

cmp-lg · Computer Science 2007-05-23 Wojciech Skut , Thorsten Brants

We describe a method for classification of handwritten Kannada characters using Hidden Markov Models (HMMs). Kannada script is agglutinative, where simple shapes are concatenated horizontally to form a character. This results in a large…

Machine Learning · Computer Science 2014-10-17 Manasij Venkatesh , Vikas Majjagi , Deepu Vijayasenan

The paper gives an example of a tree language G that is recognised by an unambiguous parity automaton and is analytic-complete as a set in Cantor space. This already shows that the unambiguous languages are topologically more complex than…

Formal Languages and Automata Theory · Computer Science 2012-10-10 Szczepan Hummel

Comparing weighted networks in neuroscience is hard, because the topological properties of a given network are necessarily dependent on the number of edges of that network. This problem arises in the analysis of both weighted and unweighted…

Applications · Statistics 2014-03-28 Cedric E. Ginestet , Arnaud P. Fournel , Andrew Simmons

This paper experiments with frequency-based corpus similarity measures across 39 languages using a register prediction task. The goal is to quantify (i) the distance between different corpora from the same language and (ii) the homogeneity…

Computation and Language · Computer Science 2022-06-10 Haipeng Li , Jonathan Dunn

We examine whether large neural language models, trained on very large collections of varied English text, learn the potentially long-distance dependency of British versus American spelling conventions, i.e., whether spelling is…

Computation and Language · Computer Science 2023-03-08 Elizabeth Nielsen , Christo Kirov , Brian Roark

Historical languages present unique challenges to the NLP community, with one prominent hurdle being the limited resources available in their closed corpora. This work describes our submission to the constrained subtask of the SIGTYP 2024…

Computation and Language · Computer Science 2024-05-31 Frederick Riemenschneider , Kevin Krahn

One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative…

We proposed new and more efficient estimators for estimating population proportion of respondents belonging to two related sensitive attributes in survey sampling by extending the work of Mangat (1994). Our proposed estimators are more…

Methodology · Statistics 2016-01-05 Olusegun S. Ewemooje , Godwin N. Amahia

The ability to compare the semantic similarity between text corpora is important in a variety of natural language processing applications. However, standard methods for evaluating these metrics have yet to be established. We propose a set…

Computation and Language · Computer Science 2022-11-30 George Kour , Samuel Ackerman , Orna Raz , Eitan Farchi , Boaz Carmeli , Ateret Anaby-Tavor

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have demonstrated significant capabilities across numerous applications. However, the performance of these models in languages with fewer resources, such…

Computation and Language · Computer Science 2024-05-24 Birger Moell

In Neural Machine Translation (NMT), the decoder can capture the features of the entire prediction history with neural connections and representations. This means that partial hypotheses with different prefixes will be regarded differently…

Computation and Language · Computer Science 2018-10-16 Zhisong Zhang , Rui Wang , Masao Utiyama , Eiichiro Sumita , Hai Zhao