English
Related papers

Related papers: Language discrimination and clustering via a neura…

200 papers

Cross-lingual language tasks typically require a substantial amount of annotated data or parallel translation data. We explore whether language representations that capture relationships among languages can be learned and subsequently…

Computation and Language · Computer Science 2021-06-07 Dian Yu , Taiqi He , Kenji Sagae

In natural speech, the speaker does not pause between words, yet a human listener somehow perceives this continuous stream of phonemes as a series of distinct words. The detection of boundaries between spoken words is an instance of a…

Computation and Language · Computer Science 2011-06-28 Jerry R. Van Aken

We present our experience in applying distributional semantics (neural word embeddings) to the problem of representing and clustering documents in a bilingual comparable corpus. Our data is a collection of Russian and Ukrainian academic…

Computation and Language · Computer Science 2016-04-20 Andrey Kutuzov , Mikhail Kopotev , Tatyana Sviridenko , Lyubov Ivanova

The evolution of languages closely resembles the evolution of haploid organisms. This similarity has been recently exploited \cite{GA,GJ} to construct language trees. The key point is the definition of a distance among all pairs of…

Physics and Society · Physics 2009-11-13 Maurizio Serva , Filippo Petroni

Language is one of the most important aspects of human cognition; it represents the way we think, act and communicate with each other. Each language has its own history, background, and form. A language represents a lot of important…

Adaptation and Self-Organizing Systems · Physics 2007-09-18 Dorina Strori , Ahmet Bombaci , Haluk Bingol

We use a c-GAN (convolutional generative adversarial) neural network to analyze transliterated text fragments of extant, dead comprehensible, and one dead non-deciphered (Cypro-Minoan) language to establish linguistic affinities. The paper…

Computation and Language · Computer Science 2025-05-23 Peter B. Lerner

Clustering news across languages enables efficient media monitoring by aggregating articles from multilingual sources into coherent stories. Doing so in an online setting allows scalable processing of massive news streams. To this end, we…

Computation and Language · Computer Science 2018-09-05 Sebastião Miranda , Artūrs Znotiņš , Shay B. Cohen , Guntis Barzdins

Several computational models have been developed to detect and analyze dialect variation in recent years. Most of these models assume a predefined set of geographical regions over which they detect and analyze dialectal variation. However,…

Computation and Language · Computer Science 2019-10-17 Hang Jiang , Haoshen Hong , Yuxing Chen , Vivek Kulkarni

A novel type of permutation tests for dendrogram data is studied with respect to two types of metrics for measuring the difference between dendrograms. First, the Frobenius norm is used, and we prove the consistency and efficiency of the…

Statistics Theory · Mathematics 2014-03-13 Kei Kobayashi , Mitsuru Orita

Segmenting text into Elemental Discourse Units (EDUs) is a fundamental task in discourse parsing. We present a new simple method for identifying EDU boundaries, and hence segmenting them, based on lexical and character n-gram features,…

Computation and Language · Computer Science 2025-01-15 Mohammadreza Sediqin , Shlomo Engelson Argamon

Subwords have become the standard units of text in NLP, enabling efficient open-vocabulary models. With algorithms like byte-pair encoding (BPE), subword segmentation is viewed as a preprocessing step applied to the corpus before training.…

Computation and Language · Computer Science 2022-10-14 Francois Meyer , Jan Buys

This article presents a complete process to extract hypernym relationships in the field of construction using two main steps: terminology extraction and detection of hypernyms from these terms. We first describe the corpus analysis method…

Artificial Intelligence · Computer Science 2025-01-15 Rémy Kessler , Nicolas Béchet

This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages. Accurate language identification is an important part of…

Computation and Language · Computer Science 2022-06-10 Jonathan Dunn , Wikke Nijhof

A clustering algorithm based on the Hausdorff distance is introduced and compared to the single and complete linkage. The three clustering procedures are applied to a toy example and to the time series of financial data. The dendrograms are…

Statistical Finance · Quantitative Finance 2010-01-30 N. Basalto , R. Bellotti , F. De Carlo , P. Facchi , E. Pantaleo , S. Pascazio

Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual evaluation of large language models. Despite…

Computation and Language · Computer Science 2026-02-20 Clara Meister , Ahmetcan Yavuz , Pietro Lesci , Tiago Pimentel

Tokenization plays a pivotal role in multilingual NLP. However, existing tokenizers are often skewed towards high-resource languages, limiting their effectiveness for linguistically diverse and morphologically rich languages such as those…

Computation and Language · Computer Science 2025-06-25 N J Karthika , Maharaj Brahma , Rohit Saluja , Ganesh Ramakrishnan , Maunendra Sankar Desarkar

Deception detection is a task with many applications both in direct physical and in computer-mediated communication. Our focus is on automatic deception detection in text across cultures. We view culture through the prism of the…

Computation and Language · Computer Science 2021-05-27 Katerina Papantoniou , Panagiotis Papadakos , Theodore Patkos , Giorgos Flouris , Ion Androutsopoulos , Dimitris Plexousakis

Recent LLM benchmarks have tested models on a range of phenomena, but are still focused primarily on natural language understanding for extraction of explicit information, such as QA or summarization, with responses often targeting…

Computation and Language · Computer Science 2025-11-11 Lanni Bu , Lauren Levine , Amir Zeldes

Speech-based clinical tools are increasingly deployed in multilingual settings, yet whether pathological speech markers remain geometrically separable from accent variation remains unclear. Systems may misclassify healthy non-native…

Sound · Computer Science 2026-02-25 Bipasha Kashyap , Pubudu N. Pathirana

We consider a language together with the subword relation, the cover relation, and regular predicates. For such structures, we consider the extension of first-order logic by threshold- and modulo-counting quantifiers. Depending on the…

Formal Languages and Automata Theory · Computer Science 2019-01-09 Dietrich Kuske , Georg Zetzsche