English
Related papers

Related papers: The origin of Mayan languages from Formosan langua…

200 papers

While large language models (LLMs) have demonstrated impressive performance across a wide range of natural language processing (NLP) tasks in high-resource languages, their capabilities in low-resource and minority languages remain…

Computation and Language · Computer Science 2025-06-30 Kaiying Kevin Lin , Hsiyu Chen , Haopeng Zhang

The origin of Malagasy DNA is half African and half Indonesian, nevertheless the Malagasy language, spoken by the entire population, belongs to the Austronesian family. The language most closely related to Malagasy is Maanyan (Greater…

Computation and Language · Computer Science 2011-02-15 M. Serva , F. Petroni , D. Volchenkov , S. Wichmann

The Malagasy language belongs to the Greater Barito East group of the Austronesian family, the language most closely connected to Malagasy dialects is Maanyan (Kalimantan), but Malay as well other Indonesian and Philippine languages are…

Populations and Evolution · Quantitative Biology 2018-03-07 Maurizio Serva

Basic vocabulary in many Sulawesi Austronesian languages includes forms resisting reconstruction to any proto-form with phonological patterns inconsistent with inherited roots, but whether this non-conforming vocabulary represents…

Computation and Language · Computer Science 2026-04-02 Mukhlis Amien , Go Frendi Gunawan

This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages. Accurate language identification is an important part of…

Computation and Language · Computer Science 2022-06-10 Jonathan Dunn , Wikke Nijhof

The Mayan languages comprise a language family with an ancient history, millions of speakers, and immense cultural value, that, nevertheless, remains severely underrepresented in terms of resources and global exposure. In this paper we…

Computation and Language · Computer Science 2024-06-18 Andrés Lou , Juan Antonio Pérez-Ortiz , Felipe Sánchez-Martínez , Víctor M. Sánchez-Cartagena

We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP competencies: understanding (NLU), reasoning (NLR), and generation (NLG). BURMESE-SAN…

Computation and Language · Computer Science 2026-05-25 Thura Aung , Jann Railey Montalan , Jian Gang Ngui , Peerat Limkonchotiwat

The dialects of Madagascar belong to the Greater Barito East group of the Austronesian family and it is widely accepted that the Island was colonized by Indonesian sailors after a maritime trek which probably took place around 650 CE. The…

Computation and Language · Computer Science 2015-05-28 Maurizio Serva

Tokenization constitutes a fundamental stage in Large Language Model (LLM) processing; however, subword-based tokenization methods optimized on English-dominant corpora may produce token fragmentation misaligned with the linguistic…

Computers and Society · Computer Science 2026-02-10 Andhika Bernard Lumbantobing , Hokky Situngkir

Phylogenetic methods have broad potential in linguistics beyond tree inference. Here, we show how a phylogenetic approach opens the possibility of gaining historical insights from entirely new kinds of linguistic data--in this instance,…

Computation and Language · Computer Science 2021-05-10 Jayden L. Macklin-Cordes , Claire Bowern , Erich R. Round

The study of spoken languages comprises phonology, morphology, and grammar. The languages can be classified as root languages, inflectional languages, and stem languages. In addition, languages continually change over time and space by…

Computation and Language · Computer Science 2025-11-05 Shreekanth M Prabhu , Abhisek Midya

The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for…

Computation and Language · Computer Science 2025-10-13 Stefan Krsteski , Matea Tashkovska , Borjan Sazdov , Hristijan Gjoreski , Branislav Gerazov

Non-M\=aori-speaking New Zealanders (NMS)are able to segment M\=aori words in a highlysimilar way to fluent speakers (Panther et al.,2024). This ability is assumed to derive through the identification and extraction of statistically…

Computation and Language · Computer Science 2024-03-22 Ashvini Varatharaj , Simon Todd

Australian Aboriginal languages are of significant cultural and linguistic value but remain severely underrepresented in modern speech AI systems. While state-of-the-art speech foundation models and automatic speech recognition excel in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Ting Dang , Trini Manoj Jeyaseelan , Eliathamby Ambikairajah , Vidhyasaharan Sethu

We present a large-scale comparative study of 242 Latin and Cyrillic-script languages using subword-based methodologies. By constructing 'glottosets' from Wikipedia lexicons, we introduce a framework for simultaneous cross-linguistic…

Computation and Language · Computer Science 2026-01-27 Iaroslav Chelombitko , Mika Hämäläinen , Aleksey Komissarov

Languages evolve over time in a process in which reproduction, mutation and extinction are all possible, similar to what happens to living organisms. Using this similarity it is possible, in principle, to build family trees which show the…

Computation and Language · Computer Science 2012-07-03 Maurizio Serva

The morphology of selected groups of sources in the FIRST (Faint Images of the Radio Sky at Twenty Centimeters) survey and catalog is examined. Sources in the FIRST catalog (April 2003 release, 811117 entries) were sorted into singles,…

Cosmology and Nongalactic Astrophysics · Physics 2015-05-27 D. D. Proctor

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces Aya, a massively…

Addressing the gap in Large Language Model pretrained from scratch with Malaysian context, We trained models with 1.1 billion, 3 billion, and 5 billion parameters on a substantial 349GB dataset, equivalent to 90 billion tokens based on our…

Computation and Language · Computer Science 2024-01-30 Husein Zolkepli , Aisyah Razak , Kamarul Adha , Ariff Nazhan

A $q$-ary maximum distance separable (MDS) code $C$ with length $n$, dimension $k$ over an alphabet $\mathcal{A}$ of size $q$ is a set of $q^k$ codewords that are elements of $\mathcal{A}^n$, such that the Hamming distance between two…

Combinatorics · Mathematics 2015-04-28 Janne I. Kokkala , Patric R. J. Östergård
‹ Prev 1 2 3 10 Next ›