English
Related papers

Related papers: Quantifying Phonosemantic Iconicity Distributional…

200 papers

In the development of Large Language Models (LLMs), considerable attention has been given to the quality of training datasets. However, the role of tokenizers in the LLM training pipeline, particularly for multilingual models, has received…

Computation and Language · Computer Science 2024-10-18 Iaroslav Chelombitko , Egor Safronov , Aleksey Komissarov

Compositionality in language refers to how much the meaning of some phrase can be decomposed into the meaning of its constituents and the way these constituents are combined. Based on the premise that substitution by synonyms is…

Computation and Language · Computer Science 2017-03-13 Christina Lioma , Niels Dalum Hansen

Text-to-speech (TTS) systems typically require high-quality studio data and accurate transcriptions for training. India has 1369 languages, with 22 official using 13 scripts. Training a TTS system for all these languages, most of which have…

Computation and Language · Computer Science 2025-06-05 Utkarsh Pathak , Chandra Sai Krishna Gunda , Anusha Prakash , Keshav Agarwal , Hema A. Murthy

Machine learning techniques have shown their competence for representing and reasoning in symbolic systems such as language and phonology. In Sinitic Historical Phonology, notable tasks that could benefit from machine learning include the…

Computation and Language · Computer Science 2023-07-06 Zhibai Jia

In cross-lingual language models, representations for many different languages live in the same space. Here, we investigate the linguistic and non-linguistic factors affecting sentence-level alignment in cross-lingual pretrained language…

Computation and Language · Computer Science 2021-09-15 Alex Jones , William Yang Wang , Kyle Mahowald

This paper connects a series of papers dealing with taxonomic word embeddings. It begins by noting that there are different types of semantic relatedness and that different lexical representations encode different forms of relatedness. A…

Computation and Language · Computer Science 2020-02-19 Magdalena Kacmajor , John D. Kelleher , Filip Klubicka , Alfredo Maldonado

Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grounded corpus of…

Computation and Language · Computer Science 2026-02-16 Adib Sakhawat , Shamim Ara Parveen , Md Ruhul Amin , Shamim Al Mahmud , Md Saiful Islam , Tahera Khatun

Non-native speakers show difficulties with spoken word processing. Many studies attribute these difficulties to imprecise phonological encoding of words in the lexical memory. We test an alternative hypothesis: that some of these…

Computation and Language · Computer Science 2021-03-12 Yevgen Matusevych , Herman Kamper , Thomas Schatz , Naomi H. Feldman , Sharon Goldwater

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Model (LLM) development. Existing tokenizers often exhibit high…

Puns represent a typical linguistic phenomenon that exploits polysemy and phonetic ambiguity to generate humour, posing unique challenges for natural language understanding. Within pun research, audio plays a central role in human…

Tone is a prosodic feature used to distinguish words in many languages, some of which are endangered and scarcely documented. In this work, we use unsupervised representation learning to identify probable clusters of syllables that share…

Sound · Computer Science 2020-05-18 Bai Li , Jing Yi Xie , Frank Rudzicz

Sycophantic response patterns in Large Language Models (LLMs) have been increasingly claimed in the literature. We review methodological challenges in measuring LLM sycophancy and identify five core operationalizations. Despite sycophancy…

Computation and Language · Computer Science 2025-12-02 Jan Batzner , Volker Stocker , Stefan Schmid , Gjergji Kasneci

Causal processes can give rise to distinctive distributions in the linguistic variables that they affect. Consequently, a secure understanding of a variable's distribution can hold a key to understanding the forces that have causally shaped…

Computation and Language · Computer Science 2021-01-01 Jayden L. Macklin-Cordes , Erich R. Round

Human languages expand vocabularies by combining existing morphemes rather than inventing arbitrary forms. Communicative efficiency shapes lexical systems at multiple levels (Gibson et al., 2019), yet morphological composition -- combining…

Computation and Language · Computer Science 2026-05-06 Fengyuan Yang , Yongqian Peng , Yuxi Ma , Chenheng Xu , Yixin Zhu

Languages are grouped into families that share common linguistic traits. While this approach has been successful in understanding genetic relations between diverse languages, more analyses are needed to accurately quantify their…

Computation and Language · Computer Science 2024-10-04 Juan De Gregorio , Raúl Toral , David Sánchez

This study addresses the gap in the literature concerning the comparative performance of LLMs in interpreting different types of figurative language across multiple languages. By evaluating LLMs using two multilingual datasets on simile and…

Computation and Language · Computer Science 2025-03-11 Paria Khoshtab , Danial Namazifard , Mostafa Masoudi , Ali Akhgary , Samin Mahdizadeh Sani , Yadollah Yaghoobzadeh

We construct a multilingual common semantic space based on distributional semantics, where words from multiple languages are projected into a shared space to enable knowledge and resource transfer across languages. Beyond word alignment, we…

Computation and Language · Computer Science 2018-04-24 Lifu Huang , Kyunghyun Cho , Boliang Zhang , Heng Ji , Kevin Knight

The ability of Large Language Models (LLMs) to encode syntactic and semantic structures of language is well examined in NLP. Additionally, analogy identification, in the form of word analogies are extensively studied in the last decade of…

Computation and Language · Computer Science 2024-02-07 Thilini Wijesiriwardene , Ruwan Wickramarachchi , Aishwarya Naresh Reganti , Vinija Jain , Aman Chadha , Amit Sheth , Amitava Das

Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplored. We conduct a comprehensive study across 96 languages to analyze the underlying structure of…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Kwanghee Choi , Eunjung Yeo , Cheol Jun Cho , David Harwath , David R. Mortensen

Conversational tones -- the manners and attitudes in which speakers communicate -- are essential to effective communication. Amidst the increasing popularization of Large Language Models (LLMs) over recent years, it becomes necessary to…

Computation and Language · Computer Science 2024-06-07 Dun-Ming Huang , Pol Van Rijn , Ilia Sucholutsky , Raja Marjieh , Nori Jacoby
‹ Prev 1 8 9 10 Next ›