中文
相关论文

相关论文: Subword-Based Comparative Linguistics across 242 L…

200 篇论文

Contextualized embeddings based on large language models (LLMs) are available for various languages, but their coverage is often limited for lower resourced languages. Using LLMs for such languages is often difficult due to a high…

计算与语言 · 计算机科学 2025-01-28 Daniil Gurgurov , Rishu Kumar , Simon Ostermann

We generalize the word analogy task across languages, to provide a new intrinsic evaluation method for cross-lingual semantic spaces. We experiment with six languages within different language families, including English, German, Spanish,…

计算与语言 · 计算机科学 2018-07-12 Tomáš Brychcín , Stephen Eugene Taylor , Lukáš Svoboda

Very low-resource languages, having only a few million tokens worth of data, are not well-supported by multilingual NLP approaches due to poor quality cross-lingual word representations. Recent work showed that good cross-lingual…

计算与语言 · 计算机科学 2023-11-22 Viktor Hangya , Silvia Severini , Radoslav Ralev , Alexander Fraser , Hinrich Schütze

The use of subword embedding has proved to be a major innovation in Neural Machine Translation (NMT). It helps NMT to learn better context vectors for Low Resource Languages (LRLs) so as to predict the target words by better modelling the…

计算与语言 · 计算机科学 2023-05-23 Amit Kumar , Shantipriya Parida , Ajay Pratap , Anil Kumar Singh

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

计算与语言 · 计算机科学 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer

Handwritten numerals of different languages have various characteristics. Similarities and dissimilarities of the languages can be measured by analyzing the extracted features of the numerals. Handwritten numeral datasets are available and…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Md. Rahat-uz-Zaman , Shadmaan Hye

We present a language independent, unsupervised method for building word embeddings using morphological expansion of text. Our model handles the problem of data sparsity and yields improved word embeddings by relying on training word…

计算与语言 · 计算机科学 2017-11-16 Syed Sarfaraz Akhtar , Arihant Gupta , Avijit Vajpayee , Arjit Srivastava , Manish Shrivastava

Words and phrases acquire meaning from the way they are used in society, from their relative semantics to other words and phrases. For computers the equivalent of `society' is `database,' and the equivalent of `use' is `way to search the…

计算与语言 · 计算机科学 2007-06-13 Rudi Cilibrasi , Paul M. B. Vitanyi

Traditionally linguists have organized languages of the world as language families modelled as trees. In this work we take a contrarian approach and question the tree-based model that is rather restrictive. For example, the affinity that…

计算与语言 · 计算机科学 2020-12-15 Priya S. Nayak , Rhythm Girdhar , Shreekanth M. Prabhu

Cross-Lingual Word Embeddings (CLWEs) are a key component to transfer linguistic information learnt from higher-resource settings into lower-resource ones. Recent research in cross-lingual representation learning has focused on offline…

计算与语言 · 计算机科学 2022-03-29 Sonal Sannigrahi , Jesse Read

We test similarity-based word alignment models (SimAlign and awesome-align) in combination with word embeddings from mBERT and XLM-R on parallel sentences in German and Romansh. Since Romansh is an unseen language, we are dealing with a…

计算与语言 · 计算机科学 2024-05-01 Eyal Liron Dolev

Traditional greedy tokenization methods have been a critical step in Natural Language Processing (NLP), influencing how text is converted into tokens and directly impacting model performance. While subword tokenizers like Byte-Pair Encoding…

计算与语言 · 计算机科学 2025-05-05 Bharath Raj , Garvit Suri , Vikrant Dewangan , Raghav Sonavane

Calculating the semantic similarity between sentences is a long dealt problem in the area of natural language processing. The semantic analysis field has a crucial role to play in the research related to the text analytics. The semantic…

计算与语言 · 计算机科学 2018-02-22 Atish Pawar , Vijay Mago

The rise of digital misinformation has heightened interest in using multilingual Large Language Models (LLMs) for fact-checking. This study systematically evaluates translation bias and the effectiveness of LLMs for cross-lingual claim…

计算与语言 · 计算机科学 2024-10-15 Aryan Singhal , Veronica Shao , Gary Sun , Ryan Ding , Jonathan Lu , Kevin Zhu

Contextual word-representations became a standard in modern natural language processing systems. These models use subword tokenization to handle large vocabularies and unknown words. Word-level usage of such systems requires a way of…

计算与语言 · 计算机科学 2021-03-30 Judit Ács , Ákos Kádár , András Kornai

Lexical-semantic resources (LSRs), such as online lexicons and wordnets, are fundamental to natural language processing applications as well as to fields such as linguistic anthropology and language preservation. In many languages, however,…

计算与语言 · 计算机科学 2025-11-21 Hadi Khalilia , Jahna Otterbacher , Gabor Bella , Shandy Darma , Fausto Giunchiglia

Emerging digital technologies are exacerbating the existing divide in Open Access Data (OAD) between high-and low-resource languages, excluding many communities from participating in the global digital transformation. In this PhD proposal,…

人工智能 · 计算机科学 2026-05-08 Ndeye-Emilie Mbengue

Bilingual lexicons and phrase tables are critical resources for modern Machine Translation systems. Although recent results show that without any seed lexicon or parallel data, highly accurate bilingual lexicons can be learned using…

计算与语言 · 计算机科学 2020-09-01 Ashiqur R. KhudaBukhsh , Shriphani Palakodety , Tom M. Mitchell

Pre-tokenization, the initial step in many modern tokenization pipelines, segments text into smaller units called pretokens, typically splitting on whitespace and punctuation. While this process encourages having full, individual words as…

计算与语言 · 计算机科学 2025-10-03 Craig W. Schmidt , Varshini Reddy , Chris Tanner , Yuval Pinter

Neural Machine Translation (NMT) in low-resource settings and of morphologically rich languages is made difficult in part by data sparsity of vocabulary words. Several methods have been used to help reduce this sparsity, notably Byte-Pair…

计算与语言 · 计算机科学 2018-09-11 Pamela Shapiro , Kevin Duh