中文
相关论文

相关论文: Adaptive Engram Memory System for Indonesian Langu…

200 篇论文

This paper presents a novel syllable-based tokenization approach for Indonesian large language models, inspired by the Gasing Literacy Learning System's pedagogical methodology. Drawing on information-theoretic principles, we develop a…

计算机与社会 · 计算机科学 2026-01-21 H. Situngkir , A. B. Lumbantobing , Y. Surya

Tokenization constitutes a fundamental stage in Large Language Model (LLM) processing; however, subword-based tokenization methods optimized on English-dominant corpora may produce token fragmentation misaligned with the linguistic…

计算机与社会 · 计算机科学 2026-02-10 Andhika Bernard Lumbantobing , Hokky Situngkir

Text-only adaptation of an end-to-end (E2E) model remains a challenging task for automatic speech recognition (ASR). Language model (LM) fusion-based approaches require an additional external LM during inference, significantly increasing…

计算与语言 · 计算机科学 2022-11-01 Zhong Meng , Yashesh Gaur , Naoyuki Kanda , Jinyu Li , Xie Chen , Yu Wu , Yifan Gong

Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-resource languages. However, their performance in low-resource languages is limited by the scarcity of both parallel and monolingual corpora,…

计算与语言 · 计算机科学 2024-10-11 William Tan , Kevin Zhu

In this work, we present a comprehensive exploration of finetuning Malaysian language models, specifically Llama2 and Mistral, on embedding tasks involving negative and positive pairs. We release two distinct models tailored for Semantic…

计算与语言 · 计算机科学 2024-02-06 Husein Zolkepli , Aisyah Razak , Kamarul Adha , Ariff Nazhan

Indonesian language is spoken by almost 200 million people and is the 10th most spoken language in the world, but it is under-represented in NLP (Natural Language Processing) research. A sparsity of language resources has hampered previous…

计算与语言 · 计算机科学 2024-10-28 Mukhlish Fuadi , Adhi Dharma Wibawa , Surya Sumpeno

We present an efficient method for adapting a monolingual Large Language Model (LLM) to another language, addressing challenges of catastrophic forgetting and tokenizer limitations. We focus this study on adapting Llama 2 to Arabic. Our…

Transformer architectures based on the attention mechanism have revolutionized natural language processing (NLP), driving major breakthroughs across virtually every NLP task. However, their substantial memory and computational requirements…

计算与语言 · 计算机科学 2026-03-25 Riccardo Bravin , Massimo Pavan , Hazem Hesham Yousef Shalby , Fabrizio Pittorino , Manuel Roveri

This paper presents a systematic benchmark of state-of-the-art multilingual large language models (LLMs) adapted via token pruning - a compression technique that eliminates tokens and embedding parameters corresponding to languages…

计算与语言 · 计算机科学 2026-04-20 Hoyeol Kim , Hyeonwoo Kim

Neural machine translation (NMT) for low-resource local languages in Indonesia faces significant challenges, including the need for a representative benchmark and limited data availability. This work addresses these challenges by…

计算与语言 · 计算机科学 2023-11-03 Lucky Susanto , Ryandito Diandaru , Adila Krisnadhi , Ayu Purwarianti , Derry Wijaya

Large language models (LLMs) deployed in user-facing applications require long-horizon consistency: the ability to remember prior interactions, respect user preferences, and ground reasoning in past events. However, contemporary memory…

多智能体系统 · 计算机科学 2026-02-04 Daivik Patel , Shrenik Patel

Large language models (LLMs) show remarkable human-like capability in various domains and languages. However, a notable quality gap arises in low-resource languages, e.g., Indonesian indigenous languages, rendering them ineffective and…

Sentence embeddings are a foundational component for semantic search, clustering, classification, and retrieval-augmented generation. This paper presents embeddingmagibu-200m, a Turkish-focused sentence embedding model that produces…

计算与语言 · 计算机科学 2026-05-29 M. Ali Bayram , Banu Diri , Savaş Yıldırım

Indonesian is an agglutinative language since it has a compounding process of word-formation. Therefore, the translation model of this language requires a mechanism that is even lower than the word level, referred to as the sub-word level.…

计算与语言 · 计算机科学 2022-07-04 Mukhlis Amien , Feng Chong , Huang Heyan

Addressing the gap in Large Language Model pretrained from scratch with Malaysian context, We trained models with 1.1 billion, 3 billion, and 5 billion parameters on a substantial 349GB dataset, equivalent to 90 billion tokens based on our…

计算与语言 · 计算机科学 2024-01-30 Husein Zolkepli , Aisyah Razak , Kamarul Adha , Ariff Nazhan

Although some linguists (Rusmali et al., 1985; Crouch, 2009) have fairly attempted to define the morphology and syntax of Minangkabau, information processing in this language is still absent due to the scarcity of the annotated resource. In…

计算与语言 · 计算机科学 2020-09-22 Fajri Koto , Ikhwan Koto

Induction head mechanism is a part of the computational circuits for in-context learning (ICL) that enable large language models (LLMs) to adapt to new tasks without fine-tuning. Most existing work explains the training dynamics behind…

计算与语言 · 计算机科学 2025-07-09 Shuo Wang , Issei Sato

Time series modeling holds significant importance in many real-world applications and has been extensively studied. While pre-trained foundation models have made impressive strides in the fields of natural language processing (NLP) and…

计算与语言 · 计算机科学 2025-02-20 Juyuan Zhang , Wei Zhu , Jiechao Gao

The recent breakthroughs in Large Language Models (LLMs) have mostly focused on languages with easily available and sufficient resources, such as English. However, there remains a significant gap for languages that lack sufficient…

计算与语言 · 计算机科学 2024-03-20 Louis Owen , Vishesh Tripathi , Abhay Kumar , Biddwan Ahmed
‹ 上一页 1 2 3 10 下一页 ›