English
Related papers

Related papers: Sailor: Open Language Models for South-East Asia

200 papers

Sailor2 is a family of cutting-edge multilingual language models for South-East Asian (SEA) languages, available in 1B, 8B, and 20B sizes to suit diverse applications. Building on Qwen2.5, Sailor2 undergoes continuous pre-training on 500B…

Recently, Large Language Models (LLMs) have dominated much of the artificial intelligence scene with their ability to process and generate natural languages. However, the majority of LLM research and development remains English-centric,…

We introduce Xmodel-1.5, a 1-billion-parameter multilingual large language model pretrained on 2 trillion tokens, designed for balanced performance and scalability. Unlike most large models that use the BPE tokenizer, Xmodel-1.5 employs a…

Computation and Language · Computer Science 2024-12-05 Wang Qun , Liu Yang , Lin Qingquan , Jiang Ling

Despite the remarkable achievements of large language models (LLMs) in various tasks, there remains a linguistic bias that favors high-resource languages, such as English, often at the expense of low-resource and regional languages. To…

Large Language Models (LLMs) have shown remarkable abilities across various tasks, yet their development has predominantly centered on high-resource languages like English and Chinese, leaving low-resource languages underserved. To address…

Computation and Language · Computer Science 2024-07-30 Wenxuan Zhang , Hou Pong Chan , Yiran Zhao , Mahani Aljunied , Jianyu Wang , Chaoqun Liu , Yue Deng , Zhiqiang Hu , Weiwen Xu , Yew Ken Chia , Xin Li , Lidong Bing

In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass encompasses three main SEA languages, eight primary tasks…

Computation and Language · Computer Science 2024-12-03 Jia Guo , Longxu Dou , Guangtao Zeng , Stanley Kok , Wei Lu , Qian Liu

In this study, we introduce Orion-14B, a collection of multilingual large language models with 14 billion parameters. We utilize a data scheduling approach to train a foundational model on a diverse corpus of 2.5 trillion tokens, sourced…

Computation and Language · Computer Science 2024-01-24 Du Chen , Yi Huang , Xiaopu Li , Yongqiang Li , Yongqiang Liu , Haihui Pan , Leichao Xu , Dacheng Zhang , Zhipeng Zhang , Kun Han

We introduce SeaLLMs-Audio, the first large audio-language model (LALM) tailored for multiple Southeast Asian (SEA) languages-Indonesian (id), Thai (th), and Vietnamese (vi)-alongside English (en) and Chinese (zh). Trained on a large-scale…

Computation and Language · Computer Science 2025-11-04 Chaoqun Liu , Mahani Aljunied , Guizhen Chen , Hou Pong Chan , Weiwen Xu , Yu Rong , Wenxuan Zhang

This technical report introduces JAI-1, a Thai-centric language model with 75B parameters. Recent Thai models have primarily relied on existing open-source models, applying additional training without structural modifications to specialize…

With the rapid emergence of novel capabilities in Large Language Models (LLMs), the need for rigorous multilingual and multicultural benchmarks that are integrated has become more pronounced. Though existing LLM benchmarks are capable of…

Large language models have exhibited significant proficiency in languages endowed with extensive linguistic resources, such as English and Chinese. Nevertheless, their effectiveness notably diminishes when applied to languages characterized…

Computation and Language · Computer Science 2024-04-16 Sophia Maria

We present Gamayun, a 1.5B-parameter multilingual language model trained entirely from scratch on 2.5T tokens. Designed for efficiency and deployment in resource-constrained environments, Gamayun addresses the lack of research on small…

Typhoon is a series of Thai large language models (LLMs) developed specifically for the Thai language. This technical report presents challenges and insights in developing Thai LLMs, including data preparation, pretraining,…

We introduce OpenJAI-v1.0, an open-source large language model for Thai and English, developed from the Qwen3-14B model. Our work focuses on boosting performance on practical tasks through carefully curated data across three key use cases:…

Computation and Language · Computer Science 2025-10-09 Pontakorn Trakuekul , Attapol T. Rutherford , Jullajak Karnjanaekarin , Narongkorn Panitsrisit , Sumana Sumanakul

This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietary and open-source…

Computation and Language · Computer Science 2025-03-27 Yangyang Meng , Jinpeng Li , Guodong Lin , Yu Pu , Guanbo Wang , Hu Du , Zhiming Shao , Yukai Huang , Ke Li , Wei-Qiang Zhang

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based…

This paper introduces Typhoon 2, a series of text and multimodal large language models optimized for the Thai language. The series includes models for text, vision, and audio. Typhoon2-Text builds on state-of-the-art open models, such as…

Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing the need for extensive feature engineering. However, most…

We introduce Solar Open, a 102B-parameter bilingual Mixture-of-Experts language model for underserved languages. Solar Open demonstrates a systematic methodology for building competitive LLMs by addressing three interconnected challenges.…

OpenThaiGPT 1.5 is an advanced Thai language chat model based on Qwen v2.5, finetuned on over 2,000,000 Thai instruction pairs. This report provides an engineering perspective on the model's development, capabilities, and performance. We…

Computation and Language · Computer Science 2025-02-26 Sumeth Yuenyong , Kobkrit Viriyayudhakorn , Apivadee Piyatumrong , Jillaphat Jaroenkantasima
‹ Prev 1 2 3 10 Next ›