中文
相关论文

相关论文: DriveThru: a Document Extraction Platform and Benc…

200 篇论文

This study provides an overview of the history of the development of Natural Language Processing (NLP) in the context of the Indonesian language, with a focus on the basic technologies, methods, and practical applications that have been…

计算与语言 · 计算机科学 2023-04-07 Mukhlis Amien

Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multilingual datasets are derived from English translations,…

计算与语言 · 计算机科学 2025-11-13 Vanessa Rebecca Wiyono , David Anugraha , Ayu Purwarianti , Genta Indra Winata

Machine Reading Comprehension (MRC) has become one of the essential tasks in Natural Language Understanding (NLU) as it is often included in several NLU benchmarks (Liang et al., 2020; Wilie et al., 2020). However, most MRC datasets only…

计算与语言 · 计算机科学 2022-10-26 Rifki Afina Putri , Alice Oh

Previous work in Indonesian part-of-speech (POS) tagging are hard to compare as they are not evaluated on a common dataset. Furthermore, in spite of the success of neural network models for English POS tagging, they are rarely explored for…

计算与语言 · 计算机科学 2019-02-27 Kemal Kurniawan , Alham Fikri Aji

Despite the existence of numerous Optical Character Recognition (OCR) tools, the lack of comprehensive open-source systems hampers the progress of document digitization in various low-resource languages, including Bengali. Low-resource…

In this paper, we introduce a large-scale Indonesian summarization dataset. We harvest articles from Liputan6.com, an online news portal, and obtain 215,827 document-summary pairs. We leverage pre-trained language models to develop…

计算与语言 · 计算机科学 2020-11-03 Fajri Koto , Jey Han Lau , Timothy Baldwin

Compared to English, the amount of labeled data for Indonesian text classification tasks is very small. Recently developed multilingual language models have shown its ability to create multilingual representations effectively. This paper…

计算与语言 · 计算机科学 2020-09-15 Ilham Firdausi Putra , Ayu Purwarianti

As one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in terms of NLP progress. We introduce LoraxBench, a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks:…

计算与语言 · 计算机科学 2025-08-19 Alham Fikri Aji , Trevor Cohn

Hate speech poses a significant threat to social harmony. Over the past two years, Indonesia has seen a ten-fold increase in the online hate speech ratio, underscoring the urgent need for effective detection mechanisms. However, progress is…

Recently there have been intensifying efforts to improve the understanding of Indonesian cultures by large language models (LLMs). An attractive source of cultural knowledge that has been largely overlooked is local journals of social…

计算与语言 · 计算机科学 2026-01-21 Adimulya Kartiyasa , Bao Gia Cao , Boyang Li

Although region-specific large language models (LLMs) are increasingly developed, their safety remains underexplored, particularly in culturally diverse settings like Indonesia, where sensitivity to local norms is essential and highly…

计算与语言 · 计算机科学 2025-06-04 Muhammad Falensi Azmi , Muhammad Dehan Al Kautsar , Alfan Farizki Wicaksono , Fajri Koto

This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as English, Hindi, and…

计算与语言 · 计算机科学 2025-05-19 Hrishit Madhavi , Jacob Cherian , Yuvraj Khamkar , Dhananjay Bhagat

An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise…

计算与语言 · 计算机科学 2024-10-15 Aulia Adila , Dessi Lestari , Ayu Purwarianti , Dipta Tanaya , Kurniawati Azizah , Sakriani Sakti

Optical Character Recognition (OCR) technology finds applications in digitizing books and unstructured documents, along with applications in other domains such as mobility statistics, law enforcement, traffic, security systems, etc. The…

计算机视觉与模式识别 · 计算机科学 2023-07-11 Aishik Rakshit , Samyak Mehta , Anirban Dasgupta

Making use of off-the-shelf resources of resource-rich languages to transfer knowledge for low-resource languages raises much attention recently. The requirements of enabling the model to reach the reliable performance lack well guided,…

计算与语言 · 计算机科学 2024-10-25 Donglin Di , Weinan Zhang , Yue Zhang , Fanglin Wang

Significant progress has been made on Indonesian NLP. Nevertheless, exploration of the code-mixing phenomenon in Indonesian is limited, despite many languages being frequently mixed with Indonesian in daily conversation. In this work, we…

计算与语言 · 计算机科学 2023-11-22 Muhammad Farid Adilazuarda , Samuel Cahyawijaya , Genta Indra Winata , Pascale Fung , Ayu Purwarianti

Document extraction is a core component of digital workflows, yet existing vision-language models (VLMs) predominantly favor high-resource languages. Thai presents additional challenges due to script complexity from non-latin letters, the…

Although large language models (LLMs) are often pre-trained on large-scale multilingual texts, their reasoning abilities and real-world knowledge are mainly evaluated based on English datasets. Assessing LLM capabilities beyond English is…

计算与语言 · 计算机科学 2023-10-24 Fajri Koto , Nurul Aisyah , Haonan Li , Timothy Baldwin

We present a novel approach to data preparation for developing multilingual Indic large language model. Our meticulous data acquisition spans open-source and proprietary sources, including Common Crawl, Indic books, news articles, and…

Lexical-semantic resources (LSRs), such as online lexicons and wordnets, are fundamental to natural language processing applications as well as to fields such as linguistic anthropology and language preservation. In many languages, however,…

计算与语言 · 计算机科学 2025-11-21 Hadi Khalilia , Jahna Otterbacher , Gabor Bella , Shandy Darma , Fausto Giunchiglia