中文
相关论文

相关论文: InDEX: Indonesian Idiom and Expression Dataset for…

200 篇论文

Idioms are special fixed phrases usually derived from stories. They are commonly used in casual conversations and literary writings. Their meanings are usually highly non-compositional. The idiom cloze task is a challenge problem in Natural…

计算与语言 · 计算机科学 2021-12-07 Ruiyang Qin , Haozheng Luo , Zheheng Fan , Ziang Ren

Although the Indonesian language is spoken by almost 200 million people and the 10th most spoken language in the world, it is under-represented in NLP research. Previous work on Indonesian has been hampered by a lack of annotated datasets,…

计算与语言 · 计算机科学 2020-11-03 Fajri Koto , Afshin Rahimi , Jey Han Lau , Timothy Baldwin

Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multilingual datasets are derived from English translations,…

计算与语言 · 计算机科学 2025-11-13 Vanessa Rebecca Wiyono , David Anugraha , Ayu Purwarianti , Genta Indra Winata

We present IndoNLI, the first human-elicited NLI dataset for Indonesian. We adapt the data collection protocol for MNLI and collect nearly 18K sentence pairs annotated by crowd workers and experts. The expert-annotated data is used…

计算与语言 · 计算机科学 2022-03-30 Rahmad Mahendra , Alham Fikri Aji , Samuel Louvan , Fahrurrozi Rahman , Clara Vania

Cloze-style reading comprehension in Chinese is still limited due to the lack of various corpora. In this paper we propose a large-scale Chinese cloze test dataset ChID, which studies the comprehension of idiom, a unique language phenomenon…

计算与语言 · 计算机科学 2020-01-28 Chujie Zheng , Minlie Huang , Aixin Sun

We present IndoBERTweet, the first large-scale pretrained model for Indonesian Twitter that is trained by extending a monolingually-trained Indonesian BERT model with additive domain-specific vocabulary. We focus in particular on efficient…

计算与语言 · 计算机科学 2021-09-13 Fajri Koto , Jey Han Lau , Timothy Baldwin

Existing Indonesian sentiment analysis models classify text in isolation, ignoring the topical context that often determines whether a statement is positive, negative, or neutral. We introduce IndoBERT-Sentiment, a context-conditioned…

In this paper, we introduce a large-scale Indonesian summarization dataset. We harvest articles from Liputan6.com, an online news portal, and obtain 215,827 document-summary pairs. We leverage pre-trained language models to develop…

计算与语言 · 计算机科学 2020-11-03 Fajri Koto , Jey Han Lau , Timothy Baldwin

We introduced KaWAT (Kata Word Analogy Task), a new word analogy task dataset for Indonesian. We evaluated on it several existing pretrained Indonesian word embeddings and embeddings trained on Indonesian online news corpus. We also tested…

计算与语言 · 计算机科学 2019-06-25 Kemal Kurniawan

Although Indonesian is known to be the fourth most frequently used language over the internet, the research progress on this language in the natural language processing (NLP) is slow-moving due to a lack of available resources. In response,…

Determining whether a piece of text is relevant to a given topic is a fundamental task in natural language processing, yet it remains largely unexplored for Bahasa Indonesia. Unlike sentiment analysis or named entity recognition, relevancy…

Idiomatic expressions have always been a bottleneck for language comprehension and natural language understanding, specifically for tasks like Machine Translation(MT). MT systems predominantly produce literal translations of idiomatic…

计算与语言 · 计算机科学 2020-06-18 Prateek Saxena , Soma Paul

Indonesian marketplace reviews mix standard vocabulary with slang, regional loanwords, numeric shorthands, and emoji, making lexicon-based sentiment tools unreliable in practice. This paper describes a two-track classification pipeline…

计算与语言 · 计算机科学 2026-04-28 Hermawan Manurung , Ibrahim Al-Kahfi , Ahmad Rizqi , Martin Clinton Tosima Manullang

Researches on Indonesian named entity (NE) tagger have been conducted since years ago. However, most did not use deep learning and instead employed traditional machine learning algorithms such as association rule, support vector machine,…

计算与语言 · 计算机科学 2020-09-15 Devin Hoesen , Ayu Purwarianti

Idiomatic expressions can be problematic for natural language processing applications as their meaning cannot be inferred from their constituting words. A lack of successful methodological approaches and sufficiently large datasets prevents…

计算与语言 · 计算机科学 2021-11-11 Tadej Škvorc , Polona Gantar , Marko Robnik-Šikonja

Automatic text summarization is generally considered as a challenging task in the NLP community. One of the challenges is the publicly available and large dataset that is relatively rare and difficult to construct. The problem is even worse…

计算与语言 · 计算机科学 2019-03-21 Kemal Kurniawan , Samuel Louvan

We introduce SCDE, a dataset to evaluate the performance of computational models through sentence prediction. SCDE is a human-created sentence cloze dataset, collected from public school English examinations. Our task requires a model to…

计算与语言 · 计算机科学 2020-04-28 Xiang Kong , Varun Gangal , Eduard Hovy

Chinese idioms are special fixed phrases usually derived from ancient stories, whose meanings are oftentimes highly idiomatic and non-compositional. The Chinese idiom prediction task is to select the correct idiom from a set of candidate…

计算与语言 · 计算机科学 2020-11-05 Minghuan Tan , Jing Jiang

Understanding emotions in the Indonesian language is essential for improving customer experiences in e-commerce. This study focuses on enhancing the accuracy of emotion classification in Indonesian by leveraging advanced language models,…

计算与语言 · 计算机科学 2025-09-19 William Christian , Daniel Adamlu , Adrian Yu , Derwin Suhartono

Evaluating the in-context learning classification performance of language models poses challenges due to small dataset sizes, extensive prompt-selection using the validation set, and intentionally difficult tasks that lead to near-random…

计算与语言 · 计算机科学 2024-11-12 Gregory Yauney , David Mimno
‹ 上一页 1 2 3 10 下一页 ›