English
Related papers

Related papers: InDEX: Indonesian Idiom and Expression Dataset for…

200 papers

We introduce XED, a multilingual fine-grained emotion dataset. The dataset consists of human-annotated Finnish (25k) and English sentences (30k), as well as projected annotations for 30 additional languages, providing new resources for many…

Computation and Language · Computer Science 2020-11-09 Emily Öhman , Marc Pàmies , Kaisla Kajava , Jörg Tiedemann

In this paper, MapReduce programming model is used to parallelize training and tagging proceess in Maximum Entropy part of speech tagging for Bahasa Indonesia. In training process, MapReduce model is implemented dictionary, tagtoken, and…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-08-16 Arif Nurwidyantoro , Edi Winarko

While contextualized word embeddings have been a de-facto standard, learning contextualized phrase embeddings is less explored and being hindered by the lack of a human-annotated benchmark that tests machine understanding of phrase…

Computation and Language · Computer Science 2023-02-03 Thang M. Pham , Seunghyun Yoon , Trung Bui , Anh Nguyen

Studies of morphological processing have shown that semantic transparency is crucial for word recognition. Its computational operationalization is still under discussion. Our primary objectives are to explore embedding-based measures of…

Computation and Language · Computer Science 2025-05-12 M. Maziyah Mohamed , R. H. Baayen

Crossword puzzles are popular linguistic games often used as tools to engage students in learning. Educational crosswords are characterized by less cryptic and more factual clues that distinguish them from traditional crossword puzzles.…

Computation and Language · Computer Science 2024-04-10 Andrea Zugarini , Kamyar Zeinalipour , Surya Sai Kadali , Marco Maggini , Marco Gori , Leonardo Rigutini

Chinese input methods are used to convert pinyin sequence or other Latin encoding systems into Chinese character sentences. For more effective pinyin-to-character conversion, typical Input Method Engines (IMEs) rely on a predefined…

Computation and Language · Computer Science 2017-12-13 Xihu Zhang , Chu Wei , Hai Zhao

Although region-specific large language models (LLMs) are increasingly developed, their safety remains underexplored, particularly in culturally diverse settings like Indonesia, where sensitivity to local norms is essential and highly…

Computation and Language · Computer Science 2025-06-04 Muhammad Falensi Azmi , Muhammad Dehan Al Kautsar , Alfan Farizki Wicaksono , Fajri Koto

Indonesian is an agglutinative language since it has a compounding process of word-formation. Therefore, the translation model of this language requires a mechanism that is even lower than the word level, referred to as the sub-word level.…

Computation and Language · Computer Science 2022-07-04 Mukhlis Amien , Feng Chong , Huang Heyan

The exponential growth of e-commerce platforms in Indonesia has generated a massive volume of user-generated product reviews. Analyzing the sentiment of these reviews is critical for measuring customer satisfaction and identifying product…

Wordnets are indispensable tools for various natural language processing applications. Unfortunately, wordnets get outdated, and producing or updating wordnets can be slow and costly in terms of time and resources. This problem intensifies…

Domain-specific dialogue systems generally determine user intents by relying on sentence level classifiers that mainly focus on single action sentences. Such classifiers are not designed to effectively handle complex queries composed of…

Computation and Language · Computer Science 2022-11-24 Aadesh Gupta , Kaustubh D. Dhole , Rahul Tarway , Swetha Prabhakar , Ashish Shrivastava

The article observes data analysis of 286 multi-word expressions (MWEs) based on 16 lexical, grammatical and other criteria described in theoretical books and papers on the notion of idiomaticity. MWEs were collected from the same…

Computation and Language · Computer Science 2026-05-20 Elena Mikhalkova , Anastasiya Vishnyakova , Anastasiya Drozdova , Polina Gavin , Aleksander Zhmykhov , Timofey Protasov

Named entity recognition (NER) is a fundamental task of natural language processing (NLP). However, most state-of-the-art research is mainly oriented to high-resource languages such as English and has not been widely applied to low-resource…

Computation and Language · Computer Science 2021-09-06 Yingwen Fu , Nankai Lin , Zhihe Yang , Shengyi Jiang

Most available semantic parsing datasets, comprising of pairs of natural utterances and logical forms, were collected solely for the purpose of training and evaluation of natural language understanding systems. As a result, they do not…

Computation and Language · Computer Science 2021-06-10 Moshe Hazoom , Vibhor Malik , Ben Bogin

Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, the first large-scale…

Computation and Language · Computer Science 2024-06-25 Alkis Koudounas , Moreno La Quatra , Lorenzo Vaiani , Luca Colomba , Giuseppe Attanasio , Eliana Pastor , Luca Cagliero , Elena Baralis

In Linguistics, a grapheme is a written unit of a writing system corresponding to a phonological sound. In Natural Language Processing tasks, written language is analysed through two different mediums, word analysis, and character analysis.…

Computation and Language · Computer Science 2024-04-03 Samuel Rose , Chandrasekhar Kambhampati

As a kind of new expression elements, Internet memes are popular and extensively used in online chatting scenarios since they manage to make dialogues vivid, moving, and interesting. However, most current dialogue researches focus on…

Computation and Language · Computer Science 2021-09-07 Zhengcong Fei , Zekang Li , Jinchao Zhang , Yang Feng , Jie Zhou

We present ClidSum, a benchmark dataset for building cross-lingual summarization systems on dialogue documents. It consists of 67k+ dialogue documents from two subsets (i.e., SAMSum and MediaSum) and 112k+ annotated summaries in different…

Computation and Language · Computer Science 2022-10-18 Jiaan Wang , Fandong Meng , Ziyao Lu , Duo Zheng , Zhixu Li , Jianfeng Qu , Jie Zhou

Learning high-quality dialogue representations is essential for solving a variety of dialogue-oriented tasks, especially considering that dialogue systems often suffer from data scarcity. In this paper, we introduce Dialogue Sentence…

Computation and Language · Computer Science 2022-07-25 Zhihan Zhou , Dejiao Zhang , Wei Xiao , Nicholas Dingwall , Xiaofei Ma , Andrew O. Arnold , Bing Xiang

We present the Siamese Continuous Bag of Words (Siamese CBOW) model, a neural network for efficient estimation of high-quality sentence embeddings. Averaging the embeddings of words in a sentence has proven to be a surprisingly successful…

Computation and Language · Computer Science 2016-06-16 Tom Kenter , Alexey Borisov , Maarten de Rijke
‹ Prev 1 3 4 5 6 7 10 Next ›