English
Related papers

Related papers: KoParadigm: A Korean Conjugation Paradigm Generato…

200 papers

Most of the post-processing methods for character recognition rely on contextual information of character and word-fragment levels. However, due to linguistic characteristics of Korean, such low-level information alone is not sufficient for…

cmp-lg · Computer Science 2008-02-03 Geunbae Lee , Jong-Hyeok Lee , JinHee Yoo

The recent emergence of Large Vision-Language Models(VLMs) has resulted in a variety of different benchmarks for evaluating such models. Despite this, we observe that most existing evaluation methods suffer from the fact that they either…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yoonshik Kim , Jaeyoon Jung

Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification models, creating such datasets presents several challenges: i)…

Computation and Language · Computer Science 2025-06-17 Minkyeong Jeon , Hyemin Jeong , Yerang Kim , Jiyoung Kim , Jae Hyeon Cho , Byung-Jun Lee

Short text classification (STC) remains a challenging task due to the scarcity of contextual information and labeled data. However, existing approaches have pre-dominantly focused on English because most benchmark datasets for the STC are…

Computation and Language · Computer Science 2026-03-05 JaeGeon Yoo , Byoungwook Kim , Yeongwook Yang , Hong-Jun Jang

This paper introduces Thunder-Tok, a new Korean tokenizer designed to reduce token fertility without compromising model performance. Our approach uses a rule-based pre-tokenization method that aligns with the linguistic structure of the…

Computation and Language · Computer Science 2025-06-19 Gyeongje Cho , Yeonkyoun So , Chanwoo Park , Sangmin Lee , Sungmok Jung , Jaejin Lee

Markedness in natural language is often associated with non-literal meanings in discourse. Differential Object Marking (DOM) in Korean is one instance of this phenomenon, where post-positional markers are selected based on both the semantic…

Computation and Language · Computer Science 2024-07-02 Hagyeong Shin , Sean Trott

Perspective differences exist among different cultures or languages. A lack of mutual understanding among different groups about their perspectives on specific values or events may lead to uninformed decisions or biased opinions.…

Computation and Language · Computer Science 2021-04-14 Yufei Tian , Tuhin Chakrabarty , Fred Morstatter , Nanyun Peng

Large language models (LLMs) use pretraining to predict the subsequent word; however, their expansion requires significant computing resources. Numerous big tech companies and research institutes have developed multilingual LLMs (MLLMs) to…

English based datasets are commonly available from Kaggle, GitHub, or recently published papers. Although benchmark tests with English datasets are sufficient to show off the performances of new models and methods, still a researcher need…

Computation and Language · Computer Science 2022-11-29 Byunghyun Ban

The development of practical (multimodal) large language model assistants for Korean weather forecasters is hindered by the absence of a multidimensional, expert-level evaluation framework grounded in authoritative sources. To address this,…

Computation and Language · Computer Science 2026-04-28 Soyeon Kim , Cheongwoong Kang , Myeongjin Lee , Eun-Chul Chang , Jaedeok Lee , Jaesik Choi

Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build smaller yet capable models, most focus on English and little…

For readability and disambiguation of the written text, appropriate word segmentation is recommended for documentation, and it also holds for the digitized texts. If the language is agglutinative while far from scriptio continua, for…

Computation and Language · Computer Science 2021-05-05 Won Ik Cho , Sung Jun Cheon , Woo Hyun Kang , Ji Won Kim , Nam Soo Kim

Large-scale natural language generation requires the integration of vast amounts of knowledge: lexical, grammatical, and conceptual. A robust generator must be able to operate well even when pieces of knowledge are missing. It must also be…

cmp-lg · Computer Science 2008-02-03 Kevin Knight , Vasileios Hatzivassiloglou

We describe a resource-based method of morphological annotation of written Korean text. Korean is an agglutinative language. The output of our system is a graph of morphemes annotated with accurate linguistic information. The language…

Computation and Language · Computer Science 2007-11-22 Hyun-Gue Huh , Eric Laporte

We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that…

Computation and Language · Computer Science 2019-11-01 Kang Min Yoo , Taeuk Kim , Sang-goo Lee

Chinese couplet is a special form of poetry composed of complex syntax with ancient Chinese language. Due to the complexity of semantic and grammatical rules, creation of a suitable couplet is a formidable challenge. This paper presents a…

Computation and Language · Computer Science 2021-12-06 Kuan-Yu Chiang , Shihao Lin , Joe Chen , Qian Yin , Qizhen Jin

Physical commonsense reasoning datasets like PIQA are predominantly English-centric and lack cultural diversity. We introduce Ko-PIQA, a Korean physical commonsense reasoning dataset that incorporates cultural context. Starting from 3.01…

Computation and Language · Computer Science 2025-09-30 Dasol Choi , Jungwhan Kim , Guijin Son

This paper attempts to analyze the Korean sentence classification system for a chatbot. Sentence classification is the task of classifying an input sentence based on predefined categories. However, spelling or space error contained in the…

Computation and Language · Computer Science 2021-06-08 DongHyun Choi , IlNam Park , Myeong Cheol Shin , EungGyun Kim , Dong Ryeol Shin

E-learning systems should deliver contents that reflect various phenomena of the language as it is used. In addition to formal Korean, e-learning systems that would include real-world Korean expressions such as those in web documents,…

Computation and Language · Computer Science 2026-05-29 Sang-Taek Park , Ae-Lim Ahn , Eric Laporte , Jee-Sun Nam

The sense analysis is still critical problem in machine translation system, especially such as English-Korean translation which the syntactical different between source and target languages is very great. We suggest a method for selecting…

Computation and Language · Computer Science 2013-05-16 Hyonil Kim , Changil Choe
‹ Prev 1 3 4 5 6 7 10 Next ›