中文
相关论文

相关论文: KoParadigm: A Korean Conjugation Paradigm Generato…

200 篇论文

Most of the post-processing methods for character recognition rely on contextual information of character and word-fragment levels. However, due to linguistic characteristics of Korean, such low-level information alone is not sufficient for…

cmp-lg · 计算机科学 2008-02-03 Geunbae Lee , Jong-Hyeok Lee , JinHee Yoo

The recent emergence of Large Vision-Language Models(VLMs) has resulted in a variety of different benchmarks for evaluating such models. Despite this, we observe that most existing evaluation methods suffer from the fact that they either…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yoonshik Kim , Jaeyoon Jung

Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification models, creating such datasets presents several challenges: i)…

计算与语言 · 计算机科学 2025-06-17 Minkyeong Jeon , Hyemin Jeong , Yerang Kim , Jiyoung Kim , Jae Hyeon Cho , Byung-Jun Lee

Short text classification (STC) remains a challenging task due to the scarcity of contextual information and labeled data. However, existing approaches have pre-dominantly focused on English because most benchmark datasets for the STC are…

计算与语言 · 计算机科学 2026-03-05 JaeGeon Yoo , Byoungwook Kim , Yeongwook Yang , Hong-Jun Jang

This paper introduces Thunder-Tok, a new Korean tokenizer designed to reduce token fertility without compromising model performance. Our approach uses a rule-based pre-tokenization method that aligns with the linguistic structure of the…

计算与语言 · 计算机科学 2025-06-19 Gyeongje Cho , Yeonkyoun So , Chanwoo Park , Sangmin Lee , Sungmok Jung , Jaejin Lee

Markedness in natural language is often associated with non-literal meanings in discourse. Differential Object Marking (DOM) in Korean is one instance of this phenomenon, where post-positional markers are selected based on both the semantic…

计算与语言 · 计算机科学 2024-07-02 Hagyeong Shin , Sean Trott

Perspective differences exist among different cultures or languages. A lack of mutual understanding among different groups about their perspectives on specific values or events may lead to uninformed decisions or biased opinions.…

计算与语言 · 计算机科学 2021-04-14 Yufei Tian , Tuhin Chakrabarty , Fred Morstatter , Nanyun Peng

Large language models (LLMs) use pretraining to predict the subsequent word; however, their expansion requires significant computing resources. Numerous big tech companies and research institutes have developed multilingual LLMs (MLLMs) to…

English based datasets are commonly available from Kaggle, GitHub, or recently published papers. Although benchmark tests with English datasets are sufficient to show off the performances of new models and methods, still a researcher need…

计算与语言 · 计算机科学 2022-11-29 Byunghyun Ban

The development of practical (multimodal) large language model assistants for Korean weather forecasters is hindered by the absence of a multidimensional, expert-level evaluation framework grounded in authoritative sources. To address this,…

计算与语言 · 计算机科学 2026-04-28 Soyeon Kim , Cheongwoong Kang , Myeongjin Lee , Eun-Chul Chang , Jaedeok Lee , Jaesik Choi

Recent frontier models employ long chain-of-thought reasoning to explore solution spaces in context and achieve stonger performance. While many works study distillation to build smaller yet capable models, most focus on English and little…

For readability and disambiguation of the written text, appropriate word segmentation is recommended for documentation, and it also holds for the digitized texts. If the language is agglutinative while far from scriptio continua, for…

计算与语言 · 计算机科学 2021-05-05 Won Ik Cho , Sung Jun Cheon , Woo Hyun Kang , Ji Won Kim , Nam Soo Kim

Large-scale natural language generation requires the integration of vast amounts of knowledge: lexical, grammatical, and conceptual. A robust generator must be able to operate well even when pieces of knowledge are missing. It must also be…

cmp-lg · 计算机科学 2008-02-03 Kevin Knight , Vasileios Hatzivassiloglou

We describe a resource-based method of morphological annotation of written Korean text. Korean is an agglutinative language. The output of our system is a graph of morphemes annotated with accurate linguistic information. The language…

计算与语言 · 计算机科学 2007-11-22 Hyun-Gue Huh , Eric Laporte

We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that…

计算与语言 · 计算机科学 2019-11-01 Kang Min Yoo , Taeuk Kim , Sang-goo Lee

Chinese couplet is a special form of poetry composed of complex syntax with ancient Chinese language. Due to the complexity of semantic and grammatical rules, creation of a suitable couplet is a formidable challenge. This paper presents a…

计算与语言 · 计算机科学 2021-12-06 Kuan-Yu Chiang , Shihao Lin , Joe Chen , Qian Yin , Qizhen Jin

Physical commonsense reasoning datasets like PIQA are predominantly English-centric and lack cultural diversity. We introduce Ko-PIQA, a Korean physical commonsense reasoning dataset that incorporates cultural context. Starting from 3.01…

计算与语言 · 计算机科学 2025-09-30 Dasol Choi , Jungwhan Kim , Guijin Son

This paper attempts to analyze the Korean sentence classification system for a chatbot. Sentence classification is the task of classifying an input sentence based on predefined categories. However, spelling or space error contained in the…

计算与语言 · 计算机科学 2021-06-08 DongHyun Choi , IlNam Park , Myeong Cheol Shin , EungGyun Kim , Dong Ryeol Shin

E-learning systems should deliver contents that reflect various phenomena of the language as it is used. In addition to formal Korean, e-learning systems that would include real-world Korean expressions such as those in web documents,…

计算与语言 · 计算机科学 2026-05-29 Sang-Taek Park , Ae-Lim Ahn , Eric Laporte , Jee-Sun Nam

The sense analysis is still critical problem in machine translation system, especially such as English-Korean translation which the syntactical different between source and target languages is very great. We suggest a method for selecting…

计算与语言 · 计算机科学 2013-05-16 Hyonil Kim , Changil Choe