English
Related papers

Related papers: KOMBO: Korean Character Representations Based on t…

200 papers

Most pretrained language models rely on subword tokenization, which processes text as a sequence of subword tokens. However, different granularities of text, such as characters, subwords, and words, can contain different kinds of…

Computation and Language · Computer Science 2024-04-09 Yilin Wang , Xinyi Hu , Matthew R. Gormley

While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea before the 20th century, whose characters were borrowed from old…

Computation and Language · Computer Science 2025-01-22 Seyoung Song , Haneul Yoo , Jiho Jin , Kyunghyun Cho , Alice Oh

E-learning systems should deliver contents that reflect various phenomena of the language as it is used. In addition to formal Korean, e-learning systems that would include real-world Korean expressions such as those in web documents,…

Computation and Language · Computer Science 2026-05-29 Sang-Taek Park , Ae-Lim Ahn , Eric Laporte , Jee-Sun Nam

Speaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions. These models primarily extract embeddings for speaker identification but…

Computation and Language · Computer Science 2025-08-26 Massa Baali , Shuo Han , Syed Abdul Hannan , Purusottam Samal , Karanveer Singh , Soham Deshmukh , Rita Singh , Bhiksha Raj

We implemented a high-performance optical character recognition model for classical handwritten documents using data augmentation with highly variable cropping within the document region. Optical character recognition in handwritten…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Joonmo Ahna , Taehong Jang , Quan Fengnyu , Hyungil Lee , Jaehyuk Lee , Sojung Lucia Kim

We present a two-stage fine-tuning approach to make the large language model Qwen3 14B "think" natively in Korean. In the first stage, supervised fine-tuning (SFT) on a high-quality Korean reasoning dataset establishes a strong foundation…

Computation and Language · Computer Science 2025-08-15 Jungyup Lee , Jemin Kim , Sang Park , SeungJae Lee

Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate more fluent and coherent content. Current multilingual…

Computation and Language · Computer Science 2024-12-11 Xiaonan Wang , Jinyoung Yeo , Joon-Ho Lim , Hansaem Kim

Existing research generally treats Chinese character as a minimum unit for representation. However, such Chinese character representation will suffer two bottlenecks: 1) Learning bottleneck, the learning cannot benefit from its rich…

Computation and Language · Computer Science 2022-11-24 Zhijun Wang , Xuebo Liu , Min Zhang

We describe a resource-based method of morphological annotation of written Korean text. Korean is an agglutinative language. The output of our system is a graph of morphemes annotated with accurate linguistic information. The language…

Computation and Language · Computer Science 2007-11-22 Hyun-Gue Huh , Eric Laporte

In this study, we propose a morpheme-based scheme for Korean dependency parsing and adopt the proposed scheme to Universal Dependencies. We present the linguistic rationale that illustrates the motivation and the necessity of adopting the…

Computation and Language · Computer Science 2022-09-21 Yige Chen , Eunkyul Leah Jo , Yundong Yao , KyungTae Lim , Miikka Silfverberg , Francis M. Tyers , Jungyeul Park

As language models become increasingly deployed in online environments, toxicity detection and detoxification have received growing attention. Existing studies primarily focus on non-obfuscated text, which limits robustness when users…

Computation and Language · Computer Science 2026-05-29 Yejin Lee , Su-Hyeon Kim , Hyundong Jin , Dayoung Kim , Yeonsoo Kim , Yo-Sub Han

A Lite BERT (ALBERT) has been introduced to scale up deep bidirectional representation learning for natural languages. Due to the lack of pretrained ALBERT models for Korean language, the best available practice is the multilingual model or…

Computation and Language · Computer Science 2021-01-28 Hyunjae Lee , Jaewoong Yoon , Bonggyu Hwang , Seongho Joe , Seungjai Min , Youngjune Gwon

We introduce KFinEval-Pilot, a benchmark suite specifically designed to evaluate large language models (LLMs) in the Korean financial domain. Addressing the limitations of existing English-centric benchmarks, KFinEval-Pilot comprises over…

Syntactic elements, such as word order and case markers, are fundamental in natural language processing. Recent studies show that syntactic information boosts language model performance and offers clues for people to understand their…

Computation and Language · Computer Science 2024-07-15 Jong Myoung Kim , Young-Jun Lee , Yong-jin Han , Sangkeun Jung , Ho-Jin Choi

The Korean wave, which denotes the global popularity of South Korea's cultural economy, contributes to the increasing demand for the Korean language. However, as there does not exist any application for foreigners to learn Korean, this…

Computation and Language · Computer Science 2022-05-05 Minjong Cheon , Minseon Kim , Hanseon Joo

Large Language Models (LLMs) demonstrate strong reasoning and self-correction abilities in high-resource languages like English, but their performance remains limited in low-resource languages such as Korean. In this study, we investigate…

Computation and Language · Computer Science 2026-01-12 Hongjin Kim , Jaewook Lee , Kiyoung Lee , Jong-hun Shin , Soojong Lim , Oh-Woog Kwon

This paper presents a HMM-based speech recognition engine and its integration into direct manipulation interfaces for Korean document editor. Speech recognition can reduce typical tedious and repetitive actions which are inevitable in…

cmp-lg · Computer Science 2016-08-31 Geunbae Lee , Jong-Hyeok Lee , Sangeok Kim

The Sejong dictionary dataset offers a valuable resource, providing extensive coverage of morphology, syntax, and semantic representation. This dataset can be utilized to explore linguistic information in greater depth. The labeled…

Computation and Language · Computer Science 2025-04-04 Seohyun Song , Eunkyul Leah Jo , Yige Chen , Jeen-Pyo Hong , Kyuwon Kim , Jin Wee , Miyoung Kang , KyungTae Lim , Jungyeul Park , Chulwoo Park

Despite their remarkable progress across diverse domains, Large Language Models (LLMs) consistently fail at simple character-level tasks, such as counting letters in words, due to a fundamental limitation: tokenization. In this work, we…

Computation and Language · Computer Science 2025-09-17 Adrian Cosma , Stefan Ruseti , Emilian Radoi , Mihai Dascalu

The rapid advancement of large language models (LLMs) increases the difficulty of distinguishing between human-written and LLM-generated text. Detecting LLM-generated text is crucial for upholding academic integrity, preventing plagiarism,…

Computation and Language · Computer Science 2025-09-22 Shinwoo Park , Shubin Kim , Do-Kyung Kim , Yo-Sub Han