English
Related papers

Related papers: RedWhale: An Adapted Korean LLM Through Efficient …

200 papers

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

Computation and Language · Computer Science 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

Pre-training Large Language Models (LLMs) on high-quality, meticulously curated datasets is widely recognized as critical for enhancing their performance and generalization capabilities. This study explores the untapped potential of Common…

Large Language Models (LLMs) demonstrate strong reasoning and self-correction abilities in high-resource languages like English, but their performance remains limited in low-resource languages such as Korean. In this study, we investigate…

Computation and Language · Computer Science 2026-01-12 Hongjin Kim , Jaewook Lee , Kiyoung Lee , Jong-hun Shin , Soojong Lim , Oh-Woog Kwon

Large language models (LLMs) use pretraining to predict the subsequent word; however, their expansion requires significant computing resources. Numerous big tech companies and research institutes have developed multilingual LLMs (MLLMs) to…

Since state-of-the-art LLMs often underperform in languages other than English or Chinese, improving the capability of LLMs in new languages has become an essential task. Moreover, LLMs' entire end-to-end training process remains largely…

Computation and Language · Computer Science 2025-06-30 Jinpyo Kim , Gyeongje Cho , Chanwoo Park , Jongwon Park , Jongmin Kim , Yeonkyoun So , Jaejin Lee

The recent progress of AI can be largely attributed to large language models (LLMs). However, their escalating memory requirements introduce challenges for machine learning (ML) researchers and engineers. Addressing this requires developers…

Machine Learning · Computer Science 2024-06-14 Bowen Tan , Yun Zhu , Lijuan Liu , Hongyi Wang , Yonghao Zhuang , Jindong Chen , Eric Xing , Zhiting Hu

We introduce Korean Language Understanding Evaluation (KLUE) benchmark. KLUE is a collection of 8 Korean natural language understanding (NLU) tasks, including Topic Classification, SemanticTextual Similarity, Natural Language Inference,…

High-resource languages such as English, enables the pretraining of high-quality large language models (LLMs). The same can not be said for most other languages as LLMs still underperform for non-English languages, likely due to a gap in…

Computation and Language · Computer Science 2025-02-20 Jiayi Wang , Yao Lu , Maurice Weber , Max Ryabinin , David Adelani , Yihong Chen , Raphael Tang , Pontus Stenetorp

Large language models (LLMs) demonstrate exceptional performance on complex reasoning tasks. However, despite their strong reasoning capabilities in high-resource languages (e.g., English and Chinese), a significant performance gap persists…

Computation and Language · Computer Science 2025-02-03 Hyunwoo Ko , Guijin Son , Dasol Choi

This work presents the first large-scale investigation into constructing a fully open bilingual large language model (LLM) for a non-English language, specifically Korean, trained predominantly on synthetic data. We introduce KORMo-10B, a…

Computation and Language · Computer Science 2025-10-13 Minjun Kim , Hyeonseok Lim , Hangyeol Yoo , Inho Won , Seungwoo Song , Minkyung Cho , Junhun Yuk , Changsu Choi , Dongjae Shin , Huige Lee , Hoyun Song , Alice Oh , Kyungtae Lim

Large Language Models (LLMs) have demonstrated profound impact on Natural Language Processing (NLP) tasks. However, their effective deployment across diverse domains often require domain-specific adaptation strategies, as generic models may…

Artificial Intelligence · Computer Science 2025-10-15 Jingyi Wang , Hongyuan Zhu , Ye Niu , Yunhui Deng

We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean,…

Computation and Language · Computer Science 2024-04-16 Kang Min Yoo , Jaegeun Han , Sookyo In , Heewon Jeon , Jisu Jeong , Jaewook Kang , Hyunwook Kim , Kyung-Min Kim , Munhyong Kim , Sungju Kim , Donghyun Kwak , Hanock Kwak , Se Jung Kwon , Bado Lee , Dongsoo Lee , Gichang Lee , Jooho Lee , Baeseong Park , Seongjin Shin , Joonsang Yu , Seolki Baek , Sumin Byeon , Eungsup Cho , Dooseok Choe , Jeesung Han , Youngkyun Jin , Hyein Jun , Jaeseung Jung , Chanwoong Kim , Jinhong Kim , Jinuk Kim , Dokyeong Lee , Dongwook Park , Jeong Min Sohn , Sujung Han , Jiae Heo , Sungju Hong , Mina Jeon , Hyunhoon Jung , Jungeun Jung , Wangkyo Jung , Chungjoon Kim , Hyeri Kim , Jonghyun Kim , Min Young Kim , Soeun Lee , Joonhee Park , Jieun Shin , Sojin Yang , Jungsoon Yoon , Hwaran Lee , Sanghwan Bae , Jeehwan Cha , Karl Gylleus , Donghoon Ham , Mihak Hong , Youngki Hong , Yunki Hong , Dahyun Jang , Hyojun Jeon , Yujin Jeon , Yeji Jeong , Myunggeun Ji , Yeguk Jin , Chansong Jo , Shinyoung Joo , Seunghwan Jung , Adrian Jungmyung Kim , Byoung Hoon Kim , Hyomin Kim , Jungwhan Kim , Minkyoung Kim , Minseung Kim , Sungdong Kim , Yonghee Kim , Youngjun Kim , Youngkwan Kim , Donghyeon Ko , Dughyun Lee , Ha Young Lee , Jaehong Lee , Jieun Lee , Jonghyun Lee , Jongjin Lee , Min Young Lee , Yehbin Lee , Taehong Min , Yuri Min , Kiyoon Moon , Hyangnam Oh , Jaesun Park , Kyuyon Park , Younghun Park , Hanbae Seo , Seunghyun Seo , Mihyun Sim , Gyubin Son , Matt Yeo , Kyung Hoon Yeom , Wonjoon Yoo , Myungin You , Doheon Ahn , Homin Ahn , Joohee Ahn , Seongmin Ahn , Chanwoo An , Hyeryun An , Junho An , Sang-Min An , Boram Byun , Eunbin Byun , Jongho Cha , Minji Chang , Seunggyu Chang , Haesong Cho , Youngdo Cho , Dalnim Choi , Daseul Choi , Hyoseok Choi , Minseong Choi , Sangho Choi , Seongjae Choi , Wooyong Choi , Sewhan Chun , Dong Young Go , Chiheon Ham , Danbi Han , Jaemin Han , Moonyoung Hong , Sung Bum Hong , Dong-Hyun Hwang , Seongchan Hwang , Jinbae Im , Hyuk Jin Jang , Jaehyung Jang , Jaeni Jang , Sihyeon Jang , Sungwon Jang , Joonha Jeon , Daun Jeong , Joonhyun Jeong , Kyeongseok Jeong , Mini Jeong , Sol Jin , Hanbyeol Jo , Hanju Jo , Minjung Jo , Chaeyoon Jung , Hyungsik Jung , Jaeuk Jung , Ju Hwan Jung , Kwangsun Jung , Seungjae Jung , Soonwon Ka , Donghan Kang , Soyoung Kang , Taeho Kil , Areum Kim , Beomyoung Kim , Byeongwook Kim , Daehee Kim , Dong-Gyun Kim , Donggook Kim , Donghyun Kim , Euna Kim , Eunchul Kim , Geewook Kim , Gyu Ri Kim , Hanbyul Kim , Heesu Kim , Isaac Kim , Jeonghoon Kim , Jihye Kim , Joonghoon Kim , Minjae Kim , Minsub Kim , Pil Hwan Kim , Sammy Kim , Seokhun Kim , Seonghyeon Kim , Soojin Kim , Soong Kim , Soyoon Kim , Sunyoung Kim , Taeho Kim , Wonho Kim , Yoonsik Kim , You Jin Kim , Yuri Kim , Beomseok Kwon , Ohsung Kwon , Yoo-Hwan Kwon , Anna Lee , Byungwook Lee , Changho Lee , Daun Lee , Dongjae Lee , Ha-Ram Lee , Hodong Lee , Hwiyeong Lee , Hyunmi Lee , Injae Lee , Jaeung Lee , Jeongsang Lee , Jisoo Lee , Jongsoo Lee , Joongjae Lee , Juhan Lee , Jung Hyun Lee , Junghoon Lee , Junwoo Lee , Se Yun Lee , Sujin Lee , Sungjae Lee , Sungwoo Lee , Wonjae Lee , Zoo Hyun Lee , Jong Kun Lim , Kun Lim , Taemin Lim , Nuri Na , Jeongyeon Nam , Kyeong-Min Nam , Yeonseog Noh , Biro Oh , Jung-Sik Oh , Solgil Oh , Yeontaek Oh , Boyoun Park , Cheonbok Park , Dongju Park , Hyeonjin Park , Hyun Tae Park , Hyunjung Park , Jihye Park , Jooseok Park , Junghwan Park , Jungsoo Park , Miru Park , Sang Hee Park , Seunghyun Park , Soyoung Park , Taerim Park , Wonkyeong Park , Hyunjoon Ryu , Jeonghun Ryu , Nahyeon Ryu , Soonshin Seo , Suk Min Seo , Yoonjeong Shim , Kyuyong Shin , Wonkwang Shin , Hyun Sim , Woongseob Sim , Hyejin Soh , Bokyong Son , Hyunjun Son , Seulah Son , Chi-Yun Song , Chiyoung Song , Ka Yeon Song , Minchul Song , Seungmin Song , Jisung Wang , Yonggoo Yeo , Myeong Yeon Yi , Moon Bin Yim , Taehwan Yoo , Youngjoon Yoo , Sungmin Yoon , Young Jin Yoon , Hangyeol Yu , Ui Seon Yu , Xingdong Zuo , Jeongin Bae , Joungeun Bae , Hyunsoo Cho , Seonghyun Cho , Yongjin Cho , Taekyoon Choi , Yera Choi , Jiwan Chung , Zhenghui Han , Byeongho Heo , Euisuk Hong , Taebaek Hwang , Seonyeol Im , Sumin Jegal , Sumin Jeon , Yelim Jeong , Yonghyun Jeong , Can Jiang , Juyong Jiang , Jiho Jin , Ara Jo , Younghyun Jo , Hoyoun Jung , Juyoung Jung , Seunghyeong Kang , Dae Hee Kim , Ginam Kim , Hangyeol Kim , Heeseung Kim , Hyojin Kim , Hyojun Kim , Hyun-Ah Kim , Jeehye Kim , Jin-Hwa Kim , Jiseon Kim , Jonghak Kim , Jung Yoon Kim , Rak Yeong Kim , Seongjin Kim , Seoyoon Kim , Sewon Kim , Sooyoung Kim , Sukyoung Kim , Taeyong Kim , Naeun Ko , Bonseung Koo , Heeyoung Kwak , Haena Kwon , Youngjin Kwon , Boram Lee , Bruce W. Lee , Dagyeong Lee , Erin Lee , Euijin Lee , Ha Gyeong Lee , Hyojin Lee , Hyunjeong Lee , Jeeyoon Lee , Jeonghyun Lee , Jongheok Lee , Joonhyung Lee , Junhyuk Lee , Mingu Lee , Nayeon Lee , Sangkyu Lee , Se Young Lee , Seulgi Lee , Seung Jin Lee , Suhyeon Lee , Yeonjae Lee , Yesol Lee , Youngbeom Lee , Yujin Lee , Shaodong Li , Tianyu Liu , Seong-Eun Moon , Taehong Moon , Max-Lasse Nihlenramstroem , Wonseok Oh , Yuri Oh , Hongbeen Park , Hyekyung Park , Jaeho Park , Nohil Park , Sangjin Park , Jiwon Ryu , Miru Ryu , Simo Ryu , Ahreum Seo , Hee Seo , Kangdeok Seo , Jamin Shin , Seungyoun Shin , Heetae Sin , Jiangping Wang , Lei Wang , Ning Xiang , Longxiang Xiao , Jing Xu , Seonyeong Yi , Haanju Yoo , Haneul Yoo , Hwanhee Yoo , Liang Yu , Youngjae Yu , Weijie Yuan , Bo Zeng , Qian Zhou , Kyunghyun Cho , Jung-Woo Ha , Joonsuk Park , Jihyun Hwang , Hyoung Jo Kwon , Soonyong Kwon , Jungyeon Lee , Seungho Lee , Seonghyeon Lim , Hyunkyung Noh , Seungho Choi , Sang-Woo Lee , Jung Hwa Lim , Nako Sung

Cross-lingual continual pre-training of large language models (LLMs) initially trained on English corpus allows us to leverage the vast amount of English language resources and reduce the pre-training cost. In this study, we constructed…

Computation and Language · Computer Science 2024-04-30 Kazuki Fujii , Taishi Nakamura , Mengsay Loem , Hiroki Iida , Masanari Ohi , Kakeru Hattori , Hirai Shota , Sakae Mizuki , Rio Yokota , Naoaki Okazaki

The advancements in the Large Language Model (LLM) have helped in solving several problems related to language processing. Most of the researches have focused on the English language only, because of its popularity and abundance on the…

Computation and Language · Computer Science 2024-12-31 Sanjay Chouhan , Shubha Brata Nath , Aparajita Dutta

Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications. Pretraining on vast web-scale data has laid the foundation for these models, yet the research community is now…

Instruction Tuning on Large Language Models is an essential process for model to function well and achieve high performance in specific tasks. Accordingly, in mainstream languages such as English, instruction-based datasets are being…

Computation and Language · Computer Science 2024-03-26 Dongjun Jang , Sungjoo Byun , Hyemi Jo , Hyopil Shin

Pre-training is crucial for large language models (LLMs), as it is when most representations and capabilities are acquired. However, natural language pre-training has problems: high-quality text is finite, it contains human biases, and it…

Machine Learning · Computer Science 2026-03-12 Dan Lee , Seungwook Han , Akarsh Kumar , Pulkit Agrawal

A Lite BERT (ALBERT) has been introduced to scale up deep bidirectional representation learning for natural languages. Due to the lack of pretrained ALBERT models for Korean language, the best available practice is the multilingual model or…

Computation and Language · Computer Science 2021-01-28 Hyunjae Lee , Jaewoong Yoon , Bonggyu Hwang , Seongho Joe , Seungjai Min , Youngjune Gwon

Large language models (LLMs) are routinely pre-trained on billions of tokens, only to start the process over again once new data becomes available. A much more efficient solution is to continually pre-train these models, saving significant…

The rapid advancement of large language models (LLMs) increases the difficulty of distinguishing between human-written and LLM-generated text. Detecting LLM-generated text is crucial for upholding academic integrity, preventing plagiarism,…

Computation and Language · Computer Science 2025-09-22 Shinwoo Park , Shubin Kim , Do-Kyung Kim , Yo-Sub Han
‹ Prev 1 2 3 10 Next ›