中文
相关论文

相关论文: CCPM: A Chinese Classical Poetry Matching Dataset

200 篇论文

The quality of natural language texts in fine-tuning datasets plays a critical role in the performance of generative models, particularly in computational creativity tasks such as poem or song lyric generation. Fluency defects in generated…

计算与语言 · 计算机科学 2025-05-08 Ilya Koziev

Chinese Spelling Check (CSC) is a task to detect and correct spelling errors in Chinese natural language. Existing methods have made attempts to incorporate the similarity knowledge between Chinese characters. However, they take the…

计算与语言 · 计算机科学 2020-05-14 Xingyi Cheng , Weidi Xu , Kunlong Chen , Shaohua Jiang , Feng Wang , Taifeng Wang , Wei Chu , Yuan Qi

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese…

计算与语言 · 计算机科学 2022-09-13 Yudong Li , Yuqing Zhang , Zhe Zhao , Linlin Shen , Weijie Liu , Weiquan Mao , Hui Zhang

Recent work in training large language models (LLMs) to follow natural language instructions has opened up exciting opportunities for natural language interface design. Building on the prior success of LLMs in the realm of computer-assisted…

计算与语言 · 计算机科学 2022-10-26 Tuhin Chakrabarty , Vishakh Padmakumar , He He

Recognizing a piece of writing as a poem or prose is usually easy for the majority of people; however, only specialists can determine which meter a poem belongs to. In this paper, we build Recurrent Neural Network (RNN) models that can…

计算与语言 · 计算机科学 2019-05-15 Waleed A. Yousef , Omar M. Ibrahime , Taha M. Madbouly , Moustafa A. Mahmoud

Cant is important for understanding advertising, comedies and dog-whistle politics. However, computational research on cant is hindered by a lack of available datasets. In this paper, we propose a large and diverse Chinese dataset for…

计算与语言 · 计算机科学 2021-06-09 Canwen Xu , Wangchunshu Zhou , Tao Ge , Ke Xu , Julian McAuley , Furu Wei

This study introduces KPoEM (Korean Poetry Emotion Mapping), a novel dataset that serves as a foundation for both emotion-centered analysis and generative applications in modern Korean poetry. Despite advancements in NLP, poetry remains…

计算与语言 · 计算机科学 2026-01-15 Iro Lim , Haein Ji , Byungjun Kim

Recent studies in sequence-to-sequence learning demonstrate that RNN encoder-decoder structure can successfully generate Chinese poetry. However, existing methods can only generate poetry with a given first line or user's intent theme. In…

计算与语言 · 计算机科学 2019-11-20 Dayiheng Liu , Quan Guo , Wubo Li , Jiancheng Lv

Large vision-language models (VLMs) have demonstrated remarkable abilities in understanding everyday content. However, their performance in the domain of art, particularly culturally rich art forms, remains less explored. As a pearl of…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Tuo Zhang , Tiantian Feng , Yibin Ni , Mengqin Cao , Ruying Liu , Katharine Butler , Yanjun Weng , Mi Zhang , Shrikanth S. Narayanan , Salman Avestimehr

We aim at segmenting words in the Complete Tang Poems (CTP). Although it is possible to do some research about CTP without doing full-scale word segmentation, we must move forward to word-level analysis of CTP for conducting advanced…

计算与语言 · 计算机科学 2019-08-29 Chao-Lin Liu

Automatic generation of natural language from images has attracted extensive attention. In this paper, we take one step further to investigate generation of poetic language (with multiple lines) to an image for automatic poetry creation.…

计算机视觉与模式识别 · 计算机科学 2018-10-11 Bei Liu , Jianlong Fu , Makoto P. Kato , Masatoshi Yoshikawa

The prevalence of rapidly evolving slang, neologisms, and highly stylized expressions in informal user-generated text, particularly on Chinese social media, poses significant challenges for Machine Translation (MT) benchmarking.…

计算与语言 · 计算机科学 2026-02-02 Kaiyan Zhao , Zheyong Xie , Zhongtao Miao , Xinze Lyu , Yao Hu , Shaosheng Cao

Classical Chinese Understanding (CCU) holds significant value in preserving and exploration of the outstanding traditional Chinese culture. Recently, researchers have attempted to leverage the potential of Large Language Models (LLMs) for…

计算与语言 · 计算机科学 2024-05-31 Jiahuan Cao , Yongxin Shi , Dezhi Peng , Yang Liu , Lianwen Jin

Whole word masking (WWM), which masks all subwords corresponding to a word at once, makes a better English BERT model. For the Chinese language, however, there is no subword because each token is an atomic character. The meaning of a word…

计算与语言 · 计算机科学 2022-03-03 Yong Dai , Linyang Li , Cong Zhou , Zhangyin Feng , Enbo Zhao , Xipeng Qiu , Piji Li , Duyu Tang

Large language models (LLMs) have obtained promising results in mathematical reasoning, which is a foundational skill for human intelligence. Most previous studies focus on improving and measuring the performance of LLMs based on textual…

计算与语言 · 计算机科学 2024-11-04 Wentao Liu , Qianjun Pan , Yi Zhang , Zhuo Liu , Ji Wu , Jie Zhou , Aimin Zhou , Qin Chen , Bo Jiang , Liang He

Metaphors are pervasive in communication, making them crucial for natural language processing (NLP). Previous research on automatic metaphor processing predominantly relies on training data consisting of English samples, which often reflect…

计算与语言 · 计算机科学 2025-06-10 Senqi Yang , Dongyu Zhang , Jing Ren , Ziqi Xu , Xiuzhen Zhang , Yiliao Song , Hongfei Lin , Feng Xia

Chinese couplet is a special form of poetry composed of complex syntax with ancient Chinese language. Due to the complexity of semantic and grammatical rules, creation of a suitable couplet is a formidable challenge. This paper presents a…

计算与语言 · 计算机科学 2021-12-06 Kuan-Yu Chiang , Shihao Lin , Joe Chen , Qian Yin , Qizhen Jin

The rapid advancement of large language models (LLMs) has reshaped the landscape of machine translation, yet challenges persist in preserving poetic intent, cultural heritage, and handling specialized terminology in Chinese-English…

计算与语言 · 计算机科学 2025-04-29 Li Weigang , Pedro Carvalho Brom

Chinese Spell Checking (CSC) aims to detect and correct Chinese spelling errors, which are mainly caused by the phonological or visual similarity. Recently, pre-trained language models (PLMs) promote the progress of CSC task. However, there…

计算与语言 · 计算机科学 2022-03-03 Yinghui Li , Qingyu Zhou , Yangning Li , Zhongli Li , Ruiyang Liu , Rongyi Sun , Zizhen Wang , Chao Li , Yunbo Cao , Hai-Tao Zheng

Machine Reading Comprehension (MRC) has become enormously popular recently and has attracted a lot of attention. However, existing reading comprehension datasets are mostly in English. To add diversity in reading comprehension datasets, in…

计算与语言 · 计算机科学 2018-03-16 Yiming Cui , Ting Liu , Zhipeng Chen , Wentao Ma , Shijin Wang , Guoping Hu