English
Related papers

Related papers: COMPILING: A Benchmark Dataset for Chinese Complex…

200 papers

Lexical complexity prediction (LCP) is the task of predicting the complexity of words in a text on a continuous scale. It plays a vital role in simplifying or annotating complex words to assist readers. To study lexical complexity in…

Computation and Language · Computer Science 2023-07-03 Yusuke Ide , Masato Mita , Adam Nohejl , Hiroki Ouchi , Taro Watanabe

Holistically measuring societal biases of large language models is crucial for detecting and reducing ethical risks in highly capable AI models. In this work, we present a Chinese Bias Benchmark dataset that consists of over 100K questions…

Computation and Language · Computer Science 2023-06-29 Yufei Huang , Deyi Xiong

This paper presents a systematic investigation into the constrained generation capabilities of large language models (LLMs) in producing Songci, a classical Chinese poetry form characterized by strict structural, tonal, and rhyme…

Computation and Language · Computer Science 2026-02-19 Zhan Qu , Shuzhou Yuan , Michael Färber

Chinese text segmentation is a well-known and difficult problem. On one side, there is not a simple notion of "word" in Chinese language making really hard to implement rule-based systems to segment written texts, thus lexicons and…

Computation and Language · Computer Science 2007-05-23 Daniel Gayo-Avello

Previous works of video captioning aim to objectively describe the video's actual content, which lacks subjective and attractive expression, limiting its practical application scenarios. Video titling is intended to achieve this goal, but…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Ziqi Zhang , Yuxin Chen , Zongyang Ma , Zhongang Qi , Chunfeng Yuan , Bing Li , Ying Shan , Weiming Hu

The ability to generate natural-language questions with controlled complexity levels is highly desirable as it further expands the applicability of question generation. In this paper, we propose an end-to-end neural complexity-controllable…

Computation and Language · Computer Science 2021-10-14 Sheng Bi , Xiya Cheng , Yuan-Fang Li , Lizhen Qu , Shirong Shen , Guilin Qi , Lu Pan , Yinlin Jiang

Idioms are an important language phenomenon in Chinese, but idiom translation is notoriously hard. Current machine translation models perform poorly on idiom translation, while idioms are sparse in many translation datasets. We present…

Computation and Language · Computer Science 2022-02-22 Kenan Tang

Modern machine learning relies on datasets to develop and validate research ideas. Given the growth of publicly available data, finding the right dataset to use is increasingly difficult. Any research question imposes explicit and implicit…

Information Retrieval · Computer Science 2023-06-08 Vijay Viswanathan , Luyu Gao , Tongshuang Wu , Pengfei Liu , Graham Neubig

We study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Dictionary definitions are an essential resource for learning word senses, but manually creating them is costly,…

Computation and Language · Computer Science 2026-01-06 Yusuke Ide , Adam Nohejl , Joshua Tanner , Hitomi Yanaka , Christopher Lindsay , Taro Watanabe

This paper introduces the first dataset for evaluating English-Chinese Bilingual Contextual Word Similarity, namely BCWS (https://github.com/MiuLab/BCWS). The dataset consists of 2,091 English-Chinese word pairs with the corresponding…

Computation and Language · Computer Science 2018-10-23 Ta-Chung Chi , Ching-Yen Shih , Yun-Nung Chen

Analogical reasoning is effective in capturing linguistic regularities. This paper proposes an analogical reasoning task on Chinese. After delving into Chinese lexical knowledge, we sketch 68 implicit morphological relations and 28 explicit…

Computation and Language · Computer Science 2018-08-27 Shen Li , Zhe Zhao , Renfen Hu , Wensi Li , Tao Liu , Xiaoyong Du

Generating coherent, grammatically correct, and meaningful text is very challenging, however, it is crucial to many modern NLP systems. So far, research has mostly focused on English language, for other languages both standardized datasets,…

Computation and Language · Computer Science 2020-05-07 Zein Shaheen , Gerhard Wohlgenannt , Bassel Zaity , Dmitry Mouromtsev , Vadim Pak

In view of the poor robustness of existing Chinese grammatical error correction models on attack test sets and large model parameters, this paper uses the method of knowledge distillation to compress model parameters and improve the…

Computation and Language · Computer Science 2022-09-01 Peng Xia , Yuechi Zhou , Ziyan Zhang , Zecheng Tang , Juntao Li

Automatic literature review generation is one of the most challenging tasks in natural language processing. Although large language models have tackled literature review generation, the absence of large-scale datasets has been a stumbling…

Computation and Language · Computer Science 2023-05-25 Tetsu Kasanishi , Masaru Isonuma , Junichiro Mori , Ichiro Sakata

This paper explores the task of Difficulty-Controllable Question Generation (DCQG), which aims at generating questions with required difficulty levels. Previous research on this task mainly defines the difficulty of a question as whether it…

Computation and Language · Computer Science 2021-05-26 Yi Cheng , Siyao Li , Bang Liu , Ruihui Zhao , Sujian Li , Chenghua Lin , Yefeng Zheng

Word embedding is a modern distributed word representations approach widely used in many natural language processing tasks. Converting the vocabulary in a legal document into a word embedding model facilitates subjecting legal documents to…

Computation and Language · Computer Science 2022-04-01 Chun-Hsien Lin , Pu-Jen Cheng

Enhancing the attribution in large language models (LLMs) is a crucial task. One feasible approach is to enable LLMs to cite external sources that support their generations. However, existing datasets and evaluation methods in this domain…

Computation and Language · Computer Science 2024-05-30 Haolin Deng , Chang Wang , Xin Li , Dezhang Yuan , Junlang Zhan , Tianhua Zhou , Jin Ma , Jun Gao , Ruifeng Xu

With the development of pre-trained models and the incorporation of phonetic and graphic information, neural models have achieved high scores in Chinese Spelling Check (CSC). However, it does not provide a comprehensive reflection of the…

Computation and Language · Computer Science 2023-07-26 Xunjian Yin , Xiaojun Wan

Abbreviation is a common phenomenon across languages, especially in Chinese. In most cases, if an expression can be abbreviated, its abbreviation is used more often than its fully expanded forms, since people tend to convey information in a…

Computation and Language · Computer Science 2017-12-19 Yi Zhang , Xu Sun

The Chinese character riddle is a unique form of cultural entertainment specific to the Chinese language. It typically comprises two parts: the riddle description and the solution. The solution to the riddle is a single character, while the…

Computation and Language · Computer Science 2023-09-26 Fan Xu , Yunxiang Zhang , Xiaojun Wan
‹ Prev 1 4 5 6 7 8 10 Next ›