English
Related papers

Related papers: A Large-Scale Chinese Short-Text Conversation Data…

200 papers

Existing rhetorical understanding and generation datasets or corpora primarily focus on single coarse-grained categories or fine-grained categories, neglecting the common interrelations between different rhetorical devices by treating them…

Computation and Language · Computer Science 2024-10-01 Nuowei Liu , Xinhao Chen , Hongyi Wu , Changzhi Sun , Man Lan , Yuanbin Wu , Xiaopeng Bai , Shaoguang Mao , Yan Xia

High-quality speech dialogue datasets are crucial for Speech-LLM development, yet existing acquisition methods face significant limitations. Human recordings incur high costs and privacy concerns, while synthetic approaches often lack…

Computation and Language · Computer Science 2025-04-01 Minghan Wang , Ye Bai , Yuxia Wang , Thuy-Trang Vu , Ehsan Shareghi , Gholamreza Haffari

We introduce CCI4.0, a large-scale bilingual pre-training dataset engineered for superior data quality and diverse human-like reasoning trajectory. CCI4.0 occupies roughly $35$ TB of disk space and comprises two sub-datasets: CCI4.0-M2-Base…

Computation and Language · Computer Science 2025-06-10 Guang Liu , Liangdong Wang , Jijie Li , Yang Yu , Yao Xu , Jiabei Chen , Yu Bai , Feng Liao , Yonghua Lin

The Qwen 2.5 3B base model was fine-tuned to generate contextually rich and engaging movie dialogue, leveraging the Cornell Movie-Dialog Corpus, a curated dataset of movie conversations. Due to the limitations in GPU computing and VRAM, the…

Computation and Language · Computer Science 2025-02-25 Kartik Gupta

Vision-Language Pre-training (VLP) models have achieved remarkable success by leveraging large-scale image-text pairs. While English-centric models like CLIP and SigLIP benefit from massive datasets (e.g., LAION-400M), the development of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Hengyu Shen , Tiancheng Gu , Bin Qin , Lan Wu , Yuling Wu , Shuo Tan , Zelong Sun , Jun Wang , Nan Wu , Xiang An , Weidong Cai , Ziyong Feng , Kaicheng Yang

The driving factors behind the development of large language models (LLMs) with impressive learning capabilities are their colossal model sizes and extensive training datasets. Along with the progress in natural language processing, LLMs…

Computation and Language · Computer Science 2023-09-19 Thuat Nguyen , Chien Van Nguyen , Viet Dac Lai , Hieu Man , Nghia Trung Ngo , Franck Dernoncourt , Ryan A. Rossi , Thien Huu Nguyen

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

Computation and Language · Computer Science 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

Conversational systems based on Large Language Models (LLMs), such as ChatGPT, show exceptional proficiency in context understanding and response generation. However, despite their impressive capabilities, they still possess limitations,…

Computation and Language · Computer Science 2023-10-17 Yang Deng , Lizi Liao , Liang Chen , Hongru Wang , Wenqiang Lei , Tat-Seng Chua

General-purpose large language models demonstrate notable capabilities in language comprehension and generation, achieving results that are comparable to, or even surpass, human performance in many natural language processing tasks.…

Computation and Language · Computer Science 2025-06-19 Shen Li , Renfen Hu , Lijun Wang

Learner corpus collects language data produced by L2 learners, that is second or foreign-language learners. This resource is of great relevance for second language acquisition research, foreign-language teaching, and automatic grammatical…

Computation and Language · Computer Science 2022-01-03 Yingying Wang , Cunliang Kong , Liner Yang , Yijun Wang , Xiaorong Lu , Renfen Hu , Shan He , Zhenghao Liu , Yun Chen , Erhong Yang , Maosong Sun

Dialogue topic shift detection is to detect whether an ongoing topic has shifted or should shift in a dialogue, which can be divided into two categories, i.e., response-known task and response-unknown task. Currently, only a few…

Computation and Language · Computer Science 2023-05-03 Jiangyi Lin , Yaxin Fan , Feng Jiang , Xiaomin Chu , Peifeng Li

A well-designed interactive human-like dialogue system is expected to take actions (e.g. smiling) and respond in a pattern similar to humans. However, due to the limitation of single-modality (only speech) or small volume of currently…

Human-Computer Interaction · Computer Science 2022-12-13 Zhiling Luo , Qiankun Shi , Sha Zhao , Wei Zhou , Haiqing Chen , Yuankai Ma , Haitao Leng

The rapid evolution of large language models (LLMs) has ushered in the need for comprehensive assessments of their performance across various dimensions. In this paper, we propose LFED, a Literary Fiction Evaluation Dataset, which aims to…

Computation and Language · Computer Science 2024-05-17 Linhao Yu , Qun Liu , Deyi Xiong

Although large language models (LLMs) demonstrate impressive performance for many language tasks, most of them can only handle texts a few thousand tokens long, limiting their applications on longer sequence inputs, such as books, reports,…

Computation and Language · Computer Science 2024-06-21 Yushi Bai , Xin Lv , Jiajie Zhang , Hongchang Lyu , Jiankai Tang , Zhidian Huang , Zhengxiao Du , Xiao Liu , Aohan Zeng , Lei Hou , Yuxiao Dong , Jie Tang , Juanzi Li

Dialogue systems, commonly known as chatbots, have gained escalating popularity in recent times due to their wide-spread applications in carrying out chit-chat conversations with users and task-oriented dialogues to accomplish various user…

Computation and Language · Computer Science 2024-06-18 Sahisnu Mazumder , Bing Liu

Large Language Models (LLMs) have attained the impressive capability to resolve a wide range of NLP tasks by fine-tuning high-quality instruction data. However, collecting human-written data of high quality, especially multi-turn dialogues,…

Computation and Language · Computer Science 2023-10-20 Dongjie Yang , Ruifeng Yuan , Yuantao Fan , Yifei Yang , Zili Wang , Shusen Wang , Hai Zhao

Large Language Models (LLMs) are becoming integral to modern software development workflows, assisting developers with code generation, API explanation, and iterative problem-solving through natural language conversations. Despite…

Software Engineering · Computer Science 2025-09-15 Suzhen Zhong , Ying Zou , Bram Adams