English
Related papers

Related papers: Synth-Empathy: Towards High-Quality Synthetic Empa…

200 papers

The scarcity of domain-specific dialogue datasets limits the development of dialogue systems across applications. Existing research is constrained by general or niche datasets that lack sufficient scale for training dialogue systems. To…

Computation and Language · Computer Science 2025-02-11 Sathya Krishnan Suresh , Wu Mengjun , Tushar Pranav , Eng Siong Chng

Modern studies of societal phenomena rely on the availability of large datasets capturing attributes and activities of synthetic, city-level, populations. For instance, in epidemiology, synthetic population datasets are necessary to study…

Databases · Computer Science 2016-02-26 Hao Wu , Yue Ning , Prithwish Chakraborty , Jilles Vreeken , Nikolaj Tatti , Naren Ramakrishnan

Product information extraction is crucial for e-commerce services, but obtaining high-quality labeled datasets remains challenging. We present a systematic approach for generating synthetic e-commerce product data using Large Language…

Computation and Language · Computer Science 2026-01-09 Virginia Negri , Víctor Martínez Gómez , Sergio A. Balanya , Subburam Rajaram

Large Language models (LLMs), while powerful, exhibit harmful social biases. Debiasing is often challenging due to computational costs, data constraints, and potential degradation of multi-task language capabilities. This work introduces a…

Computation and Language · Computer Science 2024-09-17 Pengrui Han , Rafal Kocielnik , Adhithya Saravanan , Roy Jiang , Or Sharir , Anima Anandkumar

We propose CoT-Self-Instruct, a synthetic data generation method that instructs LLMs to first reason and plan via Chain-of-Thought (CoT) based on given seed tasks, and then generate a new synthetic example of similar quality and complexity.…

Artificial Intelligence · Computer Science 2025-09-04 Ping Yu , Jack Lanchantin , Tianlu Wang , Weizhe Yuan , Olga Golovneva , Ilia Kulikov , Sainbayar Sukhbaatar , Jason Weston , Jing Xu

Generative models have been showing potential for producing data in mass. This study explores the enhancement of clinical natural language processing performance by utilizing synthetic data generated from advanced language models. Promising…

Computation and Language · Computer Science 2024-03-29 Shan Chen , Jack Gallifant , Marco Guevara , Yanjun Gao , Majid Afshar , Timothy Miller , Dmitriy Dligach , Danielle S. Bitterman

Effective toxic content detection relies heavily on high-quality and diverse data, which serve as the foundation for robust content moderation models. Synthetic data has become a common approach for training models across various NLP tasks.…

Computation and Language · Computer Science 2025-02-25 Zheng Hui , Zhaoxiao Guo , Hang Zhao , Juanyong Duan , Lin Ai , Yinheng Li , Julia Hirschberg , Congrui Huang

Supervised fine-tuning (SFT) of large language models (LLMs) for specialized tasks requires high-quality datasets, but manual curation is prohibitively expensive. Synthetic data generation offers scalability, but its effectiveness relies on…

Machine Learning · Computer Science 2025-11-13 Shuzhen Bi , Chang Song , Siyu Song , Jinze Lv , Jian Chen , Xinyun Wang , Aimin Zhou , Hao Hao

Recent advancements in large language models (LLMs) have led to the development of highly potent models like OpenAI's ChatGPT. These models have exhibited exceptional performance in a variety of tasks, such as question answering, essay…

Computation and Language · Computer Science 2023-04-12 Ruixiang Tang , Xiaotian Han , Xiaoqian Jiang , Xia Hu

Deep generative models for tabular data (GANs, diffusion models, and LLM-based generators) exhibit highly non-uniform behavior across datasets; the best-performing synthesizer family depends strongly on distributional stressors such as…

Machine Learning · Computer Science 2026-04-02 Hochan Son , Xiaofeng Lin , Jason Ni , Guang Cheng

Recent advances in large language models (LLMs) have enabled human-like social simulations at unprecedented scale and fidelity, offering new opportunities for computational social science. A key challenge, however, is the construction of…

Computation and Language · Computer Science 2025-10-07 Zhengyu Hu , Jianxun Lian , Zheyuan Xiao , Max Xiong , Yuxuan Lei , Tianfu Wang , Kaize Ding , Ziang Xiao , Nicholas Jing Yuan , Xing Xie

This research explores a hybrid approach to fine-tuning large language models (LLMs) by integrating real-world and synthetic data to boost model performance, particularly in generating accurate and contextually relevant responses. By…

Computation and Language · Computer Science 2024-10-15 Alexey Zhezherau , Alexei Yanockin

Large Language Models (LLMs) have achieved remarkable success but remain data-inefficient, especially when learning from small, specialized corpora with limited and proprietary data. Existing synthetic data generation methods for continue…

Computation and Language · Computer Science 2025-09-16 Shengjie Ma , Xuhui Jiang , Chengjin Xu , Cehao Yang , Liyu Zhang , Jian Guo

Advancements in emotion aware language processing increasingly shape vital NLP applications ranging from conversational AI and affective computing to computational psychology and creative content generation. Existing emotion datasets either…

Computation and Language · Computer Science 2025-04-14 Vishal Gandhi , Sagar Gandhi

Linear programming (LP) problems are pervasive in real-life applications. However, despite their apparent simplicity, an untrained user may find it difficult to determine the linear model of their specific problem. We envisage the creation…

Computation and Language · Computer Science 2024-02-01 Yelaman Abdullin , Diego Molla-Aliod , Bahadorreza Ofoghi , John Yearwood , Qingyang Li

Recent research shows that greater numbers of people are turning to Large Language Models (LLMs) for emotional support, and that people rate LLM responses as more empathic than human-written responses. We suggest a reason for this success:…

Computation and Language · Computer Science 2026-04-10 Emma Gueorguieva , Hongli Zhan , Jina Suh , Javier Hernandez , Tatiana Lau , Junyi Jessy Li , Desmond C. Ong

Imbalanced classification and spurious correlation are common challenges in data science and machine learning. Both issues are linked to data imbalance, with certain groups of data samples significantly underrepresented, which in turn would…

Machine Learning · Statistics 2026-02-10 Ryumei Nakada , Yichen Xu , Lexin Li , Linjun Zhang

Empathetic dialogue requires not only recognizing a user's emotional state but also making strategy-aware, context-sensitive decisions throughout response generation. However, the lack of a comprehensive empathy strategy framework, explicit…

Computation and Language · Computer Science 2026-04-21 Hongru Ji , Yuyin Fan , Meng Zhao , Xianghua Li , Lianwei Wu , Chao Gao

Synthetic data sets are used across linguistic domains and NLP tasks, particularly in scenarios where authentic data is limited (or even non-existent). One such domain is that of clinical (healthcare) contexts, where there exist significant…

Computation and Language · Computer Science 2026-03-17 Steven Bedrick , A. Seza Doğruöz , Sergiu Nisioi

We present a synthetic data approach for instruction-tuning large language models (LLMs) for low-resource languages in a data-efficient manner, specifically focusing on Thai. We identify three key properties that contribute to the…

Computation and Language · Computer Science 2024-11-26 Parinthapat Pengpun , Can Udomcharoenchaikit , Weerayut Buaphet , Peerat Limkonchotiwat