中文
相关论文

相关论文: Synthia: Scalable Grounded Persona Generation from…

200 篇论文

Large Language Models (LLMs) have democratized synthetic data generation, which in turn has the potential to simplify and broaden a wide gamut of NLP tasks. Here, we tackle a pervasive problem in synthetic data generation: its generative…

计算与语言 · 计算机科学 2023-05-25 Veniamin Veselovsky , Manoel Horta Ribeiro , Akhil Arora , Martin Josifoski , Ashton Anderson , Robert West

AI-based persona simulation -- often referred to as digital twin simulation -- is increasingly used for market research, recommender systems, and social sciences. Despite their flexibility, large language models (LLMs) often exhibit…

机器学习 · 计算机科学 2026-04-10 Grace Jiarui Fan , Chengpiao Huang , Tianyi Peng , Kaizheng Wang , Yuhang Wu

Large language models (LLMs) are known to generate biased responses where the opinions of certain groups and populations are underrepresented. Here, we present a novel approach to achieve controllable generation of specific viewpoints using…

计算与语言 · 计算机科学 2024-04-04 Junyi Li , Ninareh Mehrabi , Charith Peris , Palash Goyal , Kai-Wei Chang , Aram Galstyan , Richard Zemel , Rahul Gupta

A critical challenge in social science research is the high cost associated with experiments involving human participants. We identify Synthetic Discussion Generation (SDG), a novel Natural Language Processing (NLP) direction aimed at…

人机交互 · 计算机科学 2026-04-20 Dimitris Tsirmpas , Ion Androutsopoulos , John Pavlopoulos

Ensuring the generalisability of clinical machine learning (ML) models across diverse healthcare settings remains a significant challenge due to variability in patient demographics, disease prevalence, and institutional practices. Existing…

机器学习 · 计算机科学 2025-04-30 Bradley Segal , Joshua Fieggen , David Clifton , Lei Clifton

The ability to simulate human privacy decisions has significant implications for aligning autonomous agents with individual intent and conducting cost-effective, large-scale privacy-centric user studies. Prior approaches prompt Large…

密码学与安全 · 计算机科学 2026-05-11 Kassem Fawaz , Ren Yi , Octavian Suciu , Rishabh Khandelwal , Hamza Harkous , Nina Taft , Marco Gruteser

Recent advances enable Large Language Models (LLMs) to generate AI personas, yet their lack of deep contextual, cultural, and emotional understanding poses a significant limitation. This study quantitatively compared human responses with…

计算机与社会 · 计算机科学 2025-12-03 Tabia Tanzin Prama , Christopher M. Danforth , Peter Sheridan Dodds

Diaspora communities are disproportionately impacted by off-the-radar misinformation and often neglected by mainstream fact-checking efforts, creating a critical need to scale-up efforts of nascent fact-checking initiatives. In this paper…

信息检索 · 计算机科学 2024-05-20 Michael Shliselberg , Ashkan Kazemi , Scott A. Hale , Shiri Dori-Hacohen

Synthetic data generation has recently emerged as a promising approach for enhancing the capabilities of large language models (LLMs) without the need for expensive human annotations. However, existing methods often generate data that can…

Motivated by the remarkable progress of large language models (LLMs) in objective tasks like mathematics and coding, there is growing interest in their potential to simulate human behavior--a capability with profound implications for…

计算与语言 · 计算机科学 2026-01-23 Yuxuan Lei , Tianfu Wang , Jianxun Lian , Zhengyu Hu , Defu Lian , Xing Xie

Understanding and predicting user behavior on social media platforms is crucial for content recommendation and platform design. While existing approaches focus primarily on common actions like retweeting and liking, the prediction of rare…

计算与语言 · 计算机科学 2025-11-24 Benjamin White , Anastasia Shimorina

Personalization in LLMs often relies on costly human feedback or interaction logs, limiting scalability and neglecting deeper user attributes. To reduce the reliance on human annotations, we introduce GRAVITY (Generative Response with…

计算与语言 · 计算机科学 2025-12-11 Priyanka Dey , Daniele Rosa , Wenqing Zheng , Daniel Barcklow , Jieyu Zhao , Emilio Ferrara

What does it mean to model a person, not merely to predict isolated responses, preferences, or behaviors, but to simulate how an individual interprets events, forms opinions, makes judgments, and acts consistently across contexts? This…

计算机与社会 · 计算机科学 2026-03-31 Mao Li , Frederick G. Conrad

Online social networks have dramatically altered the landscape of public discourse, creating both opportunities for enhanced civic participation and risks of deepening social divisions. Prevalent approaches to studying online polarization…

物理与社会 · 物理学 2025-06-12 Tim Donkers , Jürgen Ziegler

The emergence of decentralized social media platforms presents new opportunities and challenges for real-time analysis of public discourse. This study introduces CognitiveSky, an open-source and scalable framework designed for sentiment,…

计算与语言 · 计算机科学 2026-05-07 Gaurab Chhetri , Anandi Dutta , Subasish Das

Developing and validating psychometric scales requires large samples, multiple testing phases, and substantial resources. Recent advances in Large Language Models (LLMs) enable the generation of synthetic participant data by prompting…

人机交互 · 计算机科学 2025-12-30 Enrico Cipriani , Pavel Okopnyi , Danilo Menicucci , Simone Grassini

Synthetic population generation is the process of combining multiple socioeconomic and demographic datasets from different sources and/or granularity levels, and downscaling them to an individual level. Although it is a fundamental step for…

机器学习 · 计算机科学 2019-11-12 Colin Wan , Zheng Li , Alicia Guo , Yue Zhao

Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with human annotators is prohibitively expensive, error-prone, and…

人工智能 · 计算机科学 2026-04-01 Tim R. Davidson , Benoit Seguin , Enrico Bacis , Cesar Ilharco , Hamza Harkous

Large Language Models (LLMs) offer promising avenues for methodological and applied innovations in survey research by using synthetic respondents to emulate human answers and behaviour, potentially mitigating measurement and representation…

计算与语言 · 计算机科学 2025-09-22 Bastián González-Bustamante , Nando Verelst , Carla Cisternas

Strict privacy regulations limit access to real transaction data, slowing open research in financial AI. Synthetic data can bridge this gap, but existing generators do not jointly achieve behavioral diversity and logical groundedness.…