English
Related papers

Related papers: LLM Predictive Scoring and Validation: Inferring E…

200 papers

This study compares the performance of AI-generated and human-written product descriptions using a multifaceted evaluation model. We analyze descriptions for 100 products generated by four AI models (Gemma 2B, LLAMA, GPT2, and ChatGPT 4)…

Computation and Language · Computer Science 2024-12-30 Sanjukta Ghosh

The rapid development and deployment of Generative AI in social settings raise important questions about how to optimally personalize them for users while maintaining accuracy and realism. Based on a Facebook public post-comment dataset,…

Computation and Language · Computer Science 2024-10-03 Kristen M. Altenburger , Hongda Jiang , Robert E. Kraut , Yi-Chia Wang , Jane Dwivedi-Yu

Classroom observation protocols standardize the assessment of teaching effectiveness and facilitate comprehension of classroom interactions. Whereas these protocols offer teachers specific feedback on their teaching practices, the manual…

Human-Computer Interaction · Computer Science 2024-07-04 Ruikun Hou , Tim Fütterer , Babette Bühler , Efe Bozkir , Peter Gerjets , Ulrich Trautwein , Enkelejda Kasneci

Human annotators frequently disagree on emotion labels, yet most evaluations of Large Language Model (LLM) emotion annotation collapse these judgments into a single gold standard, discarding the distributional information that disagreement…

Computation and Language · Computer Science 2026-05-04 Keito Inoshita , Xiaokang Zhou , Akira Kawai , Katsutoshi Yada

Online user reviews describing various products and services are now abundant on the web. While the information conveyed through review texts and ratings is easily comprehensible, there is a wealth of hidden information in them that is not…

Information Retrieval · Computer Science 2016-04-20 Rahul Kamath , Masanao Ochi , Yutaka Matsuo

This research proposes a systematic, large language model (LLM) approach for extracting product and service attributes, features, and associated sentiments from customer reviews. Grounded in marketing theory, the framework distinguishes…

Machine Learning · Statistics 2025-10-23 Khaled Boughanmi , Kamel Jedidi , Nour Jedidi

Effective personalized feedback is critical to students' literacy development. Though LLM-powered tools now promise to automate such feedback at scale, LLMs are not language-neutral: they privilege standard academic English and reproduce…

Computation and Language · Computer Science 2026-03-16 Mei Tan , Lena Phalen , Dorottya Demszky

The use of machine learning (ML) models to assess and score textual data has become increasingly pervasive in an array of contexts including natural language processing, information retrieval, search and recommendation, and credibility…

Computation and Language · Computer Science 2023-09-27 Marialena Bevilacqua , Kezia Oketch , Ruiyang Qin , Will Stamey , Xinyuan Zhang , Yi Gan , Kai Yang , Ahmed Abbasi

Automated text annotation is a compelling use case for generative large language models (LLMs) in social media research. Recent work suggests that LLMs can achieve strong performance on annotation tasks; however, these studies evaluate LLMs…

Computation and Language · Computer Science 2024-09-24 Nicholas Pangakis , Samuel Wolken

General large language models (LLMs) such as ChatGPT have shown remarkable success, but it has also raised concerns among people about the misuse of AI-generated texts. Therefore, an important question is how to detect whether the texts are…

Computation and Language · Computer Science 2023-10-24 Rongsheng Wang , Qi Li , Sihong Xie

A single digital newsletter usually contains many messages (regions). Users' reading time spent on, and read level (skip/skim/read-in-detail) of each message is important for platforms to understand their users' interests, personalize their…

Human-Computer Interaction · Computer Science 2023-06-14 Ruoyan Kong , Ruixuan Sun , Charles Chuankai Zhang , Chen Chen , Sneha Patri , Gayathri Gajjela , Joseph A. Konstan

As AI models progress beyond simple chatbots into more complex workflows, we draw ever closer to the event horizon beyond which AI systems will be utilized in autonomous, self-maintaining feedback loops. Any autonomous AI system will depend…

Artificial Intelligence · Computer Science 2026-03-06 Benjamin Feuer , Lucas Rosenblatt , Oussama Elachqar

AI-based systems such as language models have been shown to replicate and even amplify social biases reflected in their training data. Among other questionable behaviors, this can lead to AI-generated text--and text suggestions--that…

Computation and Language · Computer Science 2026-02-19 Connor Baumler , Hal Daumé

Accurately measuring consumer emotions and evaluations from unstructured text remains a core challenge for marketing research and practice. This study introduces the Linguistic eXtractor (LX), a fine-tuned, large language model trained on…

Computation and Language · Computer Science 2026-02-18 Stephan Ludwig , Peter J. Danaher , Xiaohao Yang , Yu-Ting Lin , Ehsan Abedin , Dhruv Grewal , Lan Du

Large language models (LLMs) have gained significant attention due to their ability to mimic human language. Identifying texts generated by LLMs is crucial for understanding their capabilities and mitigating potential consequences. This…

Computation and Language · Computer Science 2024-07-19 Anjali Rawal , Hui Wang , Youjia Zheng , Yu-Hsuan Lin , Shanu Sushmita

This paper investigates bias in GLLM annotations by conceptually replicating manual annotations of Boukes (2024). Using various GLLMs (Llama3.1:8b, Llama3.3:70b, GPT4o, Qwen2.5:72b) in combination with five different prompts for five…

Computation and Language · Computer Science 2025-12-10 Sjoerd B. Stolwijk , Mark Boukes , Damian Trilling

Evaluating AI-generated research ideas typically relies on LLM judges or human panels -- both subjective and disconnected from actual research impact. We introduce HindSight, a time-split evaluation framework that measures idea quality by…

Computation and Language · Computer Science 2026-03-18 Bo Jiang

Text-based measurement in political research often treats classi6ication disagreement as random noise. We examine this assumption using con6idence-weighted human annotations of 5,000 social media messages by U.S. politicians. We 6ind that…

General Economics · Economics 2026-04-27 Krishna Sharma , Khemraj Bhatt

We establish empirical bounds on behavioral inference through controlled experiments at scale: LLM-based agents assigned one of 36 behavioral profiles (9 belief systems x 4 motivations) generate over 1.5 million behavioral sequences across…

Multiagent Systems · Computer Science 2026-03-10 Jason Starace , Terence Soule

In menstrual cycle tracking apps (MCTAs), AI-based predictions and insights have become increasingly popular. These features enable users to receive personalized information about their bodies and mental states. However, there is currently…

Human-Computer Interaction · Computer Science 2026-05-14 Wendy Zhou , Pelin Karaturhan , Alexandra Weilenmann , Jichen Zhu