English
Related papers

Related papers: LLM Predictive Scoring and Validation: Inferring E…

200 papers

An earlier paper (Hong, Potteiger, and Zapata 2026) established that an unoptimized GPT 4.1 prompt predicts fan-reported experience ratings within one point 67% of the time from open-ended survey text. This paper tests the relative impact…

Computation and Language · Computer Science 2026-04-22 Andrew Hong , Jason Potteiger , Luis E. Zapata

Appraisal theories suggest that emotions arise from subjective evaluations of events, referred to as appraisals. The taxonomy of appraisals is quite diverse, and they are usually given ratings on a Likert scale to be annotated in an…

Computation and Language · Computer Science 2025-03-25 Deniss Ruder , Andero Uusberg , Kairit Sirts

The variations between in-group and out-group speech (intergroup bias) are subtle and could underlie many social phenomena like stereotype perpetuation and implicit bias. In this paper, we model the intergroup bias as a tagging task on…

Computation and Language · Computer Science 2025-11-10 Venkata S Govindarajan , Matianyu Zang , Kyle Mahowald , David Beaver , Junyi Jessy Li

Large language models (LLMs) can serve as judges that offer rapid and reliable assessments of other LLM outputs. However, models may systematically assign overly favorable ratings to their own outputs, a phenomenon known as self-bias, which…

Computation and Language · Computer Science 2025-08-12 Evangelia Spiliopoulou , Riccardo Fogliato , Hanna Burnsky , Tamer Soliman , Jie Ma , Graham Horwood , Miguel Ballesteros

Authors often struggle to interpret peer review feedback, deriving false hope from polite comments or feeling confused by specific low scores. To investigate this, we construct a dataset of over 30,000 ICLR 2021-2025 submissions and compare…

Computation and Language · Computer Science 2026-04-17 Yingxuan Wen

Interleaving is an online evaluation approach for information retrieval systems that compares the effectiveness of ranking functions in interpreting the users' implicit feedback. Previous work such as Hofmann et al (2011) has evaluated the…

Information Retrieval · Computer Science 2023-03-20 Alessandro Benedetti , Anna Ruggero

Recent research on explainable recommendation generally frames the task as a standard text generation problem, and evaluates models simply based on the textual similarity between the predicted and ground-truth explanations. However, this…

Evaluating developer satisfaction with conversational AI assistants at scale is critical but challenging. User studies provide rich insights, but are unscalable, while large-scale quantitative signals from logs or in-product ratings are…

Software Engineering · Computer Science 2025-09-24 Daye Nam , Malgorzata Salawa , Satish Chandra

Student responses in STEM assessments are often handwritten and combine symbolic expressions, calculations, and diagrams, creating substantial variation in format and interpretation. Despite their importance for evaluating students'…

Artificial Intelligence · Computer Science 2026-04-15 Xiuxiu Tang , G. Alex Ambrose , Ying Cheng

Social aspects of software projects become increasingly important for research and practice. Different approaches analyze the sentiment of a development team, ranging from simply asking the team to so-called sentiment analysis on text-based…

Software Engineering · Computer Science 2022-07-19 Marc Herrmann , Martin Obaidi , Larissa Chazette , Jil Klünder

Proactive AI writing assistants need to predict when users want drafting help, yet we lack empirical understanding of what drives preferences. Through a factorial vignette study with 50 participants making 750 pairwise comparisons, we find…

Computation and Language · Computer Science 2026-01-09 Vivian Lai , Zana Buçinca , Nil-Jana Akpinar , Mo Houtti , Hyeonsu B. Kang , Kevin Chian , Namjoon Suh , Alex C. Williams

Popularity systems, like Twitter retweets, Reddit upvotes, and Pinterest pins have the potential to guide people toward posts that others liked. That, however, creates a feedback loop that reduces their informativeness: items marked as more…

Human-Computer Interaction · Computer Science 2018-09-05 Maria Glenski , Greg Stoddard , Paul Resnick , Tim Weninger

Automatic analysis of user reviews to understand user sentiments toward app functionality (i.e. app features) helps align development efforts with user expectations and needs. Recent advances in Large Language Models (LLMs) such as ChatGPT…

Computation and Language · Computer Science 2025-02-11 Faiz Ali Shah , Ahmed Sabir , Rajesh Sharma , Dietmar Pfahl

User sentiment on social media reveals the underlying social trends, crises, and needs. Researchers have analyzed users' past messages to trace the evolution of sentiments and reconstruct sentiment dynamics. However, predicting the imminent…

Computation and Language · Computer Science 2025-12-25 Fanhang Man , Huandong Wang , Jianjie Fang , Zhaoyi Deng , Baining Zhao , Xinlei Chen , Yong Li

Due to their architecture and vast pre-training data, large language models (LLMs) demonstrate strong text classification performance. However, LLM output - here, the category assigned to a text - depends heavily on the wording of the…

Computation and Language · Computer Science 2025-12-04 Kylie L. Anglin , Stephanie Milan , Brittney Hernandez , Claudia Ventura

Emotions exert an immense influence over human behavior and cognition in both commonplace and high-stress tasks. Discussions of whether or how to integrate large language models (LLMs) into everyday life (e.g., acting as proxies for, or…

Artificial Intelligence · Computer Science 2025-08-21 Mattson Ogg , Chace Ashcraft , Ritwik Bose , Raphael Norman-Tenazas , Michael Wolmetz

In this work, we establish a baseline potential for how modern model-generated text explanations of movie recommendations may help users, and explore what different components of these text explanations that users like or dislike,…

Artificial Intelligence · Computer Science 2023-09-19 Joyce Zhou , Thorsten Joachims

Ask ChatGPT about vacation planning, and it may infer your income. Ask it about medication, and it may infer your medical history. Because such inferences can expose more information than users intend to reveal, prior work argues that they…

Human-Computer Interaction · Computer Science 2026-05-12 Kyzyl Monteiro , Minjung Park , Alexander Ioffrida , Angelina Sanna , Hao-Ping , Lee , Niloofar Mireshghallah , Yang Wang , Sauvik Das

In the post-Turing era, evaluating large language models (LLMs) involves assessing generated text based on readers' reactions rather than merely its indistinguishability from human-produced content. This paper explores how LLM-generated…

Computational Engineering, Finance, and Science · Computer Science 2024-11-26 Takehiro Takayanagi , Hiroya Takamura , Kiyoshi Izumi , Chung-Chi Chen

Can out-of-the-box pretrained Large Language Models (LLMs) detect human affect successfully when observing a video? To address this question, for the first time, we evaluate comprehensively the capacity of popular LLMs for successfully…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 David Melhart , Matthew Barthet , Georgios N. Yannakakis
‹ Prev 1 2 3 10 Next ›