English
Related papers

Related papers: LLM Predictive Scoring and Validation: Inferring E…

200 papers

Generative AI systems have become ubiquitous for all kinds of modalities, which makes the issue of the evaluation of such models more pressing. One popular approach is preference ratings, where the generated outputs of different systems are…

Artificial Intelligence · Computer Science 2024-06-04 Pius von Däniken , Jan Deriu , Don Tuggener , Mark Cieliebak

We explore automatically predicting which Wordle games Reddit users find amusing. We scrape approximately 80k reactions by Reddit users to Wordle games from Reddit, classify the reactions as expressing amusement or not using OpenAI's…

Computation and Language · Computer Science 2025-06-09 Ronaldo Luo , Gary Liang , Cindy Liu , Adam Kabbara , Minahil Bakhtawar , Kina Kim , Michael Guerzhoy

LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made. Yet standard confidence intervals ignore variability…

Computation and Language · Computer Science 2026-05-14 Solomon Messing

Large Language Models (LLMs) have been garnering significant attention of AI researchers, especially following the widespread popularity of ChatGPT. However, due to LLMs' intricate architecture and vast parameters, several concerns and…

Software Engineering · Computer Science 2023-10-10 Tinghui Ouyang , Hoang-Quoc Nguyen-Son , Huy H. Nguyen , Isao Echizen , Yoshiki Seo

In this paper, we focus on online reviews and employ artificial intelligence tools, taken from the cognitive computing field, to help understanding the relationships between the textual part of the review and the assigned numerical score.…

Computation and Language · Computer Science 2017-07-24 Michela Fazzolari , Vittoria Cozza , Marinella Petrocchi , Angelo Spognardi

Sensitive attributes are legally protected characteristics that should not be used to discriminate. Careful steps have been taken to minimize the risk of human bias regarding these fields, such as race and age. Large language models (LLMs)…

Computers and Society · Computer Science 2026-04-14 Anay Agarwalla , Simeon Sayer

User-generated texts such as reviews and social media are valuable sources of information. Online reviews are important assets for users to buy a product, see a movie, or make a decision. Therefore, rating of a review is one of the reliable…

Computation and Language · Computer Science 2018-05-23 Mohammadamir Kavousi , Sepehr Saadatmand

Large Language Models (LLMs) have demonstrated exceptional capabilities in generalizing to new tasks in a zero-shot or few-shot manner. However, the extent to which LLMs can comprehend user preferences based on their previous behavior…

Information Retrieval · Computer Science 2023-05-12 Wang-Cheng Kang , Jianmo Ni , Nikhil Mehta , Maheswaran Sathiamoorthy , Lichan Hong , Ed Chi , Derek Zhiyuan Cheng

Accommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acting as writing agents to infer a description of user…

Computation and Language · Computer Science 2025-06-02 Stéphane Aroca-Ouellette , Natalie Mackraz , Barry-John Theobald , Katherine Metcalf

Summary assessment involves evaluating how well a generated summary reflects the key ideas and meaning of the source text, requiring a deep understanding of the content. Large Language Models (LLMs) have been used to automate this process,…

Computation and Language · Computer Science 2025-12-23 Zahra Sadeghi , Evangelos Milios , Frank Rudzicz

Automated UI evaluation can be beneficial for the design process; for example, to compare different UI designs, or conduct automated heuristic evaluation. LLM-based UI evaluation, in particular, holds the promise of generalizability to a…

Human-Computer Interaction · Computer Science 2024-08-15 Peitong Duan , Chin-yi Chen , Gang Li , Bjoern Hartmann , Yang Li

Sentiment Analysis is an important algorithm in Natural Language Processing which is used to detect sentiment within some text. In our project, we had chosen to work on analyzing reviews of various drugs which have been reviewed in form of…

Computation and Language · Computer Science 2020-03-27 Sairamvinay Vijayaraghavan , Debraj Basu

We address the problem of scaling up the production of media content, including commentary and personalized news stories, for large-scale sports and music events worldwide. Our approach relies on generative AI models to transform a large…

Computation and Language · Computer Science 2024-02-29 Aaron Baughman , Stephen Hammer , Rahul Agarwal , Gozde Akay , Eduardo Morales , Tony Johnson , Leonid Karlinsky , Rogerio Feris

Strategic model selection and reasoning settings are more effective than ensembling for optimizing automated scoring with large language models (LLMs). We examined self-consistency (intra-model majority voting) and reasoning effort for…

Computers and Society · Computer Science 2026-05-01 Scott Frohn

LLMs have so far failed both to generate consistently compelling stories and to recognize this failure--on the leading creative-writing benchmark (EQ-Bench), LLM judges rank zero-shot AI stories above New Yorker short stories, a gold…

Computation and Language · Computer Science 2026-04-14 Peiqi Sui , Yutong Zhu , Tianyi Cheng , Peter West , Richard Jean So , Hoyt Long , Ari Holtzman

After the launch of ChatGPT v.4 there has been a global vivid discussion on the ability of this artificial intelligence powered platform and some other similar ones for the automatic production of all kinds of texts, including scientific…

Computation and Language · Computer Science 2024-04-16 Javier J. Sanchez-Medina

Preference learning is critical for aligning large language models (LLMs) with human values, with the quality of preference datasets playing a crucial role in this process. While existing metrics primarily assess data quality based on…

Machine Learning · Computer Science 2025-03-05 Kexin Huang , Junkang Wu , Ziqian Chen , Xue Wang , Jinyang Gao , Bolin Ding , Jiancan Wu , Xiangnan He , Xiang Wang

Personality traits are richly encoded in natural language, and large language models (LLMs) trained on human text can simulate personality when conditioned on persona descriptions. However, existing evaluations rely predominantly on…

Computation and Language · Computer Science 2026-04-08 Ben Wigler , Maria Tsfasman , Tiffany Matej Hrkalovic

Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational…

Generative AI (GenAI) is increasingly used in survey contexts to simulate human preferences. While many research endeavors evaluate the quality of synthetic GenAI data by comparing model-generated responses to gold-standard survey results,…

Machine Learning · Computer Science 2025-02-25 Sarah Ball , Simeon Allmendinger , Frauke Kreuter , Niklas Kühl
‹ Prev 1 3 4 5 6 7 10 Next ›