中文
相关论文

相关论文: Beyond Human Judgment: A Bayesian Evaluation of LL…

200 篇论文

Effective and safe human-machine collaboration requires the regulated and meaningful exchange of emotions between humans and artificial intelligence (AI). Current AI systems based on large language models (LLMs) can provide feedback that…

计算与语言 · 计算机科学 2025-06-18 Xiuwen Wu , Hao Wang , Zhiang Yan , Xiaohan Tang , Pengfei Xu , Wai-Ting Siok , Ping Li , Jia-Hong Gao , Bingjiang Lyu , Lang Qin

This study examines the ethical reasoning of six prominent generative large language models: OpenAI GPT-4o, Meta LLaMA 3.1, Perplexity, Anthropic Claude 3.5 Sonnet, Google Gemini, and Mistral 7B. The research explores how these models…

人工智能 · 计算机科学 2025-01-16 W. Russell Neuman , Chad Coleman , Manan Shah

When LLM-based multi-agent systems disagree, current practice treats this as noise to be resolved through consensus. We propose it can be signal. We focus on hate speech moderation, a domain where judgments depend on cultural context and…

多智能体系统 · 计算机科学 2026-04-07 Michał Wawer , Jarosław A. Chudziak

Systematic reviews are crucial for synthesizing scientific evidence but remain labor-intensive, especially when extracting detailed methodological information. Large language models (LLMs) offer potential for automating methodological…

计算与语言 · 计算机科学 2025-10-14 Wenqing Zhang , Trang Nguyen , Elizabeth A. Stuart , Yiqun T. Chen

With the emergence of large language models (LLMs), investigating if they can surpass humans in areas such as emotion recognition and empathetic responding has become a focal point of research. This paper presents a comprehensive study…

计算与语言 · 计算机科学 2024-06-10 Anuradha Welivita , Pearl Pu

With the rapid development and uptake of large language models (LLMs) across high-stakes settings, it is increasingly important to ensure that LLMs behave in ways that align with human values. Existing moral benchmarks prompt LLMs with…

With the recent advances in Artificial Intelligence (AI) and Large Language Models (LLMs), the automation of daily tasks, like automatic writing, is getting more and more attention. Hence, efforts have focused on aligning LLMs with human…

计算与语言 · 计算机科学 2025-06-09 Mohammadamin Shafiei , Hamidreza Saffari

Making moral judgments is an essential step toward developing ethical AI systems. Prevalent approaches are mostly implemented in a bottom-up manner, which uses a large set of annotated data to train models based on crowd-sourced opinions…

计算与语言 · 计算机科学 2024-07-02 Jingyan Zhou , Minda Hu , Junan Li , Xiaoying Zhang , Xixin Wu , Irwin King , Helen Meng

This paper explores the integration of human-like emotions and ethical considerations into Large Language Models (LLMs). We first model eight fundamental human emotions, presented as opposing pairs, and employ collaborative LLMs to…

计算与语言 · 计算机科学 2024-06-26 Edward Y. Chang

Span annotation - annotating specific text features at the span level - can be used to evaluate texts where single-score metrics fail to provide actionable feedback. Until recently, span annotation was done by human annotators or fine-tuned…

Despite their widespread use in fact-checking, moderation, and high-stakes decision-making, large language models (LLMs) remain poorly understood as judges of truth. This study presents the largest evaluation to date of LLMs' veracity…

计算与语言 · 计算机科学 2025-09-30 Emilio Barkett , Olivia Long , Madhavendra Thakur

As general-purpose artificial intelligence systems become increasingly integrated into society and are used for information seeking, content generation, problem solving, textual analysis, coding, and running processes, it is crucial to…

计算机与社会 · 计算机科学 2025-08-28 Ljubisa Bojic , Dylan Seychell , Milan Cabarkapa

In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same task. We also explore if fine-tuning several LLMs on the…

人工智能 · 计算机科学 2024-11-26 Artem Karpov , Seong Hah Cho , Austin Meek , Raymond Koopmanschap , Lucy Farnik , Bogdan-Ionut Cirstea

The growing complexity and diversity of news coverage have made framing analysis a crucial yet challenging task in computational social science. Traditional approaches, including manual annotation and fine-tuned models, remain limited by…

计算与语言 · 计算机科学 2026-05-22 Valeria Pastorino , Jasivan A. Sivakumar , Nafise Sadat Moosavi

Large language models (LLMs) have been proposed as alternatives to human experts for estimating unknown quantities with associated uncertainty, a process known as Bayesian elicitation. We test this by asking eleven LLMs to estimate…

人工智能 · 计算机科学 2026-04-03 Luka Hobor , Mario Brcic , Mihael Kovac , Kristijan Poje

Mediation is a dispute resolution method featuring a neutral third-party (mediator) who intervenes to help the individuals resolve their dispute. In this paper, we investigate to which extent large language models (LLMs) are able to act as…

If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations, leading to less reliable benchmarks for…

计算与语言 · 计算机科学 2025-07-17 Joe Needham , Giles Edkins , Govind Pimpale , Henning Bartsch , Marius Hobbhahn

Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating large language models (LLMs). However, there are still many…

Large language models (LLMs) can support democratic deliberation at scales previously constrained by turn-taking and facilitation bandwidth. Recent work shows that LLM-generated group statements are often preferred over human-mediated…

计算与语言 · 计算机科学 2026-05-27 Wajdi Zaghouani