中文
相关论文

相关论文: Predicting Human Choice Between Textually Describe…

200 篇论文

Questionnaires are a common method for detecting the personality of Large Language Models (LLMs). However, their reliability is often compromised by two main issues: hallucinations (where LLMs produce inaccurate or irrelevant responses) and…

计算与语言 · 计算机科学 2024-10-14 Baohua Zhan , Yongyi Huang , Wenyao Cui , Huaping Zhang , Jianyun Shang

What makes large language models (LLMs) impressive is also what makes them hard to evaluate: their diversity of uses. To evaluate these models, we must understand the purposes they will be used for. We consider a setting where these…

计算与语言 · 计算机科学 2024-06-04 Keyon Vafa , Ashesh Rambachan , Sendhil Mullainathan

Thematic analysis provides valuable insights into participants' experiences through coding and theme development, but its resource-intensive nature limits its use in large healthcare studies. Large language models (LLMs) can analyze text at…

Long-sequence decision-making, which is usually addressed through reinforcement learning (RL), is a critical component for optimizing strategic operations in dynamic environments, such as real-time bidding in computational advertising. The…

Humans act via a nuanced process that depends both on rational deliberation and also on identity and contextual factors. In this work, we study how large language models (LLMs) can simulate human action in the context of social dilemma…

计算与语言 · 计算机科学 2026-02-03 Suhong Moon , Minwoo Kang , Joseph Suh , Mustafa Safdari , John Canny

Measuring the generalization ability of Large Language Models (LLMs) is challenging due to data contamination. As models grow and computation becomes cheaper, ensuring tasks and test cases are unseen during training phases will become…

计算与语言 · 计算机科学 2025-07-09 Sougata Saha , Monojit Choudhury

Reliable simulation of human behavior is essential for explaining, predicting, and intervening in our society. Recent advances in large language models (LLMs) have shown promise in emulating human behaviors, interactions, and…

计算与语言 · 计算机科学 2025-10-27 Ning Bian , Xianpei Han , Hongyu Lin , Baolei Wu , Jun Wang

In the post-Turing era, evaluating large language models (LLMs) involves assessing generated text based on readers' reactions rather than merely its indistinguishability from human-produced content. This paper explores how LLM-generated…

计算工程、金融与科学 · 计算机科学 2024-11-26 Takehiro Takayanagi , Hiroya Takamura , Kiyoshi Izumi , Chung-Chi Chen

This study compares the efficacy of GPT-4 and clinalytix Medical AI in predicting the clinical risk of delirium development. Findings indicate that GPT-4 exhibited significant deficiencies in identifying positive cases and struggled to…

计算与语言 · 计算机科学 2024-09-17 Mohamed Rezk , Patricia Cabanillas Silva , Fried-Michael Dahlweid

Large language models (LLMs) show increasingly advanced emergent capabilities and are being incorporated across various societal domains. Understanding their behavior and reasoning abilities therefore holds significant importance. We argue…

A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, non-interactive question-and-answer benchmarks. However, theoretical work in the field suggests that…

计算与语言 · 计算机科学 2026-02-20 Jared Moore , Rasmus Overmark , Ned Cooper , Beba Cibralic , Nick Haber , Cameron R. Jones

We evaluate large language models (LLMs) for automatic personality prediction from text under the binary Five Factor Model (BIG5). Five models -- including GPT-4 and lightweight open-source alternatives -- are tested across three…

计算与语言 · 计算机科学 2025-12-01 Francesco Di Cursi , Chiara Boldrini , Marco Conti , Andrea Passarella

Large Language Models (LLMs) have transformed artificial intelligence by excelling in complex natural language processing tasks. Their ability to generate human-like text has opened new possibilities for market research, particularly in…

人工智能 · 计算机科学 2026-04-20 Mengxin Wang , Dennis J. Zhang , Heng Zhang

A key challenge in transportation planning is that the collective preferences of heterogeneous travelers often diverge from the policies produced by model-driven decision tools. This misalignment frequently results in implementation delays…

计算机与社会 · 计算机科学 2025-10-29 Xiaoyu Yan , Tianxing Dai , Yu Marco Nie

The recent performance leap of Large Language Models (LLMs) opens up new opportunities across numerous industrial applications and domains. However, erroneous generations, such as false predictions, misinformation, and hallucination made by…

软件工程 · 计算机科学 2025-01-07 Yuheng Huang , Jiayang Song , Zhijie Wang , Shengming Zhao , Huaming Chen , Felix Juefei-Xu , Lei Ma

Do generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions? If so, how can these biases be mitigated? Drawing on the cognitive psychology and experimental…

综合经济学 · 经济学 2026-02-11 Pietro Bini , Lin William Cong , Xing Huang , Lawrence J. Jin

Understanding the limits of language is a prerequisite for Large Language Models (LLMs) to act as theories of natural language. LLM performance in some language tasks presents both quantitative and qualitative differences from that of…

计算与语言 · 计算机科学 2025-06-30 Vittoria Dentella , Fritz Guenther , Evelina Leivada

Large Language Models are being used in conversational agents that simulate human conversations and generate social studies data. While concerns about the models' biases have been raised and discussed in the literature, much about the data…

计算机与社会 · 计算机科学 2025-10-24 Guido Ivetta , Laura Moradbakhti , Rafael A. Calvo

In recent years, Large Language Models (LLMs) have become fundamental to a broad spectrum of artificial intelligence applications. As the use of LLMs expands, precisely estimating the uncertainty in their predictions has become crucial.…

Unlike traditional time series, the action sequences of human decision making usually involve many cognitive processes such as beliefs, desires, intentions, and theory of mind, i.e., what others are thinking. This makes predicting human…

机器学习 · 计算机科学 2022-06-07 Baihan Lin , Djallel Bouneffouf , Guillermo Cecchi
‹ 上一页 1 8 9 10 下一页 ›