中文
相关论文

相关论文: Crisis-Bench: Benchmarking Strategic Ambiguity and…

200 篇论文

While Large Language Models (LLMs) demonstrate significant potential in providing accessible mental health support, their practical deployment raises critical trustworthiness concerns due to the domains high-stakes and safety-sensitive…

计算与语言 · 计算机科学 2026-03-04 Zixin Xiong , Ziteng Wang , Haotian Fan , Xinjie Zhang , Wenxuan Wang

Psychological support hotlines serve as critical lifelines for crisis intervention but encounter significant challenges due to rising demand and limited resources. Large language models (LLMs) offer potential support in crisis assessments,…

计算与语言 · 计算机科学 2025-12-19 Guifeng Deng , Shuyin Rao , Tianyu Lin , Anlu Dai , Pan Wang , Junyi Xie , Haidong Song , Ke Zhao , Dongwu Xu , Zhengdong Cheng , Tao Li , Haiteng Jiang

Prompt design significantly impacts the moral competence and safety alignment of large language models (LLMs), yet empirical comparisons remain fragmented across datasets and models.We introduce ProMoral-Bench, a unified benchmark…

The ability of large language models (LLMs) to manage and acquire economic resources remains unclear. In this paper, we introduce \textbf{Market-Bench}, a comprehensive benchmark that evaluates the capabilities of LLMs in…

人工智能 · 计算机科学 2026-04-21 Yushuo Zheng , Huiyu Duan , Zicheng Zhang , Yucheng Zhu , Xiongkuo Min , Guangtao Zhai

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have emerged as powerful tools for a myriad of applications, from natural language processing to decision-making support systems. However, as these…

计算与语言 · 计算机科学 2025-07-08 Jianchao Ji , Yutong Chen , Mingyu Jin , Wujiang Xu , Wenyue Hua , Yongfeng Zhang

Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions underexplored. This gap is particularly critical in mental health, where patient questions…

计算与语言 · 计算机科学 2026-05-15 Yahan Li , Jifan Yao , John Bosco S. Bunyi , Adam C. Frank , Angel Hsing-Chi Hwang , Ruishan Liu

Large language models (LLMs) exhibit advancing capabilities in complex tasks, such as reasoning and graduate-level question answering, yet their resilience against misuse, particularly involving scientifically sophisticated risks, remains…

The mismatch between the growing demand for psychological counseling and the limited availability of services has motivated research into the application of Large Language Models (LLMs) in this domain. Consequently, there is a need for a…

计算与语言 · 计算机科学 2025-11-13 Bichen Wang , Yixin Sun , Junzhe Wang , Hao Yang , Xing Fu , Yanyan Zhao , Si Wei , Shijin Wang , Bing Qin

Large language models (LLMs) show strong potential for simulating human social behaviors and interactions, yet lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes. To…

社会与信息网络 · 计算机科学 2025-10-14 Jia Wang , Ziyu Zhao , Tingjuntao Ni , Zhongyu Wei

People often encounter role conflicts -- social dilemmas where the expectations of multiple roles clash and cannot be simultaneously fulfilled. As large language models (LLMs) increasingly navigate these social dynamics, a critical research…

计算与语言 · 计算机科学 2026-04-20 Jisu Shin , Hoyun Song , Juhyun Oh , Changgeon Ko , Eunsu Kim , Chani Jung , Alice Oh

Large language models (LLMs) are increasingly deployed as conversational assistants in open-domain, multi-turn settings, where users often provide incomplete or ambiguous information. However, existing LLM-focused clarification benchmarks…

计算与语言 · 计算机科学 2025-12-25 Sichun Luo , Yi Huang , Mukai Li , Shichang Meng , Fengyuan Liu , Zefa Hu , Junlan Feng , Qi Liu

We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters from natural language strategy descriptions and market…

计算与语言 · 计算机科学 2026-01-22 Abhay Srivastava , Sam Jung , Spencer Mateega

Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human behaviors. Current evaluations of simulation fidelity are…

计算与语言 · 计算机科学 2026-04-14 Tiancheng Hu , Joachim Baumann , Lorenzo Lupo , Nigel Collier , Dirk Hovy , Paul Röttger

Large language models (LLMs) excel at natural language tasks but remain brittle in domains requiring precise logical and symbolic reasoning. Chaotic dynamical systems provide an especially demanding test because chaos is deterministic yet…

人工智能 · 计算机科学 2026-02-13 Noel Thomas

As large language models (LLMs) evolve from conversational assistants into autonomous agents, evaluating the safety of their actions becomes critical. Prior safety benchmarks have primarily focused on preventing generation of harmful…

计算与语言 · 计算机科学 2026-03-04 Adi Simhi , Jonathan Herzig , Martin Tutek , Itay Itzhak , Idan Szpektor , Yonatan Belinkov

Evaluating the safety alignment of LLM responses in high-risk mental health dialogues is particularly difficult due to missing gold-standard answers and the ethically sensitive nature of these interactions. To address this challenge, we…

计算与语言 · 计算机科学 2026-02-16 Yunna Cai , Fan Wang , Haowei Wang , Kun Wang , Kailai Yang , Sophia Ananiadou , Moyan Li , Mingming Fan

Evaluating Large Language Models (LLMs) for mental health support is challenging due to the emotionally and cognitively complex nature of therapeutic dialogue. Existing benchmarks are limited in scale, reliability, often relying on…

We introduce POLIS-Bench, the first rigorous, systematic evaluation suite designed for LLMs operating in governmental bilingual policy scenarios. Compared to existing benchmarks, POLIS-Bench introduces three major advancements. (i)…

计算与语言 · 计算机科学 2025-11-10 Tingyue Yang , Junchi Yao , Yuhui Guo , Chang Liu

Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal activities or unethical behavior, posing serious compliance risks. To systematically…

计算与语言 · 计算机科学 2026-05-04 Yutao Hou , Yihan Jiang , Yuhan Xie , Jian Yang , Liwen Zhang , Hailiang Huang , Guanhua Chen , Yun Chen

Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hold significant potential to enhance decision making, coordination, and operational…

人工智能 · 计算机科学 2026-05-04 Sydney Johns , Heng Jin , Chaoyu Zhang , Y. Thomas Hou , Wenjing Lou
‹ 上一页 1 2 3 10 下一页 ›