中文
相关论文

相关论文: NoisyCausal: A Benchmark for Evaluating Causal Rea…

200 篇论文

Current speech-language models (SLMs) typically use a cascade of speech encoder and large language model, treating speech understanding as a single black box. They analyze the content of speech well but reason weakly about other aspects,…

音频与语音处理 · 电气工程与系统科学 2025-12-08 Xuanru Zhou , Jiachen Lian , Henry Hong , Xinyi Yang , Gopala Anumanchipalli

We introduce seqBench, a parametrized benchmark for probing sequential reasoning limits in Large Language Models (LLMs) through precise, multi-dimensional control over several key complexity dimensions. seqBench allows systematic variation…

人工智能 · 计算机科学 2025-09-23 Mohammad Ramezanali , Mo Vazifeh , Paolo Santi

Genuine human-like causal reasoning is fundamental for strong artificial intelligence. Humans typically identify whether an event is part of the causal chain first, and then influenced by modulatory factors such as morality, normality, and…

计算与语言 · 计算机科学 2025-10-21 Yanxi Zhang , Xin Cong , Zhong Zhang , Xiao Liu , Dongyan Zhao , Yesai Wu

The widespread adoption of large language models (LLMs) in healthcare raises critical questions about their ability to interpret patient-generated narratives, which are often informal, ambiguous, and noisy. Existing benchmarks typically…

计算与语言 · 计算机科学 2025-09-16 Eden Mama , Liel Sheri , Yehudit Aperstein , Alexander Apartsin

Learning from noisy labels (LNL) is a challenge that arises in many real-world scenarios where collected training data can contain incorrect or corrupted labels. Most existing solutions identify noisy labels and adopt active learning to…

机器学习 · 计算机科学 2025-04-07 Bo Yuan , Yulin Chen , Yin Zhang , Wei Jiang

Uncovering the mechanisms behind "jailbreaks" in large language models (LLMs) is crucial for enhancing their safety and reliability, yet these mechanisms remain poorly understood. Existing studies predominantly analyze jailbreak prompts by…

机器学习 · 计算机科学 2026-02-06 Licheng Pan , Yunsheng Lu , Jiexi Liu , Jialing Tao , Haozhe Feng , Hui Xue , Zhixuan Chu , Kui Ren

Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less attention has been paid to their ability to simulate human…

计算工程、金融与科学 · 计算机科学 2025-08-25 Yuanjun Feng , Vivek Choudhary , Yash Raj Shrestha

Large Language Models (LLMs) excel at reasoning, traditionally requiring high-quality large-scale data and extensive training. Recent works reveal a very appealing Less-Is-More phenomenon where very small, carefully curated high-quality…

机器学习 · 计算机科学 2026-04-22 Rapheal Huang , Weilong Guo

Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored. We present LlaMADRS, a benchmark for structured clinical assessment from dialogue built on the…

Objective: This study investigates the potential of Large Language Models (LLMs) as an alternative to human expert elicitation for extracting structured causal knowledge and facilitating causal modeling in biometric and healthcare…

人工智能 · 计算机科学 2025-04-15 Olha Shaposhnyk , Daria Zahorska , Svetlana Yanushkevich

Causal discovery for dynamical systems poses a major challenge in fields where active interventions are infeasible. Most methods used to investigate these systems and their associated benchmarks are tailored to deterministic,…

机器学习 · 计算机科学 2025-10-13 Benjamin Herdeanu , Juan Nathaniel , Carla Roesch , Jatan Buch , Gregor Ramien , Johannes Haux , Pierre Gentine

Despite remarkable advances in the field, LLMs remain unreliable in distinguishing causation from correlation. Recent results from the Corr2Cause dataset benchmark reveal that state-of-the-art LLMs -- such as GPT-4 (F1 score: 29.08) -- only…

人工智能 · 计算机科学 2025-05-28 Wentao Sun , João Paulo Nogueira , Alonso Silva

Implicit Sentiment Analysis (ISA) aims to infer sentiment that is implied rather than explicitly stated, requiring models to perform deeper reasoning over subtle contextual cues. While recent prompting-based methods using Large Language…

计算与语言 · 计算机科学 2025-07-02 Jing Ren , Wenhao Zhou , Bowen Li , Mujie Liu , Nguyen Linh Dan Le , Jiade Cen , Liping Chen , Ziqi Xu , Xiwei Xu , Xiaodong Li

Uncovering causal relationships is a fundamental problem across science and engineering. However, most existing causal discovery methods assume acyclicity and direct access to the system variables -- assumptions that fail to hold in many…

机器学习 · 计算机科学 2026-03-24 Muralikrishnna G. Sethuraman , Faramarz Fekri

The effective utilization of structured data, integral to corporate data strategies, has been challenged by the rise of large language models (LLMs) capable of processing unstructured information. This shift prompts the question: can LLMs…

计算与语言 · 计算机科学 2024-10-22 Zhouhong Gu , Haoning Ye , Xingzhou Chen , Zeyang Zhou , Hongwei Feng , Yanghua Xiao

Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To address this, we introduce Behavior Cue Reasoning for making LLM reasoning more controllable…

人工智能 · 计算机科学 2026-05-21 Christopher Z. Cui , Taylor W. Killian , Prithviraj Ammanabrolu

Smart buildings generate vast streams of sensor and control data, but facility managers often lack clear explanations for anomalous energy usage. We propose InsightBuild, a two-stage framework that integrates causality analysis with a…

机器学习 · 计算机科学 2025-07-14 Pinaki Prasad Guha Neogi , Ahmad Mohammadshirazi , Rajiv Ramnath

This PhD thesis contains several contributions to the field of statistical causal modeling. Statistical causal models are statistical models embedded with causal assumptions that allow for the inference and reasoning about the behavior of…

机器学习 · 统计学 2021-10-05 Martin Emil Jakobsen

Understanding commonsense causality is a unique mark of intelligence for humans. It helps people understand the principles of the real world better and benefits the decision-making process related to causation. For instance, commonsense…

计算与语言 · 计算机科学 2024-08-30 Shaobo Cui , Zhijing Jin , Bernhard Schölkopf , Boi Faltings

Identifying cause-and-effect relationships is critical to understanding real-world dynamics and ultimately causal reasoning. Existing methods for identifying event causality in NLP, including those based on Large Language Models (LLMs),…

人工智能 · 计算机科学 2025-02-13 Vy Vo , Lizhen Qu , Tao Feng , Yuncheng Hua , Xiaoxi Kang , Songhai Fan , Tim Dwyer , Lay-Ki Soon , Gholamreza Haffari
‹ 上一页 1 8 9 10 下一页 ›