English
Related papers

Related papers: Detection Without Correction: A Two-Parameter Deco…

200 papers

Retrieval Augmented Generation (RAG) improves correctness of Question Answering (QA) and addresses hallucinations in Large Language Models (LLMs), yet greatly increase computational costs. Besides, RAG is not always needed as may introduce…

Understanding how latent representations evolve during generation is a central open problem in large language model interpretability. We introduce \textbf{Dynamical Manifold Evolution Theory} (DMET), a phenomenological framework that models…

Computation and Language · Computer Science 2026-05-05 Yukun Zhang , Qi Dong , Mengkang Li

This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed $\textbf{Unsolvable Problem Detection (UPD)}$. Multiple-choice question answering (MCQA) is widely used to assess…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Atsuyuki Miyai , Jingkang Yang , Jingyang Zhang , Yifei Ming , Qing Yu , Go Irie , Yixuan Li , Hai Li , Ziwei Liu , Kiyoharu Aizawa

We employ a matrix-based solver for the linear rheology of fluid-immersed disordered spring networks to reveal four distinct dynamic response regimes. One regime - completely absent in the known vacuum response - exhibits coupled fluid flow…

Soft Condensed Matter · Physics 2019-12-04 David Head , Cornelis Storm

Democratic discourse analysis systems increasingly rely on multi-agent LLM pipelines in which distinct evaluator models are assigned adversarial roles to generate structured, multi-perspective assessments of political statements. A core…

Artificial Intelligence · Computer Science 2026-05-01 Juergen Dietrich

This paper focuses on the system identification of an important class of nonlinear systems: linearly parameterized nonlinear systems, which enjoys wide applications in robotics and other mechanical systems. We consider two system…

Systems and Control · Electrical Eng. & Systems 2024-11-22 Negin Musavi , Ziyao Guo , Geir Dullerud , Yingying Li

Current reasoning paradigms for LLMs include chain-of-thought, ReAct, and post-hoc self-critique. These paradigms rely on two assumptions that fail on long-horizon, multi-stage tasks. As a result, errors accumulate silently across reasoning…

Artificial Intelligence · Computer Science 2026-05-08 Fan Huang

Ensuring data quality in cloud-native Extract-Load-Transform (ELT) pipelines is increasingly challenging due to heterogeneous data sources, evolving schemas, and multi-backend execution environments. This paper presents a unified,…

Software Engineering · Computer Science 2026-05-21 Ismail Gargouri , Hassan Reza

Multi-agent debates have been introduced to improve the accuracy of Large Language Models (LLMs) by having multiple agents discuss solutions to a problem over several rounds of debate. However, models often generate incorrect yet…

Computation and Language · Computer Science 2025-02-25 Luke Yoffe , Alfonso Amayuelas , William Yang Wang

Speculative decoding accelerates large language model inference by using smaller draft models to generate candidate tokens for parallel verification. However, current approaches are limited by sequential stage dependencies that prevent full…

Artificial Intelligence · Computer Science 2025-05-06 Bradley McDanel , Sai Qian Zhang , Yunhai Hu , Zining Liu

Humans typically use natural language to update teammates on task states. Since not all updates are communicated, discrepancies arise between the team members' mental models that negatively affect overall team performance. How can we…

Artificial Intelligence · Computer Science 2026-05-06 Katharine Kowalyshyn , Matthias Scheutz

Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and acting on needs the user has implied but not said. We call this \emph{conversational…

Machine Learning · Computer Science 2026-05-12 Sepehr Harfi , Ahmad Salimi , Dongming Shen , Alex Smola

Large Language Models (LLMs) are increasingly instantiated as interacting agents in multi-agent systems (MAS), where collective decisions emerge through social interaction rather than independent reasoning. A fundamental yet underexplored…

Multiagent Systems · Computer Science 2026-01-12 Chen Han , Jin Tan , Bohan Yu , Wenzhen Zheng , Xijin Tang

Retrieval-Augmented Generation (RAG) integrates large language models (LLMs) with external sources, but unresolved contradictions in retrieved evidence often lead to hallucinations and legally unsound outputs. Benchmarks currently used for…

Artificial Intelligence · Computer Science 2025-10-14 Ananya Mantravadi , Shivali Dalmia , Olga Pospelova , Abhishek Mukherji , Nand Dave , Anudha Mittal

Large language models (LLMs) increasingly operate as autonomous agents that reason over external APIs to perform complex tasks. However, their reliability and agreement remain poorly characterized. We present a unified benchmarking…

Information Retrieval · Computer Science 2026-04-28 Eyhab Al-Masri

Large language models (LLMs) often generate fluent but factually incorrect outputs, known as hallucinations, which undermine their reliability in real-world applications. While uncertainty estimation has emerged as a promising strategy for…

Machine Learning · Computer Science 2025-05-13 Pei-Fu Guo , Yun-Da Tsai , Shou-De Lin

Fine-tuning LLMs on narrowly harmful datasets can lead to behavior that is broadly misaligned with respect to human values. To understand when and how this emergent misalignment occurs, we develop a comprehensive framework for detecting and…

Machine Learning · Computer Science 2025-08-28 Julian Arnold , Niels Lörch

LLM-as-a-Judge has become the dominant paradigm for evaluating language model outputs, yet LLM judges exhibit systematic biases that compromise evaluation reliability. We present a comprehensive empirical study comparing nine debiasing…

Artificial Intelligence · Computer Science 2026-04-28 Sadman Kabir Soumik

Large Language Models (LLMs) have shown great potential in Natural Language Processing (NLP) tasks. However, recent literature reveals that LLMs generate nonfactual responses intermittently, which impedes the LLMs' reliability for further…

Computation and Language · Computer Science 2024-03-22 Yukun Zhao , Lingyong Yan , Weiwei Sun , Guoliang Xing , Chong Meng , Shuaiqiang Wang , Zhicong Cheng , Zhaochun Ren , Dawei Yin

Large Language Models (LLMs) demonstrate potential in complex legal tasks like argument generation, yet their reliability remains a concern. Building upon pilot work assessing LLM generation of 3-ply legal arguments using human evaluation,…

Computation and Language · Computer Science 2025-06-04 Li Zhang , Morgan Gray , Jaromir Savelka , Kevin D. Ashley