English
Related papers

Related papers: EXAGREE: Mitigating Explanation Disagreement with …

200 papers

Standard LLM evaluations only test capabilities or dispositions that evaluators designed them for, missing unexpected differences such as behavioral shifts between model revisions or emergent misaligned tendencies. Model diffing addresses…

Machine Learning · Computer Science 2026-02-12 Elias Kempf , Simon Schrodi , Bartosz Cywiński , Thomas Brox , Neel Nanda , Arthur Conmy

Explainable artificial intelligence provides tools to better understand predictive models and their decisions, but many such methods are limited to producing insights with respect to a single class. When generating explanations for several…

Machine Learning · Computer Science 2025-02-27 Kacper Sokol , Peter Flach

We present ACAR (Adaptive Complexity and Attribution Routing), a measurement framework for studying multi-model orchestration under auditable conditions. ACAR uses self-consistency variance (sigma) computed from N=3 probe samples to route…

Machine Learning · Computer Science 2026-02-26 Ramchand Kumaresan

User Satisfaction Estimation is an important task and increasingly being applied in goal-oriented dialogue systems to estimate whether the user is satisfied with the service. It is observed that whether the user's needs are met often…

Computation and Language · Computer Science 2024-10-15 Kaisong Song , Yangyang Kang , Jiawei Liu , Xurui Li , Changlong Sun , Xiaozhong Liu

Retrieval-augmented generation (RAG) has become a widely adopted paradigm for enabling knowledge-grounded large language models (LLMs). However, standard RAG pipelines often fail to ensure that model reasoning remains consistent with the…

Artificial Intelligence · Computer Science 2025-10-14 Jiaqi Wei , Hao Zhou , Xiang Zhang , Di Zhang , Zijie Qiu , Wei Wei , Jinzhe Li , Wanli Ouyang , Siqi Sun

The opaque nature of deep learning models presents significant challenges for the ethical deployment of hate speech detection systems. To address this limitation, we introduce Supervised Rational Attention (SRA), a framework that explicitly…

Computation and Language · Computer Science 2025-11-11 Brage Eilertsen , Røskva Bjørgfinsdóttir , Francielle Vargas , Ali Ramezani-Kebrya

As Retrieval-Augmented Generation (RAG) systems evolve toward more sophisticated architectures, ensuring their trustworthiness through explainable and robust evaluation becomes critical. Existing scalar metrics suffer from limited…

Artificial Intelligence · Computer Science 2025-12-30 Shiyan Liu , Jian Ma , Rui Qu

As machine learning models grow more complex and their applications become more high-stakes, tools for explaining model predictions have become increasingly important. This has spurred a flurry of research in model explainability and has…

Machine Learning · Computer Science 2021-11-08 Yang Liu , Sujay Khandagale , Colin White , Willie Neiswanger

Retrieval-Augmented Generation (RAG) systems for Large Language Models (LLMs) hold promise in knowledge-intensive tasks but face limitations in complex multi-step reasoning. While recent methods have integrated RAG with chain-of-thought…

Computation and Language · Computer Science 2025-01-15 Zhongxiang Sun , Qipeng Wang , Weijie Yu , Xiaoxue Zang , Kai Zheng , Jun Xu , Xiao Zhang , Song Yang , Han Li

Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of components identified via XAI offers a path toward actionable…

Artificial Intelligence · Computer Science 2026-05-27 Tobias Labarta , Maximilian Dreyer , Katharina Weitz , Wojciech Samek , Sebastian Lapuschkin

Despite significant progress in alignment, large language models (LLMs) remain vulnerable to adversarial attacks that elicit harmful behaviors. Activation steering techniques offer a promising inference-time intervention approach, but…

Machine Learning · Computer Science 2026-01-28 Quy-Anh Dang , Chris Ngo

The democratization of LLMs has accelerated the generation and circulation of highly fluent disinformation, making traditional syntax-semantic verification increasingly insufficient. Such deception rarely relies solely on surface-level…

Computation and Language · Computer Science 2026-05-27 Shang Luo , Yingguang Yang , Zhenchen Sun , Yang Liu , Bin Chong , Jingru Chen , Yancheng Chen , Jiayu Liang , Kefu Xu , Hao Peng , Philip S. Yu

With the recent proliferation of artificial intelligence systems, there has been a surge in the demand for explainability of these systems. Explanations help to reduce system opacity, support transparency, and increase stakeholder trust. In…

Software Engineering · Computer Science 2024-09-12 Umm-e-Habiba , Justus Bogner , Stefan Wagner

Analysts wishing to explore multivariate data spaces, typically pose queries involving selection operators, i.e., range or radius queries, which define data subspaces of possible interest and then use aggregation functions, the results of…

Databases · Computer Science 2020-03-13 Fotis Savva , Christos Anagnostopoulos , Peter Triantafillou

Explainable Artificial Intelligence (XAI)has received a great deal of attention recently. Explainability is being presented as a remedy for the distrust of complex and opaque models. Model agnostic methods such as LIME, SHAP, or Break Down…

Machine Learning · Computer Science 2020-05-11 Alicja Gosiewska , Przemyslaw Biecek

Stakeholders often struggle to accurately express their requirements due to articulation barriers arising from limited domain knowledge or from cognitive constraints. This can cause misalignment between expressed and intended requirements,…

Software Engineering · Computer Science 2026-01-26 Michael Mircea , Emre Gevrek , Elisa Schmid , Kurt Schneider

In many real world applications of machine learning, models have to meet certain domain-based requirements that can be expressed as constraints (e.g., safety-critical constraints in autonomous driving systems). Such constraints are often…

Machine Learning · Computer Science 2022-06-20 Kshitij Goyal , Sebastijan Dumancic , Hendrik Blockeel

Machine learning models are primarily judged by predictive performance, especially in applied settings. Once a model reaches high accuracy, its explanation is often assumed to be correct and trustworthy. This assumption raises an overlooked…

Machine Learning · Computer Science 2026-02-12 Chama Bensmail

Explainable AI (XAI) methods have mostly been built to investigate and shed light on single machine learning models and are not designed to capture and explain differences between multiple models effectively. This paper addresses the…

Machine Learning · Computer Science 2025-04-01 Adam Rida , Marie-Jeanne Lesot , Xavier Renard , Christophe Marsala

Counterfactual Explanations (CFEs) interpret machine learning models by identifying the smallest change to input features needed to change the model's prediction to a desired output. For classification tasks, CFEs determine how close a…

Machine Learning · Computer Science 2025-10-01 Margarita A. Guerrero , Cristian R. Rojas