English
Related papers

Related papers: Prefix-Safe Bayesian Belief Tracking for LLM Reaso…

200 papers

This paper develops a new approach to post-selection inference for screening high-dimensional predictors of survival outcomes. Post-selection inference for right-censored outcome data has been investigated in the literature, but much…

Methodology · Statistics 2021-12-22 Tzu-Jung Huang , Alex Luedtke , Ian W. McKeague

Background. Systematic reviews in comparative effectiveness research require timely evidence synthesis. Preprints accelerate knowledge dissemination but vary in quality, posing challenges for systematic reviews. Methods. We propose…

Computation and Language · Computer Science 2025-07-14 Rui Yang , Jiayi Tong , Haoyuan Wang , Hui Huang , Ziyang Hu , Peiyu Li , Nan Liu , Christopher J. Lindsell , Michael J. Pencina , Yong Chen , Chuan Hong

Language model (LM) "reasoning", commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics underlying this process remain poorly understood. We study these dynamics through the lens…

The advent of large language models (LLMs) has dramatically advanced the state-of-the-art in numerous natural language generation tasks. For LLMs to be applied reliably, it is essential to have an accurate measure of their confidence.…

Computation and Language · Computer Science 2024-06-05 Zhen Lin , Shubhendu Trivedi , Jimeng Sun

Large language models (LLMs) inherently operate over a large generation space, yet conventional usage typically reports the most likely generation (MLG) as a point prediction, which underestimates the model's capability: although the…

Computation and Language · Computer Science 2026-03-25 Ye Li , Anqi Hu , Yuanchang Ye , Shiyan Tong , Zhiyuan Wang , Bo Fu

Large Deep Learning models are often compressed before being deployed in a resource-constrained environment. Can we trust the prediction of compressed models just as we trust the prediction of the original large model? Existing work has…

Computation and Language · Computer Science 2025-08-20 Rohit Raj Rai , Chirag Kothari , Siddhesh Shelke , Amit Awekar

Suffix-based jailbreak attacks append an adversarial suffix, i.e., a short token sequence, to steer aligned LLMs into unsafe outputs. Since suffixes are free-form text, they admit endlessly many surface forms, making jailbreak mitigation…

Cryptography and Security · Computer Science 2026-02-09 Mengyao Du , Han Fang , Haokai Ma , Gang Yang , Quanjun Yin , Shouling Ji , Ee-Chien Chang

Recently, e-learning platforms have grown as a place where students can post doubts (as a snap taken with smart phones) and get them resolved in minutes. However, the significant increase in the number of student-posted doubts with high…

Machine Learning · Computer Science 2022-08-23 Vedant Sandeep Joshi , Sivanagaraja Tatinati , Yubo Wang

Large language models (LLMs) achieve higher accuracy on challenging reasoning tasks by scaling test-time compute through multiple trajectory sampling. However, standard aggregation methods like majority voting or individual confidence-based…

Machine Learning · Computer Science 2026-02-04 Yingchuan Zhang , Terry Ma , Wenxuan Zhong , Ping Ma

With hundreds of multilingual embedding models available, practitioners lack clear guidance on which provide genuine cross-lingual semantic alignment versus task performance through language-specific patterns. Task-driven benchmarks (MTEB)…

Computation and Language · Computer Science 2026-01-16 Wen G. Gong

Classifying sequential data as early and as accurately as possible is a challenging yet critical problem, especially when a sampling cost is high. One algorithm that achieves this goal is the sequential probability ratio test (SPRT), which…

Machine Learning · Computer Science 2021-02-09 Akinori F. Ebihara , Taiki Miyagawa , Kazuyuki Sakurai , Hitoshi Imaoka

Pass$@k$ is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especially when the number of trials (samples) is limited and computational resources are constrained.…

Artificial Intelligence · Computer Science 2026-05-13 Mohsen Hariri , Amirhossein Samandar , Michael Hinczewski , Vipin Chaudhary

We study the source of uncertainty in DeepSeek R1-32B by analyzing its self-reported verbal confidence on question answering (QA) tasks. In the default answer-then-confidence setting, the model is regularly over-confident, whereas semantic…

Computation and Language · Computer Science 2025-11-06 Jakub Podolak , Rajeev Verma

Although pretrained language models (PTLMs) contain significant amounts of world knowledge, they can still produce inconsistent answers to questions when probed, even after specialized training. As a result, it can be hard to identify what…

Computation and Language · Computer Science 2021-10-01 Nora Kassner , Oyvind Tafjord , Hinrich Schütze , Peter Clark

Large Language Models (LLMs) remain vulnerable to adaptive jailbreaks that easily bypass empirical defenses like GCG. We propose a framework for certifiable robustness that shifts safety guarantees from single-pass inference to the…

Computation and Language · Computer Science 2026-02-03 Zehua Cheng , Jianwei Yang , Wei Dai , Jiahao Sun

LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence signal carries item-level information before building on it. We transfer the validity…

Computation and Language · Computer Science 2026-04-21 Jon-Paul Cacioli

Bayesian inference provides a natural framework for updating knowledge as new information becomes available, often in a sequential manner by incorporating datasets in stages or reusing previous posteriors as priors. In practice, this is…

Nuclear Theory · Physics 2026-05-22 Lipei Du

Reasoning has emerged as the next major frontier for language models (LMs), with rapid advances from both academic and industrial labs. However, this progress often outpaces methodological rigor, with many evaluations relying on…

Machine Learning · Computer Science 2025-10-08 Andreas Hochlehnert , Hardik Bhatnagar , Vishaal Udandarao , Samuel Albanie , Ameya Prabhu , Matthias Bethge

Automated Essay Scoring (AES) systems now reach near human agreement on some public benchmarks, yet real-world adoption, especially in high-stakes examinations, remains limited. A principal obstacle is that most models output a single score…

Computation and Language · Computer Science 2025-09-22 Ahmed Karim , Qiao Wang , Zheng Yuan

Large language models (LLMs) achieve strong average performance yet remain unreliable at the instance level, with frequent hallucinations, brittle failures, and poorly calibrated confidence. We study reliability through the lens of…

Artificial Intelligence · Computer Science 2026-01-13 Pranav Kallem
‹ Prev 1 4 5 6 7 8 10 Next ›