Artificial Intelligence · Computer Science
Reasoning Models Better Express Their Confidence
Dongkeun Yoon, Seungone Kim, Sohee Yang, Sunkyoung Kim +5
2025-10-23
Computation and Language · Computer Science
Thinking Out Loud: Do Reasoning Models Know When They're Right?
Qingcheng Zeng, Weihao Xuan, Leyang Cui, Rob Voigt
2025-10-21
Artificial Intelligence · Computer Science
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
Shu Zhou, Rui Ling, Junan Chen, Xin Wang +2
2026-04-14
Computation and Language · Computer Science
Increasing the Thinking Budget is Not All You Need
Ignacio Iacobacci, Zhaozhi Qian, Faroq AL-Tam, Muhammad AL-Qurishi +1
2025-12-23
Computation and Language · Computer Science
Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
Jiannan Guan, Qiguang Chen, Libo Qin, Dengyun Peng +4
2025-12-02
Machine Learning · Computer Science
Thought calibration: Efficient and confident test-time scaling
Menghua Wu, Cai Zhou, Stephen Bates, Tommi Jaakkola
2025-05-27
Artificial Intelligence · Computer Science
An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint
Yi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu +7
2025-05-22
Computation and Language · Computer Science
Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision
Chaeyun Jang, Moonseok Choi, Yegon Kim, Hyungi Lee +1
2025-06-05
Computation and Language · Computer Science
Identifying Influential N-grams in Confidence Calibration via Regression Analysis
Shintaro Ozaki, Wataru Hashimoto, Hidetaka Kamigaito, Katsuhiko Hayashi +1
2026-04-08
Computation and Language · Computer Science
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
Junlin Wang, Siddhartha Jain, Dejiao Zhang, Baishakhi Ray +2
2024-06-18
Artificial Intelligence · Computer Science
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
Jingchu Gai, Guanning Zeng, Christina Baek, Chen Wu +3
2026-05-26
Computation and Language · Computer Science
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models
Gaurav Srivastava, Aafiya Hussain, Sriram Srinivasan, Xuan Wang
2026-04-24
Computation and Language · Computer Science
On Calibration of Large Language Models: From Response To Capability
Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin, Yun-Nung Chen +2
2026-02-17
Artificial Intelligence · Computer Science
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton +2
2025-11-21
Machine Learning · Computer Science
What Large Language Models Know and What People Think They Know
Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem +4
2025-02-14
Artificial Intelligence · Computer Science
Bayesian Elicitation with LLMs: Model Size Helps, Extra "Reasoning" Doesn't Always
Luka Hobor, Mario Brcic, Mihael Kovac, Kristijan Poje
2026-04-03
Computation and Language · Computer Science
Calibrating Long-form Generations from Large Language Models
Yukun Huang, Yixin Liu, Raghuveer Thirukovalluru, Arman Cohan +1
2024-10-29
Computation and Language · Computer Science
Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
Liangjie Huang, Dawei Li, Huan Liu, Lu Cheng
2025-04-07
Human-Computer Interaction · Computer Science
Seeing the Reasoning: How LLM Rationales Influence User Trust and Decision-Making in Factual Verification Tasks
Xin Sun, Shu Wei, Jos A Bosch, Isao Echizen +2
2026-03-10
Computation and Language · Computer Science
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li +3
2024-03-19
Computation and Language · Computer Science
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
Tairan Fu, Javier Conde, Gonzalo Martínez, María Grandury +1
2026-05-05
Computation and Language · Computer Science
A Survey of Confidence Estimation and Calibration in Large Language Models
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl +2
2024-03-26