Computation and Language · Computer Science
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
Yixin Cao, Shibo Hong, Xinze Li, Jiahao Ying +23
2025-04-29
Computation and Language · Computer Science
A Survey on Enhancing Causal Reasoning Ability of Large Language Models
Xin Li, Zhuo Cai, Shoujin Wang, Kun Yu +1
2025-03-13
Computation and Language · Computer Science
Evaluation Hallucination in Multi-Round Incomplete Information Lateral-Driven Reasoning Tasks
Wenhan Dong, Tianyi Hu, Jingyi Zheng, Zhen Sun +4
2025-06-02
Computers and Society · Computer Science
Large Language Models as Partners in Student Essay Evaluation
Toru Ishida, Tongxi Liu, Hailong Wang, William K. Cheung
2024-05-30
Computation and Language · Computer Science
Case Study: Testing Model Capabilities in Some Reasoning Tasks
Min Zhang, Sato Takumi, Jack Zhang, Jun Wang
2024-02-16
Artificial Intelligence · Computer Science
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
Marco AF Pimentel, Clément Christophe, Tathagata Raha, Prateek Munjal +2
2024-08-01
Computation and Language · Computer Science
Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models
Gonzalo Martínez, Javier Conde, Elena Merino-Gómez, Beatriz Bermúdez-Margaretto +3
2024-01-30
Artificial Intelligence · Computer Science
Algorithmic Thinking Theory
MohammadHossein Bateni, Vincent Cohen-Addad, Yuzhou Gu, Silvio Lattanzi +2
2025-12-05
Computation and Language · Computer Science
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman +9
2024-10-04
Artificial Intelligence · Computer Science
Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity
Erez Yosef, Oron Anschel, Shunit Haviv Hakimi, Asaf Gendler +3
2026-04-27
Machine Learning · Computer Science
Can Large Language Models Learn Formal Logic? A Data-Driven Training and Evaluation Framework
Yuan Xia, Akanksha Atrey, Fadoua Khmaissia, Kedar S. Namjoshi
2025-04-30
Machine Learning · Computer Science
Large Language Model Enhanced Machine Learning Estimators for Classification
Yuhang Wu, Yingfei Wang, Chu Wang, Zeyu Zheng
2024-05-10
Computation and Language · Computer Science
How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison
Jiayin Wang, Zhiquang Guo, Weizhi Ma, Min Zhang
2025-08-07
Computation and Language · Computer Science
A Principled Framework for Knowledge-enhanced Large Language Model
Saizhuo Wang, Zhihan Liu, Zhaoran Wang, Jian Guo
2023-11-21
Computation and Language · Computer Science
Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges
Xiao Xiao, Yu Su, Sijing Zhang, Zhang Chen +2
2025-05-01
Computation and Language · Computer Science
DnA-Eval: Enhancing Large Language Model Evaluation through Decomposition and Aggregation
Minzhi Li, Zhengyuan Liu, Shumin Deng, Shafiq Joty +2
2024-12-10
Computation and Language · Computer Science
From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
Andreas Stephan, Dawei Zhu, Matthias Aßenmacher, Xiaoyu Shen +1
2025-05-14
Computation and Language · Computer Science
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin +3
2024-11-07