Computation and Language · Computer Science
The Glass Ceiling of Automatic Evaluation in Natural Language Generation
Pierre Colombo, Maxime Peyrard, Nathan Noiry, Robert West +1
2022-10-10
Computation and Language · Computer Science
Dynamic Human Evaluation for Relative Model Comparisons
Thórhildur Thorleiksdóttir, Cedric Renggli, Nora Hollenstein, Ce Zhang
2022-04-29
Computation and Language · Computer Science
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
Patrícia Schmidtová, Saad Mahamood, Simone Balloccu, Ondřej Dušek +5
2024-08-20
Computation and Language · Computer Science
Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation
Zhaokun Jiang, Qianxi Lv, Ziyin Zhang, Lei Lei
2024-10-15
Human-Computer Interaction · Computer Science
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge
Aparna Elangovan, Lei Xu, Jongwoo Ko, Mahsa Elyasi +3
2025-01-28
Computation and Language · Computer Science
Correction of Errors in Preference Ratings from Automated Metrics for Text Generation
Jan Deriu, Pius von Däniken, Don Tuggener, Mark Cieliebak
2023-06-07
Computation and Language · Computer Science
NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist
Iftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola Pechenizkiy
2023-05-29
Computation and Language · Computer Science
Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation
Cristina Garbacea, Samuel Carton, Shiyan Yan, Qiaozhu Mei
2019-09-09
Computation and Language · Computer Science
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
Xuansheng Wu, Padmaja Pravin Saraf, Gyeonggeon Lee, Ehsan Latif +2
2025-02-24
Computation and Language · Computer Science
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara
2025-08-22
Audio and Speech Processing · Electrical Eng. & Systems
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Huan Zhang, Jinhua Liang, Huy Phan, Wenwu Wang +1
2025-05-01
Artificial Intelligence · Computer Science
Comparing Human and Automated Evaluation of Open-Ended Student Responses to Questions of Evolution
Michael J Wiser, Louise S Mead, James J Smith, Robert T Pennock
2018-05-08