Computation and Language · Computer Science
Human Judgement as a Compass to Navigate Automatic Metrics for Formality Transfer
Huiyuan Lai, Jiali Mao, Antonio Toral, Malvina Nissim
2022-04-18
Software Engineering · Computer Science
The role of formalism in system requirements (full version)
Jean-Michel Bruel, Sophie Ebersold, Florian Galinier, Alexandr Naumchev +2
2020-04-17
Computation and Language · Computer Science
Emotion Ratings: How Intensity, Annotation Confidence and Agreements are Entangled
Enrica Troiano, Sebastian Padó, Roman Klinger
2021-03-03
Computation and Language · Computer Science
"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm +1
2022-11-10
Computation and Language · Computer Science
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer
Eleftheria Briakou, Sweta Agrawal, Joel Tetreault, Marine Carpuat
2021-10-22
Computation and Language · Computer Science
Investigating the Nature of Disagreements on Mid-Scale Ratings: A Case Study on the Abstractness-Concreteness Continuum
Urban Knupleš, Diego Frassinelli, Sabine Schulte im Walde
2024-04-18
Computation and Language · Computer Science
FormalAlign: Automated Alignment Evaluation for Autoformalization
Jianqiao Lu, Yingjia Wan, Yinya Huang, Jing Xiong +2
2024-10-15
Machine Learning · Computer Science
Instance-level Randomization: Toward More Stable LLM Evaluations
Yiyang Li, Yonghuang Wu, Ying Luo, Liangtai Sun +4
2025-09-17
Computation and Language · Computer Science
Review-Level Sentiment Classification with Sentence-Level Polarity Correction
Sylvester Olubolu Orimaye, Saadat M. Alhashmi, Eu-Gene Siew, Sang Jung Kang
2015-11-10
Computation and Language · Computer Science
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
Weiyue Li, Minda Zhao, Weixuan Dong, Jiahui Cai +11
2026-01-08
Computers and Society · Computer Science
Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education
Andrea Gaggioli, Giuseppe Casaburi, Leonardo Ercolani, Francesco Collova' +2
2025-08-05
Computation and Language · Computer Science
Coherency through formalisations of Structured Natural Language, A case study on FRETish
Joost J. Joosten, Marina López Chamosa, Sofía Santiago Fernández
2026-05-12