中文
相关论文

相关论文: EvalCards: A Framework for Standardized Evaluation…

200 篇论文

The rapid advancement and impressive capabilities of large language models (LLMs) have given rise to the field of prompt engineering, the practice of crafting inputs to guide LLMs toward high-quality, task-relevant outputs. A critical…

计算机与社会 · 计算机科学 2026-03-16 Amandine M. Caut , Beimnet Zenebe , Amy Rouillard , David J. T. Sumpter

The increasing complexity of software systems and the influence of software-supported decisions in our society have sparked the need for software that is safe, reliable, and fair. Explainability has been identified as a means to achieve…

软件工程 · 计算机科学 2022-09-02 Timo Speith

Recently, there have been increasing calls for computer science curricula to complement existing technical training with topics related to Fairness, Accountability, Transparency, and Ethics. In this paper, we present Value Card, an…

计算机与社会 · 计算机科学 2023-01-11 Hong Shen , Wesley Hanwen Deng , Aditi Chattopadhyay , Zhiwei Steven Wu , Xu Wang , Haiyi Zhu

We introduce Dynaboard, an evaluation-as-a-service framework for hosting benchmarks and conducting holistic model comparison, integrated with the Dynabench platform. Our platform evaluates NLP models directly instead of relying on…

计算与语言 · 计算机科学 2021-06-14 Zhiyi Ma , Kawin Ethayarajh , Tristan Thrush , Somya Jain , Ledell Wu , Robin Jia , Christopher Potts , Adina Williams , Douwe Kiela

State-of-the-art models in NLP are now predominantly based on deep neural networks that are opaque in terms of how they come to make predictions. This limitation has increased interest in designing more interpretable deep models for NLP…

计算与语言 · 计算机科学 2020-04-27 Jay DeYoung , Sarthak Jain , Nazneen Fatema Rajani , Eric Lehman , Caiming Xiong , Richard Socher , Byron C. Wallace

The pursuit of leaderboard rankings in Large Language Models (LLMs) has created a fundamental paradox: models excel at standardized tests while failing to demonstrate genuine language understanding and adaptability. Our systematic analysis…

计算与语言 · 计算机科学 2024-12-06 Sourav Banerjee , Ayushi Agarwal , Eishkaran Singh

The documentation practice for machine-learned (ML) models often falls short of established practices for traditional software, which impedes model accountability and inadvertently abets inappropriate or misuse of models. Recently, model…

软件工程 · 计算机科学 2023-02-10 Avinash Bhat , Austin Coursey , Grace Hu , Sixian Li , Nadia Nahar , Shurui Zhou , Christian Kästner , Jin L. C. Guo

Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit…

计算机与社会 · 计算机科学 2026-05-22 Naveen Raman , Santiago Cortes-Gomez , Mateo Dulce Rubio , Fei Fang , Bryan Wilder

The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- which benefit from structured documentation frameworks like…

机器学习 · 计算机科学 2025-12-04 Florian Bordes , Candace Ross , Justine T Kao , Evangelia Spiliopoulou , Adina Williams

Evaluation is the central means for assessing, understanding, and communicating about NLP models. In this position paper, we argue evaluation should be more than that: it is a force for driving change, carrying a sociological and political…

计算与语言 · 计算机科学 2022-12-23 Rishi Bommasani

As research and industry moves towards large-scale models capable of numerous downstream tasks, the complexity of understanding multi-modal datasets that give nuance to models rapidly increases. A clear and thorough understanding of a…

人机交互 · 计算机科学 2022-04-05 Mahima Pushkarna , Andrew Zaldivar , Oddur Kjartansson

By design, large language models (LLMs) are static general-purpose models, expensive to retrain or update frequently. As they are increasingly adopted for knowledge-intensive tasks, it becomes evident that these design choices lead to…

计算与语言 · 计算机科学 2024-03-25 Shangbin Feng , Weijia Shi , Yuyang Bai , Vidhisha Balachandran , Tianxing He , Yulia Tsvetkov

Benchmarking is seen as critical to assessing progress in NLP. However, creating a benchmark involves many design decisions (e.g., which datasets to include, which metrics to use) that often rely on tacit, untested assumptions about what…

计算与语言 · 计算机科学 2024-06-14 Yu Lu Liu , Su Lin Blodgett , Jackie Chi Kit Cheung , Q. Vera Liao , Alexandra Olteanu , Ziang Xiao

Questions of fairness, robustness, and transparency are paramount to address before deploying NLP systems. Central to these concerns is the question of reliability: Can NLP systems reliably treat different demographics fairly and function…

机器学习 · 计算机科学 2021-06-02 Samson Tan , Shafiq Joty , Kathy Baxter , Araz Taeihagh , Gregory A. Bennett , Min-Yen Kan

Developing documentation guidelines and easy-to-use templates for datasets and models is a challenging task, especially given the variety of backgrounds, skills, and incentives of the people involved in the building of natural language…

In NLP, models are usually evaluated by reporting single-number performance scores on a number of readily available benchmarks, without much deeper analysis. Here, we argue that - especially given the well-known fact that benchmarks often…

计算与语言 · 计算机科学 2022-10-05 Daniel Simig , Tianlu Wang , Verna Dankers , Peter Henderson , Khuyagbaatar Batsuren , Dieuwke Hupkes , Mona Diab

Financial question answering (QA) over long corporate filings requires evidence to satisfy strict constraints on entities, financial metrics, fiscal periods, and numeric values. However, existing LLM-based rerankers primarily optimize…

信息检索 · 计算机科学 2026-05-01 Yixi Zhou , Fan Zhang , Yu Chen , Haipeng Zhang , Preslav Nakov , Zhuohan Xie

In 2019, the paper entitled "Model Cards for Model Reporting" introduced a new tool for documenting model performance and encouraged the practice of transparent reporting for a defined list of categories. One of the categories detailed in…

计算机与社会 · 计算机科学 2024-03-26 DeBrae Kennedy-Mayo , Jake Gord

To address a looming crisis of unreproducible evaluation for named entity recognition, we propose guidelines and introduce SeqScore, a software package to improve reproducibility. The guidelines we propose are extremely simple and center…

计算与语言 · 计算机科学 2021-11-08 Chester Palen-Michel , Nolan Holley , Constantine Lignos

Objectives: To evaluate the current limitations of large language models (LLMs) in medical question answering, focusing on the quality of datasets used for their evaluation. Materials and Methods: Widely-used benchmark datasets, including…

计算与语言 · 计算机科学 2025-07-15 Mahmoud Alwakeel , Aditya Nagori , Vijay Krishnamoorthy , Rishikesan Kamaleswaran