中文
相关论文

相关论文: EvalCards: A Framework for Standardized Evaluation…

200 篇论文

Peer review is our best tool for judging the quality of conference submissions, but it is becoming increasingly spurious. We argue that a part of the problem is that the reviewers and area chairs face a poorly defined task forcing…

计算与语言 · 计算机科学 2020-10-09 Anna Rogers , Isabelle Augenstein

In the NLP community, recent years have seen a surge of research activities that address machines' ability to perform deep language understanding which goes beyond what is explicitly stated in text, rather relying on reasoning and knowledge…

计算与语言 · 计算机科学 2020-02-27 Shane Storks , Qiaozi Gao , Joyce Y. Chai

Scientific progress in NLP rests on the reproducibility of researchers' claims. The *CL conferences created the NLP Reproducibility Checklist in 2020 to be completed by authors at submission to remind them of key information to include. We…

计算与语言 · 计算机科学 2023-06-19 Ian Magnusson , Noah A. Smith , Jesse Dodge

Recent natural language processing (NLP) techniques have accomplished high performance on benchmark datasets, primarily due to the significant improvement in the performance of deep learning. The advances in the research community have led…

计算与语言 · 计算机科学 2022-10-24 Marwan Omar , Soohyeon Choi , DaeHun Nyang , David Mohaisen

A key part of the NLP ethics movement is responsible use of data, but exactly what that means or how it can be best achieved remain unclear. This position paper discusses the core legal and ethical principles for collection and sharing of…

计算与语言 · 计算机科学 2021-09-15 Anna Rogers , Tim Baldwin , Kobi Leins

Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element ensuring reliable and reproducible human assessment, has…

计算与语言 · 计算机科学 2024-06-13 Jie Ruan , Wenqing Wang , Xiaojun Wan

In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and…

Financial disclosure analysis and Knowledge extraction is an important financial analysis problem. Prevailing methods depend predominantly on quantitative ratios and techniques, which suffer from limitations like window dressing and past…

交易与市场微观结构 · 定量金融 2021-01-13 Sridhar Ravula

The development of LLM agents has led to a growing body of work on knowledge-work AI, including coding, research, and healthcare. However, current knowledge-work evaluation and benchmark design still largely follow the logic of traditional…

人工智能 · 计算机科学 2026-05-25 Yining Hua , Hongbin Na , Cyrus Ayubcha , Levi Lian

Research funding allocation remains a critical bottleneck in scientific advancement, yet the review process for funding proposals lacks the transparency that has revolutionized academic paper peer review. Traditional funding agencies…

计算机与社会 · 计算机科学 2025-12-17 Sakshi Ahuja , Subhankar Mishra

Neural networks are a prevalent and effective machine learning component, and their application is leading to significant scientific progress in many domains. As the field of neural network systems is fast growing, it is important to…

人机交互 · 计算机科学 2022-11-22 Guy Clarke Marshall , André Freitas , Caroline Jay

In this paper we present an exploratory research on quantifying the impact that data distribution has on the performance and evaluation of NLP models. We propose an automated framework that measures the data point distribution across 6…

计算与语言 · 计算机科学 2024-04-02 Venelin Kovatchev , Matthew Lease

Objectives: This paper presents a brief review on Aadhaar card, and discusses the scope and advantages of linking Aadhaar card to various systems. Further we present various cases in which Aadhaar card may pose security threats. The…

计算机与社会 · 计算机科学 2017-08-18 Raja Siddharth Raju , Sukhdev Singh , Kiran Khatter

Despite the rising popularity of saliency-based explanations, the research community remains at an impasse, facing doubts concerning their purpose, efficacy, and tendency to contradict each other. Seeking to unite the community's efforts…

计算与语言 · 计算机科学 2023-08-29 Jennifer Hsia , Danish Pruthi , Aarti Singh , Zachary C. Lipton

A popular approach to unveiling the black box of neural NLP models is to leverage saliency methods, which assign scalar importance scores to each input component. A common practice for evaluating whether an interpretability method is…

计算与语言 · 计算机科学 2023-05-12 Josip Jukić , Martin Tutek , Jan Šnajder

Sustainability reports are critical for ESG assessment, yet greenwashing and vague claims often undermine their reliability. Existing NLP models lack robustness to these practices, typically relying on surface-level patterns that generalize…

计算与语言 · 计算机科学 2026-01-30 Neil Heinrich Braun , Keane Ong , Rui Mao , Erik Cambria , Gianmarco Mengaldo

Effective evaluation of language models remains an open challenge in NLP. Researchers and engineers face methodological issues such as the sensitivity of models to evaluation setup, difficulty of proper comparisons across methods, and the…

Leaderboard systems allow researchers to objectively evaluate Natural Language Processing (NLP) models and are typically used to identify models that exhibit superior performance on a given task in a predetermined setting. However, we argue…

计算与语言 · 计算机科学 2023-03-21 Chanjun Park , Hyeonseok Moon , Seolhwa Lee , Jaehyung Seo , Sugyeong Eo , Heuiseok Lim

The rapid integration of conversational AI systems into educational settings has intensified ethical concerns about academic integrity, fairness, and students' cognitive development. Institutional responses have largely centered on AI…

计算机与社会 · 计算机科学 2026-03-10 Eduardo Davalos , Yike Zhang

Autonomous and semi-autonomous systems are using deep learning models to improve decision-making. However, deep classifiers can be overly confident in their incorrect predictions, a major issue especially in safety-critical domains. The…

机器学习 · 计算机科学 2024-12-05 Murat Sensoy , Lance M. Kaplan , Simon Julier , Maryam Saleki , Federico Cerutti