中文
相关论文

相关论文: Developing and Maintaining an Open-Source Reposito…

200 篇论文

As AI systems appear to exhibit ever-increasing capability and generality, assessing their true potential and safety becomes paramount. This paper contends that the prevalent evaluation methods for these systems are fundamentally…

人工智能 · 计算机科学 2024-07-15 John Burden

As AI systems advance and integrate into society, well-designed and transparent evaluations are becoming essential tools in AI governance, informing decisions by providing evidence about system capabilities and risks. Yet there remains a…

The development of conversational AI assistants is an iterative process with multiple components. As such, the evaluation and continual improvement of these assistants is a complex and multifaceted problem. This paper introduces the…

Artificial Intelligence systems, which benefit from the availability of large-scale datasets and increasing computational power, have become effective solutions to various critical tasks, such as natural language understanding, speech…

软件工程 · 计算机科学 2023-03-20 Zhou Yang , Chenyu Wang , Jieke Shi , Thong Hoang , Pavneet Kochhar , Qinghua Lu , Zhenchang Xing , David Lo

The advent of advanced AI underscores the urgent need for comprehensive safety evaluations, necessitating collaboration across communities (i.e., AI, software engineering, and governance). However, divergent practices and terminologies…

软件工程 · 计算机科学 2024-05-17 Boming Xia , Qinghua Lu , Liming Zhu , Zhenchang Xing

To evaluate software maintenance techniques and tools in controlled experiments with human participants, researchers currently use projects and tasks selected on an ad-hoc basis. This can unrealistically favor their tool, and it makes the…

软件工程 · 计算机科学 2020-12-01 Matúš Sulír

Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an…

Audits are critical mechanisms for identifying the risks and limitations of deployed artificial intelligence (AI) systems. However, the effective execution of AI audits remains incredibly difficult, and practitioners often need to make use…

计算机与社会 · 计算机科学 2025-03-03 Victor Ojewale , Ryan Steed , Briana Vecchione , Abeba Birhane , Inioluwa Deborah Raji

Open-source software (OSS) is foundational to modern digital infrastructure, yet this context for group work continues to struggle to ensure sufficient contributions in many critical cases. This literature review explores how artificial…

软件工程 · 计算机科学 2026-02-10 S M Rakib UI Karim , Wenyi Lu , Sean Goggins

At the heart of improving conversational AI is the open problem of how to evaluate conversations. Issues with automatic metrics are well known (Liu et al., 2016, arXiv:1603.08023), with human evaluations still considered the gold standard.…

计算与语言 · 计算机科学 2022-01-14 Eric Michael Smith , Orion Hsu , Rebecca Qian , Stephen Roller , Y-Lan Boureau , Jason Weston

The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- which benefit from structured documentation frameworks like…

机器学习 · 计算机科学 2025-12-04 Florian Bordes , Candace Ross , Justine T Kao , Evangelia Spiliopoulou , Adina Williams

As generative large model capabilities advance, safety concerns become more pronounced in their outputs. To ensure the sustainable growth of the AI ecosystem, it's imperative to undertake a holistic evaluation and refinement of associated…

人工智能 · 计算机科学 2023-12-01 Jiawen Deng , Jiale Cheng , Hao Sun , Zhexin Zhang , Minlie Huang

The external evaluation of AI systems is increasingly recognised as a crucial approach for understanding their potential risks. However, facilitating external evaluation in practice faces significant challenges in balancing evaluators' need…

计算机与社会 · 计算机科学 2025-03-04 Ben Bucknall , Robert F. Trager , Michael A. Osborne

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can…

计算机与社会 · 计算机科学 2024-12-13 Peter Barnett , Lisa Thiergart

Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them…

Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, and first-hand user…

Evaluating teaching effectiveness at scale remains a persistent challenge for large universities, particularly within engineering programs that enroll tens of thousands of students. Traditional methods, such as manual review of student…

The number and importance of AI-based systems in all domains is growing. With the pervasive use and the dependence on AI-based systems, the quality of these systems becomes essential for their practical usage. However, quality assurance for…

软件工程 · 计算机科学 2023-08-02 Michael Felderer , Rudolf Ramler

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks…

‹ 上一页 1 2 3 10 下一页 ›