中文
相关论文

相关论文: AI-Assisted Systematization for Evaluating GenAI S…

200 篇论文

Evaluation of potential AGI systems and methods is difficult due to the breadth of the engineering goal. We have no methods for perfect evaluation of the end state, and instead measure performance on small tests designed to provide…

人工智能 · 计算机科学 2025-10-03 John Hawkins

In social service, administrative burdens and decision-making challenges often hinder practitioners from performing effective casework. Generative AI (GenAI) offers significant potential to streamline these tasks, yet exacerbates concerns…

人机交互 · 计算机科学 2025-02-28 Yugin Tan , Kai Xin Soh , Renwen Zhang , Jungup Lee , Han Meng , Biswadeep Sen , Yi-Chieh Lee

Generative artificial intelligence (GenAI) is increasingly used to support a wide range of human tasks, yet empirical evidence on its effect on creativity remains scattered. Can GenAI generate ideas that are creative? To what extent can it…

人机交互 · 计算机科学 2025-05-26 Niklas Holzner , Sebastian Maier , Stefan Feuerriegel

As generative artificial intelligence (AI) continues to transform education, most existing AI evaluations rely primarily on technical performance metrics such as accuracy or task efficiency while overlooking human identity, learner agency,…

计算机与社会 · 计算机科学 2025-12-05 Shi Ding , Brian Magerko

Generative artificial intelligence (GenAI) holds the potential to transform the delivery, cultivation, and evaluation of human learning. This Perspective examines the integration of GenAI as a tool for human learning, addressing its…

人机交互 · 计算机科学 2024-09-06 Lixiang Yan , Samuel Greiff , Ziwen Teuber , Dragan Gašević

Generative AI (GenAI) has revolutionized content generation, offering transformative capabilities for improving language coherence, readability, and overall quality. This manuscript explores the application of qualitative, quantitative, and…

计算与语言 · 计算机科学 2024-11-28 Saman Sarraf

The widespread use of generative AI systems is coupled with significant ethical and social challenges. As a result, policymakers, academic researchers, and social advocacy groups have all called for such systems to be audited. However,…

计算机与社会 · 计算机科学 2024-07-09 Jakob Mokander , Justin Curl , Mihir Kshirsagar

The advent of advanced AI underscores the urgent need for comprehensive safety evaluations, necessitating collaboration across communities (i.e., AI, software engineering, and governance). However, divergent practices and terminologies…

软件工程 · 计算机科学 2024-05-17 Boming Xia , Qinghua Lu , Liming Zhu , Zhenchang Xing

Despite the potential of generative AI (GenAI) design tools to enhance design processes, professionals often struggle to integrate AI into their workflows. Fundamental cognitive challenges include the need to specify all design criteria as…

人机交互 · 计算机科学 2025-06-17 Frederic Gmeiner , Kaitao Luo , Ye Wang , Kenneth Holstein , Nikolas Martelaro

Fact verification is a critical yet underexplored component of non-litigation legal practice. While existing research has examined automation in legal workflow and human-AI collaboration in high-stakes domains, little is known about how…

人机交互 · 计算机科学 2026-02-10 Sirui Han , Yuyao Zhang , Yidan Huang , Xueyan Li , Chengzhong Liu , Yike Guo

Measurement is essential to improving AI performance and mitigating harms for marginalized groups. As generative AI systems are rapidly deployed across geographies and contexts, AI measurement practices must be designed to support…

Recent debates on artificial intelligence increasingly emphasise questions of AI consciousness and moral status, yet there remains little agreement on how such properties should be evaluated. In this paper, we argue that awareness offers a…

人工智能 · 计算机科学 2026-01-22 Nadine Meertens , Suet Lee , Ophelia Deroy

As full AI-based automation remains out of reach in most real-world applications, the focus has instead shifted to leveraging the strengths of both human and AI agents, creating effective collaborative systems. The rapid advances in this…

人机交互 · 计算机科学 2024-04-19 Steffen Holter , Mennatallah El-Assady

Through a systematization of generative AI (GenAI) stakeholder goals and expectations, this work seeks to uncover what value different stakeholders see in their contributions to the GenAI supply line. This valuation enables us to understand…

人工智能 · 计算机科学 2024-08-02 Amruta Mahuli , Asia Biega

Technical standards, or simply standards, are established documented guidelines and rules that facilitate the interoperability, quality, and accuracy of systems and processes. In recent years, we have witnessed an emerging paradigm shift…

计算机与社会 · 计算机科学 2025-03-10 Joseph Marvin Imperial , Matthew D. Jones , Harish Tayyar Madabushi

Generative AI (GAI) tools have seen rapid adoption in educational settings, yet their role in fostering critical thinking remains underexplored. While previous studies have examined GAI as a tutor for specific lessons or as a tool for…

计算机与社会 · 计算机科学 2025-09-10 W. F. Lamberti , S. R. Lawrence , D. White , S. Kim , S. Abdullah

Sensemaking in collaborative work and learning is increasingly supported by GenAI systems, however, emerging evidence suggests that poorly designed GenAI systems tend to provide explicit instruction that groups passively follow, fostering…

人机交互 · 计算机科学 2026-03-12 Yihang Zhao , Wenxin Zhang , Amy Rechkemmer , Albert Meroño Peñuela , Elena Simperl

This paper proposes a comprehensive analysis of existing concepts coming from different disciplines tackling the notion of intelligence, namely psychology and engineering, and from disciplines aiming to regulate AI innovations, namely AI…

人工智能 · 计算机科学 2021-05-10 Gauthier Chassang , Mogens Thomsen , Pierre Rumeau , Florence Sèdes , Alejandra Delfin

AI assistants can increasingly generate and evolve test cases. The challenge is no longer merely to produce them, but also to help engineers understand why a generated artefact exists and what supports it. Existing work has focused on…

软件工程 · 计算机科学 2026-04-27 Eduard Paul Enoiu , Robert Feldt

Societal stereotypes are at the center of a myriad of responsible AI interventions targeted at reducing the generation and propagation of potentially harmful outcomes. While these efforts are much needed, they tend to be fragmented and…

计算机与社会 · 计算机科学 2025-10-02 Aida Davani , Sunipa Dev , Héctor Pérez-Urbina , Vinodkumar Prabhakaran