English
Related papers

Related papers: Societal Impacts Research Requires Benchmarks for …

200 papers

Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning research and an increasingly prominent method for scientific…

Machine Learning · Computer Science 2025-10-28 Timo Freiesleben , Sebastian Zezulka

Recently, the relationship between automated and human evaluation of topic models has been called into question. Method developers have staked the efficacy of new topic model variants on automated measures, and their failure to approximate…

Computation and Language · Computer Science 2022-10-31 Alexander Hoyle , Pranav Goel , Rupak Sarkar , Philip Resnik

Intelligent systems for the annotation of media content are increasingly being used for the automation of parts of social science research. In this domain the problem of integrating various Artificial Intelligence (AI) algorithms into a…

Multiagent Systems · Computer Science 2018-06-05 Ilias Flaounas , Thomas Lansdall-Welfare , Panagiota Antonakaki , Nello Cristianini

Generative AI (GenAI) tools are radically expanding the scope and capability of automation in knowledge work such as academic research. While promising for augmenting cognition and streamlining processes, AI-assisted research tools may also…

Human-Computer Interaction · Computer Science 2025-04-22 Runlong Ye , Matthew Varona , Oliver Huang , Patrick Yung Kang Lee , Michael Liut , Carolina Nobre

Disagreements are pervasive in human communication. In this paper we investigate what makes disagreement constructive. To this end, we construct WikiDisputes, a corpus of 7 425 Wikipedia Talk page conversations that contain content…

Computation and Language · Computer Science 2021-01-27 Christine de Kock , Andreas Vlachos

The co creativity community is making significant progress in developing more sophisticated and tailored systems to support and enhance human creativity. Design considerations from prior work can serve as a valuable and efficient foundation…

Human-Computer Interaction · Computer Science 2025-06-30 Saloni Singh , Koen Hindriks , Dirk Heylen , Kim Baraka

We analyze the challenges of benchmarking scientific (multi)-agentic systems, including the difficulty of distinguishing reasoning from retrieval, the risks of data/model contamination, the lack of reliable ground truth for novel research…

Computers and Society · Computer Science 2026-04-07 Marcin Abram

With the increasing pervasiveness of algorithms across industry and government, a growing body of work has grappled with how to understand their societal impact and ethical implications. Various methods have been used at different stages of…

Computers and Society · Computer Science 2022-07-21 Julia Barnett , Nicholas Diakopoulos

The rapid proliferation of AI-generated image tools is transforming the art and design fields, challenging traditional notions of creativity and impacting both professional and non-professional users. For the purposes of this paper, we…

Human-Computer Interaction · Computer Science 2024-06-18 Yuying Tang , Ningning Zhang , Mariana Ciancia , Zhigang Wang

Generative AI models are capable of performing a wide variety of tasks that have traditionally required creativity and human understanding. During training, they learn patterns from existing data and can subsequently generate new content…

The accelerated development, deployment and adoption of artificial intelligence systems has been fuelled by the increasing involvement of big tech. This has been accompanied by increasing ethical concerns and intensified societal and…

Computers and Society · Computer Science 2025-12-04 Alex Hernandez-Garcia , Alexandra Volokhova , Ezekiel Williams , Dounia Shaaban Kabakibo

With significant advances in generative AI, new technologies are rapidly being deployed with generative components. Generative models are typically trained on large datasets, resulting in model behaviors that can mimic the worst of the…

Machine Learning · Computer Science 2023-06-13 Susan Hao , Piyush Kumar , Sarah Laszlo , Shivani Poddar , Bhaktipriya Radharapu , Renee Shelby

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized…

Computation and Language · Computer Science 2026-05-28 Katharina Deckenbach , Haritz Puerto , Jonas Geiping , Sahar Abdelnabi

The amount of text generated daily on social media is gigantic and analyzing this text is useful for many purposes. To understand what lies beneath a huge amount of text, we need dependable and effective computing techniques from…

Information Retrieval · Computer Science 2025-08-04 Ngozichukwuka Onah , Nadine Steinmetz , Hani Al-Sayeh , Kai-Uwe Sattler

An occupation is comprised of interconnected tasks, and it is these tasks, not occupations themselves, that are affected by AI. To evaluate how tasks may be impacted, previous approaches utilized manual annotations or coarse-grained…

Computers and Society · Computer Science 2024-07-31 Ali Akbar Septiandri , Marios Constantinides , Daniele Quercia

Language-based AI systems are diffusing into society, bringing positive and negative impacts. Mitigating negative impacts depends on accurate impact assessments, drawn from an empirical evidence base that makes causal connections between AI…

Computers and Society · Computer Science 2024-10-08 Merlin Stein , Jamie Bernardi , Connor Dunlop

Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their capabilities. Current evaluation methodology, mostly based…

Artificial Intelligence · Computer Science 2024-01-01 Xiting Wang , Liming Jiang , Jose Hernandez-Orallo , David Stillwell , Luning Sun , Fang Luo , Xing Xie

The rapid proliferation of AI-generated content, driven by advances in generative adversarial networks, diffusion models, and multimodal large language models, has made the creation and dissemination of synthetic media effortless,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Guangyu Lin , Li Lin , Christina P. Walker , Daniel S. Schiff , Shu Hu

Cultural evaluation of large language models has become increasingly important, yet current benchmarks often reduce culture to static facts or homogeneous values. This view conflicts with anthropological accounts that emphasize culture as…

Computation and Language · Computer Science 2025-10-23 Mai AlKhamissi , Yunze Xiao , Badr AlKhamissi , Mona Diab

How should we evaluate the quality of generative models? Many existing metrics focus on a model's producibility, i.e. the quality and breadth of outputs it can generate. However, the actual value from using a generative model stems not just…

Machine Learning · Computer Science 2025-11-13 Keyon Vafa , Sarah Bentley , Jon Kleinberg , Sendhil Mullainathan