English
Related papers

Related papers: Training AI to be Loyal

200 papers

Large language models that exhibit instruction-following behaviour represent one of the biggest recent upheavals in conversational interfaces, a trend in large part fuelled by the release of OpenAI's ChatGPT, a proprietary large language…

Computation and Language · Computer Science 2023-07-13 Andreas Liesenfeld , Alianda Lopez , Mark Dingemanse

Big models have greatly advanced AI's ability to understand, generate, and manipulate information and content, enabling numerous applications. However, as these models become increasingly integrated into everyday life, their inherent…

Computers and Society · Computer Science 2023-10-27 Xiaoyuan Yi , Jing Yao , Xiting Wang , Xing Xie

Since 2023, generative AI has rapidly advanced in the music domain. Despite significant technological advancements, music-generative models raise critical ethical challenges, including a lack of transparency and accountability, along with…

With the growing adoption of AI systems, reasoning about how society can exert control over AI becomes an increasingly urgent problem. Existing work on democratic control largely focuses on macro-level governance. In contrast, we propose a…

Computers and Society · Computer Science 2026-05-19 Paul Anton Bachmann , Niclas Boehmer , Lukas Daniel Klausner , Martin Lackner

Recent advances in Large Language Models (LLMs) have enabled human-like responses across various tasks, raising questions about their ethical decision-making capabilities and potential biases. This study systematically evaluates how nine…

Computers and Society · Computer Science 2025-11-03 Wentao Xu , Yile Yan , Yuqi Zhu

Recent studies on software tool manipulation with large language models (LLMs) mostly rely on closed model APIs. The industrial adoption of these models is substantially constrained due to the security and robustness risks in exposing…

Computation and Language · Computer Science 2023-05-29 Qiantong Xu , Fenglu Hong , Bo Li , Changran Hu , Zhengyu Chen , Jian Zhang

Security and ethics are both core to ensuring that a machine learning system can be trusted. In production machine learning, there is generally a hand-off from those who build a model to those who deploy a model. In this hand-off, the…

Computers and Society · Computer Science 2020-07-10 Abhishek Gupta , Erick Galinkin

Artificial intelligence (AI) provides many opportunities to improve private and public life. Discovering patterns and structures in large troves of data in an automated manner is a core component of data science, and currently drives…

Machine Learning · Computer Science 2020-09-25 Vaishak Belle , Ioannis Papantonis

Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship. Here, we show how this creates a significant…

Computation and Language · Computer Science 2025-07-31 Lujain Ibrahim , Franziska Sofia Hafner , Luc Rocher

Regulatory frameworks, such as the EU AI Act, encourage openness of general-purpose AI models by offering legal exemptions for "open-source" models. Despite this legislative attention on openness, the definition of open-source foundation…

Computer Science and Game Theory · Computer Science 2025-10-27 Tori Qiu , Benjamin Laufer , Jon Kleinberg , Hoda Heidari

The rapid growth of submissions to top-tier Artificial Intelligence (AI) and Machine Learning (ML) conferences has prompted many venues to transition from closed to open review platforms. Some have fully embraced open peer reviews, allowing…

Digital Libraries · Computer Science 2025-10-16 Jing Yang

Growing concerns over the lack of transparency in AI, particularly in high-stakes fields like healthcare and finance, drive the need for explainable and trustworthy systems. While Large Language Models (LLMs) perform exceptionally well in…

Artificial Intelligence · Computer Science 2025-06-10 Fadi Al Machot , Martin Thomas Horsch , Habib Ullah

Conventional machine learning studies generally assume close-environment scenarios where important factors of the learning process hold invariant. With the great success of machine learning, nowadays, more and more practical tasks,…

Machine Learning · Computer Science 2022-08-10 Zhi-Hua Zhou

A core part of human intelligence is the ability to work flexibly with others to achieve goals. The incorporation of artificial agents into human spaces is making increasing demands on artificial intelligence (AI) to demonstrate and…

Human-Computer Interaction · Computer Science 2026-03-30 William J. Bingley , S. Alexander Haslam , Janet Wiles

This study critically examines the commonly held assumption that explicability in artificial intelligence (AI) systems inherently boosts user trust. Utilizing a meta-analytical approach, we conducted a comprehensive examination of the…

Artificial Intelligence · Computer Science 2025-04-18 Zahra Atf , Peter R. Lewis

We critically examine the limitations of current AI models in achieving autonomous learning and propose a learning architecture inspired by human and animal cognition. The proposed framework integrates learning from observation (System A)…

Artificial Intelligence · Computer Science 2026-03-17 Emmanuel Dupoux , Yann LeCun , Jitendra Malik

The increasing prevalence of Artificial Intelligence (AI) in safety-critical contexts such as air-traffic control leads to systems that are practical and efficient, and to some extent explainable to humans to be trusted and accepted. The…

Computers and Society · Computer Science 2023-06-28 Sabine Theis , Sophie Jentzsch , Fotini Deligiannaki , Charles Berro , Arne Peter Raulf , Carmen Bruder

Democratization of AI means not only that people can freely use AI, but also that people can collectively decide how AI is to be used. In particular, collective decision-making power is required to redress the negative externalities from…

Computers and Society · Computer Science 2023-05-23 Alan Chan , Herbie Bradley , Nitarshan Rajkumar

As Artificial Intelligence (AI), particularly Large Language Models (LLMs), becomes increasingly embedded in education systems worldwide, ensuring their ethical, legal, and contextually appropriate deployment has become a critical policy…

Computers and Society · Computer Science 2026-05-27 Sara Alaswad , Tatiana Kalganova , Wasan Awad

Federated Learning (FL) has emerged as a promising paradigm to train machine learning models collaboratively while preserving data privacy. However, its widespread adoption faces several challenges, including scalability, heterogeneous data…

Machine Learning · Computer Science 2024-05-14 Ilir Murturi , Praveen Kumar Donta , Schahram Dustdar