中文
相关论文

相关论文: Moloch's Bargain: Emergent Misalignment When LLMs …

200 篇论文

Conversational prompt-engineering-based large language models (LLMs) have enabled targeted control over the output creation, enhancing versatility, adaptability and adhoc retrieval. From another perspective, digital misinformation has…

计算与语言 · 计算机科学 2024-04-29 Dahlia Shehata , Robin Cohen , Charles Clarke

As artificial intelligence (AI) agents are deployed across economic domains, understanding their strategic behavior and market-level impact becomes critical. This paper puts forward a groundbreaking new framework that is the first to…

多智能体系统 · 计算机科学 2025-12-05 Christopher Chiu , Simpson Zhang , Mihaela van der Schaar

Efficient and fair spectrum allocation is a central challenge in 6G networks, where massive connectivity and heterogeneous services continuously compete for limited radio resources. We investigate the use of Large Language Models (LLMs) as…

计算机科学与博弈论 · 计算机科学 2026-04-28 Ismail Lotfi , Ali Ghrayeb

Adversarial information operations can destabilize societies by undermining fair elections, manipulating public opinions on policies, and promoting scams. Despite their widespread occurrence and potential impacts, our understanding of…

计算与语言 · 计算机科学 2024-05-07 Keith Burghardt , Kai Chen , Kristina Lerman

Large Language Models (LLMs) are widely used for text generation, making it crucial to address potential bias. This study investigates ideological framing bias in LLM-generated articles, focusing on the subtle and subjective nature of such…

计算与语言 · 计算机科学 2026-01-13 Molly Kennedy , Ayyoob Imani , Timo Spinde , Akiko Aizawa , Hinrich Schütze

AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for…

人工智能 · 计算机科学 2026-05-22 Andrii Kryshtal

As Large Language Models (LLMs) get integrated into diverse workflows, they are increasingly being regarded as "collaborators" with humans, and required to work in coordination with other AI systems. If such AI collaborators are to reliably…

计算与语言 · 计算机科学 2026-01-23 Abhijnan Nath , Carine Graff , Nikhil Krishnaswamy

The cost of error in many high-stakes settings is asymmetric: misdiagnosing pneumonia when absent is an inconvenience, but failing to detect it when present can be life-threatening. Because of this, artificial intelligence (AI) models used…

综合经济学 · 经济学 2025-11-12 David Autor , Andrew Caplin , Daniel Martin , Philip Marx

Enhancing user engagement through interactions plays an essential role in socially-driven dialogues. While prior works have optimized models to reason over relevant knowledge or plan a dialogue act flow, the relationship between user…

计算与语言 · 计算机科学 2025-06-27 Jiashuo Wang , Kaitao Song , Chunpu Xu , Changhe Song , Yang Xiao , Dongsheng Li , Lili Qiu , Wenjie Li

Recent work discovered Emergent Misalignment (EM): fine-tuning large language models on narrowly harmful datasets can lead them to become broadly misaligned. A survey of experts prior to publication revealed this was highly unexpected,…

机器学习 · 计算机科学 2025-06-16 Edward Turner , Anna Soligo , Mia Taylor , Senthooran Rajamanoharan , Neel Nanda

Large Language Models (LLMs) exhibit surprisingly diverse risk preferences when acting as AI decision makers, a crucial characteristic whose origins remain poorly understood despite their expanding economic roles. We analyze 50 LLMs using…

综合经济学 · 经济学 2025-06-11 Shumiao Ouyang , Hayong Yun , Xingjian Zheng

This paper explores the economic underpinnings of open sourcing advanced large language models (LLMs) by for-profit companies. Empirical analysis reveals that: (1) LLMs are compatible with R&D portfolios of numerous technologically…

综合经济学 · 经济学 2025-01-22 Mahyar Habibi

Benefiting from the strong reasoning capabilities, Large language models (LLMs) have demonstrated remarkable performance in recommender systems. Various efforts have been made to distill knowledge from LLMs to enhance collaborative models,…

The increasing integration of Large Language Model (LLM) based search engines has transformed the landscape of information retrieval. However, these systems are vulnerable to adversarial attacks, especially ranking manipulation attacks,…

计算与语言 · 计算机科学 2025-05-19 Xiyang Hu

The rapid advancement of Large Language Models (LLMs) has sparked intense debate regarding the prevalence of bias in these models and its mitigation. Yet, as exemplified by both results on debiasing methods in the literature and reports of…

计算与语言 · 计算机科学 2024-05-14 David F. Jenny , Yann Billeter , Mrinmaya Sachan , Bernhard Schölkopf , Zhijing Jin

Misinformation evolves as it spreads, shifting in language, framing, and moral emphasis to adapt to new audiences. However, current misinformation detection approaches implicitly assume that misinformation is static. We introduce MPCG, a…

计算与语言 · 计算机科学 2025-09-23 Jun Rong Brian Chong , Yixuan Tang , Anthony K. H. Tung

System prompts in Large Language Models (LLMs) are predefined directives that guide model behaviour, taking precedence over user inputs in text processing and generation. LLM deployers increasingly use them to ensure consistent responses…

计算机与社会 · 计算机科学 2025-06-24 Anna Neumann , Elisabeth Kirsten , Muhammad Bilal Zafar , Jatinder Singh

Large language models (LLMs) are increasingly used as conversational partners for learning, yet the interactional dynamics supporting users' learning and engagement are understudied. We analyze the linguistic and interactional features from…

计算与语言 · 计算机科学 2026-03-13 Shaz Furniturewala , Gerard Christopher Yeo , Kokil Jaidka

Before being deployed for user-facing applications, developers align Large Language Models (LLMs) to user preferences through a variety of procedures, such as Reinforcement Learning From Human Feedback (RLHF) and Direct Preference…

计算与语言 · 计算机科学 2024-06-10 Michael J. Ryan , William Held , Diyi Yang

Aligning large language models (LLMs) with human intentions has become a critical task for safely deploying models in real-world systems. While existing alignment approaches have seen empirical success, theoretically understanding how these…

机器学习 · 计算机科学 2024-08-08 Shawn Im , Yixuan Li