English
Related papers

Related papers: The 2025 Foundation Model Transparency Index

200 papers

We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs). Our benchmark comprises 660 prompts across six…

Computation and Language · Computer Science 2025-03-17 Esben Kran , Hieu Minh "Jord" Nguyen , Akash Kundu , Sami Jawhar , Jinsuk Park , Mateusz Maria Jurewicz

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking…

Artificial Intelligence · Computer Science 2024-11-21 Anka Reuel , Amelia Hardy , Chandler Smith , Max Lamparth , Malcolm Hardy , Mykel J. Kochenderfer

The rapid emergence of large language models (LLMs) has raised urgent questions across the modern workforce about this new technology's strengths, weaknesses, and capabilities. For privacy professionals, the question is whether these AI…

Computers and Society · Computer Science 2025-08-13 Zane Witherspoon , Thet Mon Aye , YingYing Hao

The recent surge in open-source Large Language Models (LLMs), such as LLaMA, Falcon, and Mistral, provides diverse options for AI practitioners and researchers. However, most LLMs have only released partial artifacts, such as the final…

Recent advances in Foundation Models such as Large Language Models (LLMs) have propelled them to the forefront of Recommender Systems (RS). Despite their utility, there is a growing concern that LLMs might inadvertently perpetuate societal…

Information Retrieval · Computer Science 2024-05-30 Wenyue Hua , Yingqiang Ge , Shuyuan Xu , Jianchao Ji , Yongfeng Zhang

Do leading LLM developers possess a proprietary ``secret sauce'', or is LLM performance driven by scaling up compute? Using training and benchmark data for 809 models released between 2022 and 2025, we estimate scaling-law regressions with…

Artificial Intelligence · Computer Science 2026-05-05 Matthias Mertens , Natalia Fischl-Lanzoni , Neil Thompson

The advent of foundation models (FMs), large-scale pre-trained models with strong generalization capabilities, has opened new frontiers for financial engineering. While general-purpose FMs such as GPT-4 and Gemini have demonstrated…

Computational Finance · Quantitative Finance 2025-12-16 Liyuan Chen , Shuoling Liu , Jiangpeng Yan , Xiaoyu Wang , Henglin Liu , Chuang Li , Kecheng Jiao , Jixuan Ying , Yang Veronica Liu , Qiang Yang , Xiu Li

When encountering increasingly frequent performance improvements or cost reductions from a new large language model (LLM), developers of applications leveraging LLMs must decide whether to take advantage of these improvements or stay with…

Computation and Language · Computer Science 2025-02-20 Rubing Li , João Sedoc , Arun Sundararajan

As concerns surrounding AI-driven labor displacement intensify in knowledge-intensive sectors, existing benchmarks fail to measure performance on tasks that define practical professional expertise. Finance, in particular, has been…

Whether and how data scientists, statisticians and modellers should be accountable for the AI systems they develop remains a controversial and highly debated topic, especially given the complexity of AI systems and the difficulties in…

Artificial Intelligence · Computer Science 2023-09-12 Cassandra Bird , Daniel Williamson , Sabina Leonelli

Despite the development of effective deepfake detectors in recent years, recent studies have demonstrated that biases in the data used to train these detectors can lead to disparities in detection accuracy across different races and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Yan Ju , Shu Hu , Shan Jia , George H. Chen , Siwei Lyu

Business processes underpin a large number of enterprise operations including processing loan applications, managing invoices, and insurance claims. There is a large opportunity for infusing AI to reduce cost or provide better customer…

Artificial Intelligence · Computer Science 2020-01-22 Steve T. K. Jan , Vatche Ishakian , Vinod Muthusamy

Developing AI tools that preserve fairness is of critical importance, specifically in high-stakes applications such as those in healthcare. However, health AI models' overall prediction performance is often prioritized over the possible…

Machine Learning · Computer Science 2023-05-22 Raphael Poulain , Mirza Farhan Bin Tarek , Rahmatollah Beheshti

Federated Learning is a promising machine learning paradigm when multiple parties collaborate to build a high-quality machine learning model. Nonetheless, these parties are only willing to participate when given enough incentives, such as a…

Machine Learning · Computer Science 2021-06-10 Shuaicheng Ma , Yang Cao , Li Xiong

Over the past few years, a tremendous growth of machine learning was brought about by a significant increase in adoption of cloud-based services. As a result, various solutions have been proposed in which the machine learning models run on…

Cryptography and Security · Computer Science 2021-08-02 Tanveer Khan , Alexandros Bakas , Antonis Michalas

Bias in Foundation Models (FMs) - trained on vast datasets spanning societal and historical knowledge - poses significant challenges for fairness and equity across fields such as healthcare, education, and finance. These biases, rooted in…

Machine Learning · Computer Science 2025-01-22 Shuzhou Sun , Li Liu , Yongxiang Liu , Zhen Liu , Shuanghui Zhang , Janne Heikkilä , Xiang Li

Federated Learning (FL) is a distributed training paradigm that enables clients scattered across the world to cooperatively learn a global model without divulging confidential data. However, FL faces a significant challenge in the form of…

Machine Learning · Computer Science 2023-11-16 Xidong Wu , Wan-Yi Lin , Devin Willmott , Filipe Condessa , Yufei Huang , Zhenzhen Li , Madan Ravi Ganesh

If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations, leading to less reliable benchmarks for…

Computation and Language · Computer Science 2025-07-17 Joe Needham , Giles Edkins , Govind Pimpale , Henning Bartsch , Marius Hobbhahn

Foundation models have gained growing interest in the IoT domain due to their reduced reliance on labeled data and strong generalizability across tasks, which address key limitations of traditional machine learning approaches. However, most…

Machine Learning · Computer Science 2025-10-10 Hui Wei , Dong Yoon Lee , Shubham Rohal , Zhizhang Hu , Ryan Rossi , Shiwei Fang , Shijia Pan

In today's mobile application marketplace, the ability of consumers to make informed choices regarding their privacy is extremely limited. Consumers largely rely on privacy policies and app permission mechanisms, but these do an inadequate…

Computers and Society · Computer Science 2015-01-05 Steven C. Isley