English
Related papers

Related papers: What do model reports say about their ChemBio benc…

200 papers

The downstream use cases, benefits, and risks of AI models depend significantly on what sort of access is provided to the model, and who it is provided to. Though existing safety frameworks and AI developer usage policies recognise that the…

Computers and Society · Computer Science 2024-12-03 Edward Kembery , Tom Reed

DeepSeek v3, developed in China, was released in December 2024, followed by Alibaba's Qwen 2.5 Max in January 2025 and Qwen3 235B in April 2025. These free and open-source models offer significant potential for academic writing and content…

Computers and Society · Computer Science 2025-05-05 Omer Aydin , Enis Karaarslan , Fatih Safa Erenay , Nebojsa Bacanin

Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions. These limitations drive the proliferation of $\textit{shadow APIs}$, third-party…

Cryptography and Security · Computer Science 2026-03-06 Yage Zhang , Yukun Jiang , Zeyuan Chen , Michael Backes , Xinyue Shen , Yang Zhang

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for…

Artificial Intelligence · Computer Science 2025-05-12 Markov Grey , Charbel-Raphaël Segerie

With the advancement of AI models, more software systems are adopting AI as a component to facilitate automation. Pre-trained models (PTMs) have become a cornerstone of AI-based software, allowing for rapid integration and development with…

Software Engineering · Computer Science 2026-05-01 Haoyu Gao , Mansooreh Zahedi , Wenxin Jiang , Hong Yi Lin , James Davis , Christoph Treude

Self-replication with no human intervention is broadly recognized as one of the principal red lines associated with frontier AI systems. While leading corporations such as OpenAI and Google DeepMind have assessed GPT-o3-mini and Gemini on…

Artificial Intelligence · Computer Science 2025-03-26 Xudong Pan , Jiarun Dai , Yihe Fan , Minyuan Luo , Changyi Li , Min Yang

We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench is curated from repeatedly used, authentic university…

Artificial Intelligence · Computer Science 2026-03-04 Chongyang Gao , Diji Yang , Shuyan Zhou , Xichen Yan , Luchuan Song , Shuo Li , Kezhen Chen

AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard training and validation datasets were never designed to capture. Evaluating these systems…

Artificial Intelligence · Computer Science 2026-05-12 Prasanna Desikan , Harshit Rajgarhia , Shivali Dalmia , Ananya Mantravadi

Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark…

Artificial Intelligence · Computer Science 2026-05-28 Aakash Pant , Kavya Shah , Apoorv Agnihotri , Sneha Nikam , Prasaanth Balraj , Nakul Jain

The success of OpenAI's ChatGPT in 2023 has spurred financial enterprises into exploring Generative AI applications to reduce costs or drive revenue within different lines of businesses in the Financial Industry. While these applications…

Risk Management · Quantitative Finance 2025-03-21 Anwesha Bhattacharyya , Ye Yu , Hanyu Yang , Rahul Singh , Tarun Joshi , Jie Chen , Kiran Yalavarthy

The range of application of artificial intelligence (AI) is vast, as is the potential for harm. Growing awareness of potential risks from AI systems has spurred action to address those risks, while eroding confidence in AI systems and the…

TREs are widely, and increasingly used to support statistical analysis of sensitive data across a range of sectors (e.g., health, police, tax and education) as they enable secure and transparent research whilst protecting data…

The rapid evolution of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has driven major gains in reasoning, perception, and generation across language and vision, yet whether these advances translate into…

As the deployment of artificial intelligence (AI) is changing many fields and industries, there are concerns about AI systems making decisions and recommendations without adequately considering various ethical aspects, such as…

Computers and Society · Computer Science 2023-10-02 Conrad Sanderson , Qinghua Lu , David Douglas , Xiwei Xu , Liming Zhu , Jon Whittle

Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluating large language…

Nova Premier is Amazon's most capable multimodal foundation model and teacher for model distillation. It processes text, images, and video with a one-million-token context window, enabling analysis of large codebases, 400-page documents,…

Cryptography and Security · Computer Science 2025-07-10 Satyapriya Krishna , Ninareh Mehrabi , Abhinav Mohanty , Matteo Memelli , Vincent Ponzo , Payal Motwani , Rahul Gupta

Organizational leaders are being asked to make high-stakes decisions about AI deployment without dependable evidence of what these systems actually do in the environments they oversee. The predominant AI evaluation ecosystem yields scalable…

Computers and Society · Computer Science 2026-03-31 Reva Schwartz , Gabriella Waters

While Large Language Models (LLMs) are rapidly integrating into daily life, research on their risks often remains lab-based and disconnected from the problems users encounter "in the wild." While recent HCI research has begun to explore…

Computers and Society · Computer Science 2025-09-12 Lingyao Li , Renkai Ma , Zhaoqian Xue , Junjie Xiong

Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's Model Spec (OpenAI, 2025a), integrated into post-training via methods like character…

Artificial Intelligence · Computer Science 2026-05-26 Arya Jakkli , Senthooran Rajamanoharan , Neel Nanda