中文
相关论文

相关论文: What do model reports say about their ChemBio benc…

200 篇论文

As frontier AI models become more capable, evaluating their potential to enable cyberattacks is crucial for ensuring the safe development of Artificial General Intelligence (AGI). Current cyber evaluation efforts are often ad-hoc, lacking…

密码学与安全 · 计算机科学 2025-04-23 Mikel Rodriguez , Raluca Ada Popa , Four Flynn , Lihao Liang , Allan Dafoe , Anna Wang

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer…

Companies that develop foundation models publish behavioral guidelines they pledge their models will follow, but it remains unclear if models actually do so. While providers such as OpenAI, Anthropic, and Google have published detailed…

计算与语言 · 计算机科学 2025-10-24 Ahmed Ahmed , Kevin Klyman , Yi Zeng , Sanmi Koyejo , Percy Liang

Generative AI (GenAI) models have become vital across industries, yet current evaluation methods have not adapted to their widespread use. Traditional evaluations often rely on benchmarks and fixed datasets, frequently failing to reflect…

Companies like OpenAI, Google DeepMind, and Anthropic have the stated goal of building artificial general intelligence (AGI) - AI systems that perform as well as or better than humans on a wide variety of cognitive tasks. However, there are…

计算机与社会 · 计算机科学 2023-07-19 Leonie Koessler , Jonas Schuett

Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequentially different, question: what older, cheaper, less-elicited models could do months or…

计算机与社会 · 计算机科学 2026-05-07 David Gringras , Misha Salahshoor

Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable foundation models that…

As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how…

计算机与社会 · 计算机科学 2025-07-10 Ayrton San Joaquin , Rokas Gipiškis , Leon Staufer , Ariel Gil

Generative artificial intelligence, and large language models in particular, have emerged as one of the most transformative paradigms in modern computer science. This automated survey provides an accessible treatment of the field as of…

机器学习 · 计算机科学 2026-04-10 Eduardo C. Garrido-Merchán , Álvaro López López

EVMbench, released by OpenAI, Paradigm, and OtterSec, is the first large-scale benchmark for AI agents on smart contract security. Its results -- agents detect up to 45.6% of vulnerabilities and exploit 72.2% of a curated subset -- have…

密码学与安全 · 计算机科学 2026-03-12 Chaoyuan Peng , Lei Wu , Yajin Zhou

Knowing more about the data used to build AI systems is critical for allowing different stakeholders to play their part in ensuring responsible and appropriate deployment and use. Meanwhile, a 2023 report shows that data transparency lags…

计算机与社会 · 计算机科学 2024-09-06 Sophia Worth , Ben Snaith , Arunav Das , Gefion Thuermer , Elena Simperl

AI coding assistants are rapidly becoming integral to modern software development. A key challenge in this space is the continual need to migrate and modernize codebases in response to evolving software ecosystems. Traditionally, such…

软件工程 · 计算机科学 2025-10-14 Victor May , Diganta Misra , Yanqi Luo , Anjali Sridhar , Justine Gehring , Silvio Soares Ribeiro Junior

Prominent AI experts have suggested that companies developing high-risk AI systems should be required to show that such systems are safe before they can be developed or deployed. The goal of this paper is to expand on this idea and explore…

计算机与社会 · 计算机科学 2024-06-25 Akash R. Wasil , Joshua Clymer , David Krueger , Emily Dardaman , Simeon Campos , Evan R. Murphy

The rapid evolution of artificial intelligence (AI), especially in the domain of Large Language Models (LLMs) and generative AI, has opened new avenues for application across various fields, yet its role in business education remains…

计算与语言 · 计算机科学 2024-01-09 Vahid Ashrafimoghari , Necdet Gürkan , Jordan W. Suchow

Currently, nearly all evaluations of foundation models focus on objective metrics, emphasizing quiz performance to define model capabilities. While this model-centric approach enables rapid performance assessment, it fails to reflect…

计算与语言 · 计算机科学 2025-06-03 Yijin Guo , Kaiyuan Ji , Xiaorong Zhu , Junying Wang , Farong Wen , Chunyi Li , Zicheng Zhang , Guangtao Zhai

Since late 2022, generative AI has taken the world by storm, with widespread use of tools including ChatGPT, Gemini, and Claude. Generative AI and large language model (LLM) applications are transforming how individuals find and access data…

人工智能 · 计算机科学 2024-05-08 Hannah Chafetz , Sampriti Saxena , Stefaan G. Verhulst

A comprehensive approach to addressing catastrophic risks from AI models should cover the full model lifecycle. This paper explores contingency plans for cases where pre-deployment risk management falls short: where either very dangerous…

计算机与社会 · 计算机科学 2023-10-03 Joe O'Brien , Shaun Ee , Zoe Williams

The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases and company blog posts, where model builders highlight results on selected benchmarks.…

人工智能 · 计算机科学 2026-05-15 Stefan Baack , Christo Buschek , Maty Bohacek

The impact of frontier AI (i.e., AI agents and foundation models) in cybersecurity is rapidly increasing. In this paper, we comprehensively analyze this trend through multiple aspects: quantitative benchmarks, qualitative literature review,…

密码学与安全 · 计算机科学 2025-12-01 Yujin Potter , Wenbo Guo , Zhun Wang , Tianneng Shi , Hongwei Li , Andy Zhang , Patrick Gage Kelley , Kurt Thomas , Dawn Song

Increased adoption of artificial intelligence (AI) systems into scientific workflows will result in an increasing technical debt as the distance between the data scientists and engineers who develop AI system components and scientists,…

人工智能 · 计算机科学 2021-03-08 Iain Barclay , Harrison Taylor , Alun Preece , Ian Taylor , Dinesh Verma , Geeth de Mel