中文
相关论文

相关论文: SciHorizon: Benchmarking AI-for-Science Readiness …

200 篇论文

Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face…

Large language models (LLMs) have exhibited exceptional capabilities in natural language understanding and generation, image recognition, and multimodal tasks, charting a course towards AGI and emerging as a central issue in the global…

计算与语言 · 计算机科学 2025-11-20 Guoqiang Liang , Jingqian Gong , Mengxuan Li , Gege Lin , Shuo Zhang

The rapid advancements in large language models (LLMs), particularly in their reasoning capabilities, hold transformative potential for addressing complex challenges and boosting scientific discovery in atmospheric science. However,…

机器学习 · 计算机科学 2025-10-07 Chenyue Li , Wen Deng , Mengqian Lu , Binhang Yuan

Recent advances in large language models (LLMs) have enabled agentic systems that translate natural language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and…

Large language models (LLMs) have demonstrated significant potential to accelerate scientific discovery as valuable tools for analyzing data, generating hypotheses, and supporting innovative approaches in various scientific fields. In this…

计算与语言 · 计算机科学 2025-10-30 Jin Huang , Silviu Cucerzan , Sujay Kumar Jauhar , Ryen W. White

Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research…

The explosive growth of AI research has created unprecedented information overload, increasing the demand for scientific summarization at multiple levels of granularity beyond traditional abstracts. While LLMs are increasingly adopted for…

计算与语言 · 计算机科学 2026-03-18 Han Jang , Junhyeok Lee , Kyu Sung Choi

In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scientific exploration, bringing significant advancements across…

人工智能 · 计算机科学 2023-10-13 Shuaiwen Leon Song , Bonnie Kruft , Minjia Zhang , Conglong Li , Shiyang Chen , Chengming Zhang , Masahiro Tanaka , Xiaoxia Wu , Jeff Rasley , Ammar Ahmad Awan , Connor Holmes , Martin Cai , Adam Ghanem , Zhongzhu Zhou , Yuxiong He , Pete Luferenko , Divya Kumar , Jonathan Weyn , Ruixiong Zhang , Sylwester Klocek , Volodymyr Vragov , Mohammed AlQuraishi , Gustaf Ahdritz , Christina Floristean , Cristina Negri , Rao Kotamarthi , Venkatram Vishwanath , Arvind Ramanathan , Sam Foreman , Kyle Hippe , Troy Arcomano , Romit Maulik , Maxim Zvyagin , Alexander Brace , Bin Zhang , Cindy Orozco Bohorquez , Austin Clyde , Bharat Kale , Danilo Perez-Rivera , Heng Ma , Carla M. Mann , Michael Irvin , J. Gregory Pauloski , Logan Ward , Valerie Hayot , Murali Emani , Zhen Xie , Diangen Lin , Maulik Shukla , Ian Foster , James J. Davis , Michael E. Papka , Thomas Brettin , Prasanna Balaprakash , Gina Tourassi , John Gounley , Heidi Hanson , Thomas E Potok , Massimiliano Lupo Pasini , Kate Evans , Dan Lu , Dalton Lunga , Junqi Yin , Sajal Dash , Feiyi Wang , Mallikarjun Shankar , Isaac Lyngaas , Xiao Wang , Guojing Cong , Pei Zhang , Ming Fan , Siyan Liu , Adolfy Hoisie , Shinjae Yoo , Yihui Ren , William Tang , Kyle Felker , Alexey Svyatkovskiy , Hang Liu , Ashwin Aji , Angela Dalton , Michael Schulte , Karl Schulz , Yuntian Deng , Weili Nie , Josh Romero , Christian Dallago , Arash Vahdat , Chaowei Xiao , Thomas Gibbs , Anima Anandkumar , Rick Stevens

Large language models (LLMs) excel across many natural language processing tasks but face challenges in domain-specific, analytical tasks such as conducting research surveys. This study introduces ResearchArena, a benchmark designed to…

人工智能 · 计算机科学 2025-09-09 Hao Kang , Chenyan Xiong

Large language models and autonomous AI agents have evolved rapidly, resulting in a diverse array of evaluation benchmarks, frameworks, and collaboration protocols. Driven by the growing need for standardized evaluation and integration, we…

人工智能 · 计算机科学 2026-03-10 Mohamed Amine Ferrag , Norbert Tihanyi , Merouane Debbah

Artificial intelligence (AI) is transforming the practice of science. Machine learning and large language models (LLMs) can generate hypotheses at a scale and speed far exceeding traditional methods, offering the potential to accelerate…

人工智能 · 计算机科学 2025-12-18 Cristina Cornelio , Takuya Ito , Ryan Cory-Wright , Sanjeeb Dash , Lior Horesh

As large language models (LLMs) advance, the ultimate vision for their role in science is emerging: we could build an AI collaborator to effectively assist human beings throughout the entire scientific research process. We refer to this…

The rapid rise in popularity of Large Language Models (LLMs) with emerging capabilities has spurred public curiosity to evaluate and compare different LLMs, leading many researchers to propose their own LLM benchmarks. Noticing preliminary…

人工智能 · 计算机科学 2025-05-15 Timothy R. McIntosh , Teo Susnjak , Nalin Arachchilage , Tong Liu , Paul Watters , Malka N. Halgamuge

Artificial intelligence and machine learning are reshaping how we approach scientific discovery, not by replacing established methods but by extending what researchers can probe, predict, and design. In this roadmap we provide a…

Do large language models (LLMs) exhibit any forms of awareness similar to humans? In this paper, we introduce AwareBench, a benchmark designed to evaluate awareness in LLMs. Drawing from theories in psychology and philosophy, we define…

计算与语言 · 计算机科学 2024-02-19 Yuan Li , Yue Huang , Yuli Lin , Siyuan Wu , Yao Wan , Lichao Sun

With the rapid development of Large Language Models (LLMs), AI agents have demonstrated increasing proficiency in scientific tasks, ranging from hypothesis generation and experimental design to manuscript writing. Such agent systems are…

Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck: whether Large Language…

机器学习 · 计算机科学 2026-05-29 Sy-Tuyen Ho , Minghui Liu , Huy Nghiem , Furong Huang

Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However, concerns persist regarding the reliability of these benchmarks, which often lack clinical…

计算与语言 · 计算机科学 2026-04-30 Wenting Chen , Guo Yu , Yiu-Fai Cheung , Meidan Ding , Jie Liu , Zizhan Ma , Wenxuan Wang , Linlin Shen

With recent Nobel Prizes recognising AI contributions to science, Large Language Models (LLMs) are transforming scientific research by enhancing productivity and reshaping the scientific method. LLMs are now involved in experimental design,…

We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language models (LLMs). SciQAG consists of a QA generator and a QA…

计算与语言 · 计算机科学 2024-07-11 Yuwei Wan , Yixuan Liu , Aswathy Ajith , Clara Grazian , Bram Hoex , Wenjie Zhang , Chunyu Kit , Tong Xie , Ian Foster