English
Related papers

Related papers: FrontierScience: Evaluating AI's Ability to Perfor…

200 papers

In this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we evaluate (i) capabilities for in-context identification and…

Computation and Language · Computer Science 2024-11-01 Andrew M. Bean , Simi Hellsten , Harry Mayne , Jabez Magomere , Ethan A. Chi , Ryan Chi , Scott A. Hale , Hannah Rose Kirk

We present FedScale, a federated learning (FL) benchmarking suite with realistic datasets and a scalable runtime to enable reproducible FL research. FedScale datasets encompass a wide range of critical FL tasks, ranging from image…

Machine Learning · Computer Science 2022-06-22 Fan Lai , Yinwei Dai , Sanjay S. Singapuram , Jiachen Liu , Xiangfeng Zhu , Harsha V. Madhyastha , Mosharaf Chowdhury

Recent advancements in artificial intelligence (AI), particularly in large language models (LLMs) such as OpenAI-o1 and DeepSeek-R1, have demonstrated remarkable capabilities in complex domains such as logical reasoning and experimental…

Recent advances in large language models have enabled the emergence of AI scientists that aim to autonomously analyze biological data and assist scientific discovery. Despite rapid progress, it remains unclear to what extent these systems…

Artificial Intelligence · Computer Science 2026-01-21 Erpai Luo , Jinmeng Jia , Yifan Xiong , Xiangyu Li , Xiaobo Guo , Baoqi Yu , Minsheng Hao , Lei Wei , Xuegong Zhang

Scientific and technological frontiers advance through punctuated dynamics, yet the principles governing these dynamics remain poorly understood. Here we collect and analyze datasets tracking the evolution of frontiers across 9 different…

Physics and Society · Physics 2026-05-19 Yian Yin , Dashun Wang

While the exploration for embodied AI has spanned multiple decades, it remains a persistent challenge to endow agents with human-level intelligence, including perception, learning, reasoning, decision-making, control, and generalization…

Robotics · Computer Science 2024-02-07 Zhiyuan Xu , Kun Wu , Junjie Wen , Jinming Li , Ning Liu , Zhengping Che , Jian Tang

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and…

Artificial Intelligence · Computer Science 2025-05-27 Xinyu Zhang , Yuxuan Dong , Yanrui Wu , Jiaxing Huang , Chengyou Jia , Basura Fernando , Mike Zheng Shou , Lingling Zhang , Jun Liu

As automated reasoning systems advance rapidly, there is a growing need for research-level formal mathematical problems to accurately evaluate their capabilities. To address this, we present Formal Conjectures, an evolving benchmark of…

While Large Language Models (LLMs) have demonstrated proficiency in Deep Research or Wide Search, their capacity to solve highly complex questions-those requiring long-horizon planning, massive evidence gathering, and synthesis across…

Computation and Language · Computer Science 2026-03-04 Yubo Dong , Nianhao You , Yuxuan Hou , Zixun Sun , Yue Zhang , Liang Zhang , Siyuan Zhao , Hehe Fan

Artificial intelligence (AI) has significantly advanced Earth sciences, yet its full potential in to comprehensively modeling Earth's complex dynamics remains unrealized. Geoscience foundation models (GFMs) emerge as a paradigm-shifting…

Artificial Intelligence · Computer Science 2024-11-13 Hao Zhang , Jin-Jian Xu , Hong-Wei Cui , Lin Li , Yaowen Yang , Chao-Sheng Tang , Niklas Boers

Formalising informal mathematical reasoning into formally verifiable code is a significant challenge for large language models. In scientific fields such as physics, domain-specific machinery (\textit{e.g.} Dirac notation, vector calculus)…

Artificial Intelligence · Computer Science 2026-04-28 Jordan Meadows , Lan Zhang , Andre Freitas

Keeping track of scientific challenges, advances and emerging directions is a fundamental part of research. However, researchers face a flood of papers that hinders discovery of important knowledge. In biomedicine, this directly impacts…

In this position paper, I argue that standardized tests for elementary science such as SAT or Regents tests are not very good benchmarks for measuring the progress of artificial intelligence systems in understanding basic science. The…

Artificial Intelligence · Computer Science 2015-10-20 Ernest Davis

Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, climate adaptation and environmental protection. Although…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yushuo Zheng , Zicheng Zhang , Huiyu Duan , Chunyi Li , Zijian Chen , Ziheng Jia , Yue Shi , Ke Gu , Xiongkuo Min , Guangtao Zhai

Large Language Models (LLMs) have recently achieved impressive performance in math and reasoning benchmarks. However, they often struggle with logic problems and puzzles that are relatively easy for humans. To further investigate this, we…

Artificial Intelligence · Computer Science 2025-09-16 Nasim Borazjanizadeh , Roei Herzig , Trevor Darrell , Rogerio Feris , Leonid Karlinsky

Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical…

Large language models (LLMs) have shown promise in automating scientific hypothesis generation, yet existing approaches primarily yield coarse-grained hypotheses lacking critical methodological and experimental details. We introduce and…

Computation and Language · Computer Science 2025-10-28 Zonglin Yang , Wanhao Liu , Ben Gao , Yujie Liu , Wei Li , Tong Xie , Lidong Bing , Wanli Ouyang , Erik Cambria , Dongzhan Zhou

Many real-world problems require the combined application of multiple reasoning abilities employing suitable abstractions, commonsense knowledge, and creative synthesis of problem-solving strategies. To help advance AI systems towards such…

Computation and Language · Computer Science 2021-12-22 Ashwin Kalyan , Abhinav Kumar , Arjun Chandrasekaran , Ashish Sabharwal , Peter Clark

International institutions may have an important role to play in ensuring advanced AI systems benefit humanity. International collaborations can unlock AI's ability to further sustainable development, and coordination of regulatory efforts…

The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific knowledge. A core obstacle is the lack of established…

‹ Prev 1 8 9 10 Next ›