中文
相关论文

相关论文: InfiniteScienceGym: An Unbounded, Procedurally-Gen…

200 篇论文

The powerful reasoning capabilities of Large Language Models (LLMs) in mathematics and coding, combined with their ability to automate complex tasks through agentic frameworks, present unprecedented opportunities for accelerating scientific…

人工智能 · 计算机科学 2025-05-27 Jiabin Tang , Lianghao Xia , Zhonghang Li , Chao Huang

Artificial intelligence is increasingly catalyzing scientific automation, with multimodal large language model (MLLM) agents evolving from lab assistants into self-driving lab operators. This transition imposes stringent safety requirements…

人工智能 · 计算机科学 2026-03-13 Qianpu Sun , Xiaowei Chi , Yuhan Rui , Ying Li , Kuangzhi Ge , Jiajun Li , Sirui Han , Shanghang Zhang

Large Language Models (LLMs) have demonstrated remarkable potential in advancing scientific knowledge and addressing complex challenges. In this work, we introduce OmniScience, a specialized large reasoning model for general science,…

I propose a method for tracking and assessing scientific progress using a prediction consensus algorithm designed for the purpose. The protocol obviates the need for centralized referees to generate scientific questions, gather predictions,…

物理与社会 · 物理学 2022-05-11 Ted C. Rogers

Recent years have witnessed a rapid development of deep generative models for creating synthetic media, such as images and videos. While the practical applications of these models in everyday tasks are enticing, it is crucial to assess the…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Mike Laszkiewicz , Imant Daunhawer , Julia E. Vogt , Asja Fischer , Johannes Lederer

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Andong Deng , Taojiannan Yang , Shoubin Yu , Lincoln Spencer , Mohit Bansal , Chen Chen , Serena Yeung-Levy , Xiaohan Wang

As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced…

The rapid evolution of artificial intelligence has led to expectations of transformative impact on science, yet current systems remain fundamentally limited in enabling genuine scientific discovery. This perspective contends that progress…

人工智能 · 计算机科学 2025-12-16 Karthik Duraisamy

This demo paper presents UnScientify, an interactive system designed to detect scientific uncertainty in scholarly full text. The system utilizes a weakly supervised technique that employs a fine-grained annotation scheme to identify…

计算与语言 · 计算机科学 2023-07-27 Panggih Kusuma Ningrum , Philipp Mayr , Iana Atanassova

Existing automatic scientific question generation studies mainly focus on single-document factoid QA, overlooking the inter-document reasoning crucial for scientific understanding. We present AIM-SciQA, an automated framework for generating…

计算与语言 · 计算机科学 2026-03-17 Seungmin Lee , Dongha Kim , Yuni Jeon , Junyoung Koh , Min Song

Recent advancements in Chain-of-Thoughts (CoT) and Program-of-Thoughts (PoT) methods have greatly enhanced language models' mathematical reasoning capabilities, facilitating their integration into instruction tuning datasets with LLMs.…

机器学习 · 计算机科学 2024-08-15 Bo-Wen Zhang , Yan Yan , Lin Li , Guang Liu

Evaluating language models fairly is increasingly difficult as static benchmarks risk contamination by training data, obscuring whether models truly reason or recall. We introduce BeyondBench, an evaluation framework using algorithmic…

计算与语言 · 计算机科学 2026-03-06 Gaurav Srivastava , Aafiya Hussain , Zhenyu Bi , Swastik Roy , Priya Pitre , Meng Lu , Morteza Ziyadi , Xuan Wang

As LLM agents are increasingly built around reusable skills, a central challenge is no longer only whether agents can use provided skills, but whether they can generate correct, reusable, and executable skills from repositories and…

As large language models (LLMs) increasingly exhibit human-like capabilities, a fundamental question emerges: How can we enable LLMs to learn the underlying patterns from limited examples in entirely novel environments and apply them…

计算与语言 · 计算机科学 2025-09-23 Brian S. Lin , Jiaxin Yuan , Zihan Zhou , Shouli Wang , Shuo Wang , Cunliang Kong , Qi Shi , Yuxuan Li , Liner Yang , Zhiyuan Liu , Maosong Sun

The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLMs). Many existing benchmarks suffer from fragmentation and…

Existing memory benchmarks for LLM agents evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavior without conscious retrieval. This gap is critical: effective assistants must automatically…

人工智能 · 计算机科学 2026-04-16 Chonghan Qin , Xiachong Feng , Weitao Ma , Xiaocheng Feng , Lingpeng Kong

With the rapid growth of scientific literature, scientific question answering (SciQA) has become increasingly critical for exploring and utilizing scientific knowledge. Retrieval-Augmented Generation (RAG) enhances LLMs by incorporating…

计算与语言 · 计算机科学 2026-01-26 Haotian Chen , Qingqing Long , Siyu Pu , Xiao Luo , Wei Ju , Meng Xiao , Yuanchun Zhou , Jianghua Zhao , Xuezhi Wang

Conducting literature reviews for scientific papers is essential for understanding research, its limitations, and building on existing work. It is a tedious task which makes an automatic literature review generator appealing. Unfortunately,…

Recent advances in Large Language Models (LLMs) have demonstrated strong potential in code generation, yet their effectiveness in quantum computing remains underexplored. This paper benchmarks LLMs for PennyLane-based quantum code…

Large language models (LLMs) are playing an increasingly important role in scientific research, yet there remains a lack of comprehensive benchmarks to evaluate the breadth and depth of scientific knowledge embedded in these models. To…

计算与语言 · 计算机科学 2025-10-08 Kehua Feng , Xinyi Shen , Weijie Wang , Xiang Zhuang , Yuqi Tang , Qiang Zhang , Keyan Ding
‹ 上一页 1 8 9 10 下一页 ›