中文
相关论文

相关论文: Wiki Live Challenge: Challenging Deep Research Age…

200 篇论文

With the rapid growth of scientific literature, scientific question answering (SciQA) has become increasingly critical for exploring and utilizing scientific knowledge. Retrieval-Augmented Generation (RAG) enhances LLMs by incorporating…

计算与语言 · 计算机科学 2026-01-26 Haotian Chen , Qingqing Long , Siyu Pu , Xiao Luo , Wei Ju , Meng Xiao , Yuanchun Zhou , Jianghua Zhao , Xuezhi Wang

Multi-entity question answering (MEQA) poses significant challenges for large language models (LLMs), which often struggle to consolidate scattered information across multiple documents. An example question might be "What is the…

计算与语言 · 计算机科学 2025-03-07 Teng Lin , Yizhang Zhu , Yuyu Luo , Nan Tang

Deep-research agents are capable of executing multi-step web exploration, targeted retrieval, and sophisticated question answering. Despite their powerful capabilities, deep-research agents face two critical bottlenecks: (1) the lack of…

人工智能 · 计算机科学 2026-03-03 Tongzhou Wu , Yuhao Wang , Xinyu Ma , Xiuqiang He , Shuaiqiang Wang , Dawei Yin , Xiangyu Zhao

Wikipedia is edited by volunteer editors around the world. Considering the large amount of existing content (e.g. over 5M articles in English Wikipedia), deciding what to edit next can be difficult, both for experienced users that usually…

信息检索 · 计算机科学 2020-09-25 Oleksii Moskalenko , Denis Parra , Diego Saez-Trumper

We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs). In LLM-Wikirace, models must efficiently navigate Wikipedia hyperlinks step by step to reach a target page from…

人工智能 · 计算机科学 2026-02-24 Juliusz Ziomek , William Bankes , Lorenz Wolf , Shyam Sundhar Ramesh , Xiaohang Tang , Ilija Bogunovic

Millions of people irrespective of socioeconomic and demographic backgrounds, depend on Wikipedia articles everyday for keeping themselves informed regarding popular as well as obscure topics. Articles have been categorized by editors into…

社会与信息网络 · 计算机科学 2020-10-15 Bhanu Prakash Reddy , Sasi Bhusan , Soumya Sarkar , Animesh Mukherjee

Wikipedia plays a crucial role in the integrity of the Web. This work analyzes the reliability of this global encyclopedia through the lens of its references. We operationalize the notion of reference quality by defining reference need…

As large language models (LLMs) grow larger and more sophisticated, assessing their "reasoning" capabilities in natural language grows more challenging. Recent question answering (QA) benchmarks that attempt to assess reasoning are often…

计算与语言 · 计算机科学 2022-12-01 Matthew Ho , Aditya Sharma , Justin Chang , Michael Saxon , Sharon Levy , Yujie Lu , William Yang Wang

Retrieval-augmented generation (RAG) methods are viable solutions for addressing the static memory limits of pre-trained language models. Nevertheless, encountering conflicting sources of information within the retrieval context is an…

计算与语言 · 计算机科学 2025-06-05 Quang Hieu Pham , Hoang Ngo , Anh Tuan Luu , Dat Quoc Nguyen

Despite significant advances in large language models (LLMs), their knowledge memorization capabilities remain underexplored, due to the lack of standardized and high-quality test ground. In this paper, we introduce a novel, real-world and…

计算与语言 · 计算机科学 2025-05-20 Yuwei Zhang , Wenhao Yu , Shangbin Feng , Yifan Zhu , Letian Peng , Jayanth Srinivasa , Gaowen Liu , Jingbo Shang

Deep research agents autonomously conduct open-ended investigations, integrating complex information retrieval with multi-step reasoning across diverse sources to solve real-world problems. To sustain this capability on long-horizon tasks,…

计算与语言 · 计算机科学 2026-03-31 Bin Zhu , Qianghuai Jia , Tian Lan , Junyang Ren , Feng Gu , Feihu Jiang , Longyue Wang , Zhao Xu , Weihua Luo

We present a long-horizon, hierarchical deep research (DR) agent designed for complex materials and device discovery problems that exceed the scope of existing Machine Learning (ML) surrogates and closed-source commercial agents. Our…

机器学习 · 计算机科学 2025-12-04 Rui Ding , Rodrigo Pires Ferreira , Yuxin Chen , Junhong Chen

In knowledge-intensive tasks such as open-domain question answering (OpenQA), large language models (LLMs) often struggle to generate factual answers, relying solely on their internal (parametric) knowledge. To address this limitation,…

计算与语言 · 计算机科学 2025-04-29 Jinming Nian , Zhiyuan Peng , Qifan Wang , Yi Fang

Wikipedia, a paradigmatic example of online knowledge space is organized in a collaborative, bottom-up way with voluntary contributions, yet it maintains a level of reliability comparable to that of traditional encyclopedias. The lack of…

物理与社会 · 物理学 2021-05-24 Fumiko Ogushi , János Kertész , Kimmo Kaski , Takashi Shimada

Wikipedia is the world's largest online encyclopedia, but maintaining article quality through collaboration is challenging. Wikipedia designed a quality scale, but with such a manual assessment process, many articles remain unassessed. We…

计算与语言 · 计算机科学 2023-10-04 Pedro Miguel Moás , Carla Teixeira Lopes

Despite the integration of search tools, Deep Search Agents often suffer from a misalignment between reasoning-driven queries and the underlying web indexing structures. Existing frameworks treat the search engine as a static utility,…

机器学习 · 计算机科学 2026-03-10 Zixuan Yu , Zhenheng Tang , Tongliang Liu , Chengqi Zhang , Xiaowen Chu , Bo Han

Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but methods are only tested on small-scale or synthetic edit…

计算与语言 · 计算机科学 2025-09-23 Lukas Thede , Karsten Roth , Matthias Bethge , Zeynep Akata , Tom Hartvigsen

Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these strides, evaluating…

While Large Language Models (LLMs) have demonstrated proficiency in Deep Research or Wide Search, their capacity to solve highly complex questions-those requiring long-horizon planning, massive evidence gathering, and synthesis across…

计算与语言 · 计算机科学 2026-03-04 Yubo Dong , Nianhao You , Yuxuan Hou , Zixun Sun , Yue Zhang , Liang Zhang , Siyuan Zhao , Hehe Fan

We introduce LongDA, a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. In contrast to existing benchmarks that assume well-specified schemas and inputs, LongDA targets real-world…

数字图书馆 · 计算机科学 2026-01-13 Yiyang Li , Zheyuan Zhang , Tianyi Ma , Zehong Wang , Keerthiram Murugesan , Chuxu Zhang , Yanfang Ye