中文
相关论文

相关论文: VeriTrip: A Verifiable Benchmark for Travel Planni…

200 篇论文

Scientific and Technical Intelligence (S&TI) analysis requires verifying complex technical claims across rapidly growing literature, where existing approaches fail to bridge the verification gap between surface-level accuracy and deeper…

人工智能 · 计算机科学 2026-04-06 Yuntao Du , Minh Dinh , Kaiyuan Zhang , Ninghui Li

Agentic search, as a more autonomous and adaptive paradigm of retrieval augmentation, is driving the evolution of intelligent search systems. However, existing evaluation frameworks fail to align well with the goals of agentic search.…

计算与语言 · 计算机科学 2025-08-01 Yilong Xu , Xiang Long , Zhi Zheng , Jinhua Gao

Vision-and-Language Navigation (VLN) requires an embodied agent to navigate in a complex 3D environment according to natural language instructions. Recent progress in large language models (LLMs) has enabled language-driven navigation with…

机器人学 · 计算机科学 2026-01-27 Zijun Li , Shijie Li , Zhenxi Zhang , Bin Li , Shoujun Zhou

This research foregrounds general practices in travel demand research, emphasizing the need to change our ways. A critical barrier preventing travel demand literature from effectively informing policy is the volume of publications without…

机器学习 · 计算机科学 2024-07-16 Juan D. Caicedo , Carlos Guirado , Marta C. González , Joan L. Walker

Web agents powered by Large Language Models (LLMs) have demonstrated remarkable abilities in planning and executing multi-step interactions within complex web-based environments, fulfilling a wide range of web navigation tasks. Despite…

计算与语言 · 计算机科学 2024-02-26 Yang Deng , Xuan Zhang , Wenxuan Zhang , Yifei Yuan , See-Kiong Ng , Tat-Seng Chua

Agents powered by large language models (LLMs) are increasingly deployed in settings where communication shapes high-stakes decisions, making a principled understanding of strategic communication essential. Prior work largely studies either…

计算与语言 · 计算机科学 2026-02-03 Saaduddin Mahmud , Eugene Bagdasarian , Shlomo Zilberstein

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall…

Even when instructed to adhere to source material, language models often generate unsubstantiated content - a phenomenon known as "closed-domain hallucination." This risk is amplified in processes with multiple generative steps (MGS),…

计算与语言 · 计算机科学 2026-03-03 Dasha Metropolitansky , Jonathan Larson

Large language models (LLMs) with advanced cognitive capabilities are emerging as agents for various reasoning and planning tasks. Traditional evaluations often focus on specific reasoning or planning questions within controlled…

人工智能 · 计算机科学 2026-03-23 Tianlong Wang , Pinqiao Wang , Weili Shi , Sheng li

Large Language Models (LLMs) show promise as planners for embodied AI, but their stochastic nature lacks formal reasoning, preventing strict safety guarantees for physical deployment. Current approaches often rely on unreliable LLMs for…

人工智能 · 计算机科学 2026-04-30 Feiyu Wu , Xu Zheng , Yue Qu , Zhuocheng Wang , Zicheng Feng , Hui Li

Contemporary large language model (LLM)-based multi-agent systems exhibit systematic advantages in deep research tasks, which emphasize iterative, vertically structured information seeking. However, when confronted with wide search tasks…

多智能体系统 · 计算机科学 2026-02-03 Mingju Chen , Guibin Zhang , Heng Chang , Yuchen Guo , Shiji Zhou

The paper presents a framework of microservices-based architecture dedicated to enhancing the performance of real-time travel reservation systems using the power of predictive analytics. Traditional monolithic systems are bad at scaling and…

信息论 · 计算机科学 2024-12-23 Biman Barua , M. Shamim Kaiser

Human decision-making often involves constrained optimization. As LLM agents are deployed to assist with real-world tasks like travel planning, shopping, and scheduling, they must mirror this capability. We introduce COMPASS, a benchmark…

The rapid proliferation of multimodal misinformation presents significant challenges for automated fact-checking systems, especially when claims are ambiguous or lack sufficient context. We introduce RAMA, a novel retrieval-augmented…

计算与语言 · 计算机科学 2025-07-15 Shuo Yang , Zijian Yu , Zhenzhe Ying , Yuqin Dai , Guoqing Wang , Jun Lan , Jinfeng Xu , Jinze Li , Edith C. H. Ngai

Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this limitation, we propose Self-Verification and Self-Rectification (SVSR), a unified…

人工智能 · 计算机科学 2026-05-29 Zhe Qian , Nianbing Su , Zhonghua Wang , Hebei Li , Zhongxing Xu , Yueying Li , Fei Luo , Zhuohan Ouyang , Yanbiao Ma

Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, these sources can include unreliable or adversarial content, and the robustness of agents to adversarial…

人工智能 · 计算机科学 2026-03-03 Shrey Shah , Levent Ozgur

Travel planning serves as a critical task for long-horizon reasoning, exposing significant deficits in LLMs. However, existing benchmarks and evaluations primarily assess final plans in an end-to-end manner, which lacks interpretability and…

人工智能 · 计算机科学 2026-05-06 Bo-Wen Zhang , Jin Ye , Peng-Yu Hua , Jia-Wei Cao , Jie-Jing Shao , Yu-Feng Li , Lan-Zhe Guo

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomings limit the insight that the benchmarks provide into…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Sihao Lin , Zerui Li , Xunyi Zhao , Gengze Zhou , Liuyi Wang , Rong Wei , Rui Tang , Juncheng Li , Hanqing Wang , Jiangmiao Pang , Anton van den Hengel , Jiajun Liu , Qi Wu

Benchmarks are paramount for gauging progress in the domain of Mobile GUI Agents. In practical scenarios, users frequently fail to articulate precise directives containing full task details at the onset, and their expressions are typically…

Current validation methods often rely on recorded data and basic functional checks, which may not be sufficient to encompass the scenarios an autonomous vehicle might encounter. In addition, there is a growing need for complex scenarios…

机器人学 · 计算机科学 2024-02-08 Marc Kaufeld , Rainer Trauth , Johannes Betz