中文
相关论文

相关论文: Surfer-H Meets Holo1: Cost-Efficient Web Agent Pow…

200 篇论文

The design space of agentic AI inference spans two extremes: frontier large language models (LLMs), typically hosted in the cloud and offering strong performance across a wide range of tasks at substantially high cost, and more…

多智能体系统 · 计算机科学 2026-05-29 Corrado Rainone , Davide Belli , Bence Major , Arash Behboodi

Vehicle-to-Vehicle technologies have enabled autonomous vehicles to share information to see through occlusions, greatly enhancing perception performance. Nevertheless, existing works all focused on homogeneous traffic where vehicles are…

计算机视觉与模式识别 · 计算机科学 2023-04-24 Hao Xiang , Runsheng Xu , Jiaqi Ma

Modern web scraping struggles with dynamic, interactive websites that require more than static HTML parsing. Current methods are often brittle and require manual customization for each site. To address this, we introduce Webscraper, a…

人工智能 · 计算机科学 2026-04-01 Guan-Lun Huang , Yuh-Jzer Joung

Vision-Language Navigation (VLN) systems are fundamentally constrained by partial observability, as an agent can only accumulate knowledge from locations it has personally visited. As multiple robots increasingly coexist in shared…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Qunchao Jin , Yiliao Song , Qi Wu

The rise of Large Language Models~(LLMs) revolutionizes information retrieval, allowing users to obtain required answers through complex instructions within conversations. However, publicly available services remain inadequate in addressing…

计算与语言 · 计算机科学 2025-05-14 Mingxu Tao , Bowen Tang , Mingxuan Ma , Yining Zhang , Hourun Li , Feifan Wen , Hao Ma , Jia Yang

Web agents such as Deep Research have demonstrated superhuman cognitive abilities, capable of solving highly challenging information-seeking problems. However, most research remains primarily text-centric, overlooking visual information in…

Large language model (LLM) agents have demonstrated impressive capabilities in utilizing external tools and knowledge to boost accuracy and reduce hallucinations. However, developing prompting techniques that enable LLM agents to…

Recent advancements in Large Language Models (LLMs) have largely focused on depth scaling, where a single agent solves long-horizon problems with multi-turn reasoning and tool use. However, as tasks grow broader, the key bottleneck shifts…

人工智能 · 计算机科学 2026-03-13 Zelai Xu , Zhexuan Xu , Ruize Zhang , Chunyang Zhu , Shi Yu , Weilin Liu , Quanlu Zhang , Wenbo Ding , Chao Yu , Yu Wang

Current machine algorithms for analysis of unstructured data typically show low accuracies due to the need for human-like intelligence. Conversely, though humans are much better than machine algorithms on analyzing unstructured data, they…

人机交互 · 计算机科学 2016-06-16 Koushik Sinha , Geetha Manjunath , Bidyut Gupta , Shahram Rahimi

Homing and navigation are fundamental behaviors in biological systems that enable agents to reliably reach a target under uncertainty. We present a Reinforcement Learning (RL) framework to model adaptive homing in continuous two-dimensional…

软凝聚态物质 · 物理学 2026-02-10 Riya Singh , Pratikshya Jena , Anish Kumar , Shradha Mishra

We present Adversarial Object Fusion (AdvOF), a novel attack framework targeting vision-and-language navigation (VLN) agents in service-oriented environments by generating adversarial 3D objects. While foundational models like Large…

密码学与安全 · 计算机科学 2025-05-30 Chunlong Xie , Jialing He , Shangwei Guo , Jiacheng Wang , Shudong Zhang , Tianwei Zhang , Tao Xiang

Enabling large language models to scale and reliably use hundreds of tools is critical for real-world applications, yet challenging due to the inefficiency and error accumulation inherent in flat tool-calling architectures. To address this,…

计算与语言 · 计算机科学 2026-04-14 Chengrui Huang , Junshuo Zhang , Zhiyuan Ma , Xikun Wang , Ximeng Wang , Menghua Jiang , Gang Zeng , Zhaobing Han , Shen Gao , Shuo Shang

Sources of complementary information are connected when we link user accounts belonging to the same user across different platforms or devices. The expanded information promotes the development of a wide range of applications, such as…

社会与信息网络 · 计算机科学 2022-01-11 Wei Chen , Weiqing Wang , Hongzhi Yin , Lei Zhao , Xiaofang Zhou

In this report, we introduce Falcon-H1, a new series of large language models (LLMs) featuring hybrid architecture designs optimized for both high performance and efficiency across diverse use cases. Unlike earlier Falcon models built…

Autonomous agents powered by large language models (LLMs) have the potential to significantly enhance human productivity by reasoning, using tools, and executing complex tasks in diverse environments. However, current approaches to…

The predominant approach for training web navigation agents is to gather human demonstrations for a set of popular websites and hand-written tasks, but it is becoming clear that human data is an inefficient resource. We develop a pipeline…

机器学习 · 计算机科学 2025-05-23 Brandon Trabucco , Gunnar Sigurdsson , Robinson Piramuthu , Ruslan Salakhutdinov

This paper presents a Large Language Model (LLM) based conversational agent system designed to enhance human-machine collaboration in Machine Learning Operations (MLOps). We introduce the Swarm Agent, an extensible architecture that…

Intelligent navigation-helper agents are critical as they can navigate users in unknown areas through environmental awareness and conversational ability, serving as potential accessibility tools for individuals with disabilities. In this…

计算与语言 · 计算机科学 2023-10-18 Yue Fan , Jing Gu , Kaizhi Zheng , Xin Eric Wang

Ever-increasing ubiquity of data and computational resources in the last decade have propelled a notable transition in the machine learning paradigm towards more distributed approaches. Such a transition seeks to not only tackle the…

分布式、并行与集群计算 · 计算机科学 2024-01-22 Ahmad Esmaeili , Zahra Ghorrati , Eric T. Matson

While Large Language Model (LLM) agents show great potential for automated UI navigation such as automated UI testing and AI assistants, their efficiency has been largely overlooked. Our motivating study reveals that inefficient UI…