中文
相关论文

相关论文: MM-WebAgent: A Hierarchical Multimodal Web Agent f…

200 篇论文

Despite seemingly performant web agents on the task-completion benchmarks, most existing methods evaluate the agents based on a presupposition: the web navigation task consists of linear sequence of actions with an end state that marks task…

人工智能 · 计算机科学 2024-10-28 Revanth Gangi Reddy , Sagnik Mukherjee , Jeonghwan Kim , Zhenhailong Wang , Dilek Hakkani-Tur , Heng Ji

With the recent fast development of generative models, instruction-based image editing has shown great potential in generating high-quality images. However, the quality of editing highly depends on carefully designed instructions, placing…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Mingde Yao , Zhiyuan You , King-Man Tam , Menglu Wang , Tianfan Xue

Web 3.0 represents the next generation of the Internet, which is widely recognized as a decentralized ecosystem that focuses on value expression and data ownership. By leveraging blockchain and artificial intelligence technologies, Web 3.0…

人工智能 · 计算机科学 2025-10-07 Jinbo Wen , Jiawen Kang , Linfeng Zhang , Xiaoying Tang , Jianhang Tang , Yang Zhang , Zhaohui Yang , Dusit Niyato

Building effective clinical decision support systems requires the synthesis of complex heterogeneous multimodal data. Such modalities include temporal electronic health records data, medical images, radiology reports, and clinical notes.…

人工智能 · 计算机科学 2026-05-12 Baraa Al Jorf , Farah E. Shamout

The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous,…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Aniket Rege , Arka Sadhu , Yuliang Li , Kejie Li , Ramya Korlakai Vinayak , Yuning Chai , Yong Jae Lee , Hyo Jin Kim

In automated web testing, generating test scripts from natural language task descriptions is crucial for enhancing the test generation process. This activity involves creating the correct sequences of actions to form test scripts for future…

软件工程 · 计算机科学 2025-09-12 Duy Cao , Phu Nguyen , Vy Le , Tien N. Nguyen , Vu Nguyen

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Keda Tao , Wenjie Du , Bohan Yu , Weiqiang Wang , Jian Liu , Huan Wang

Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack the ability to maintain consistency with the provided…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Zhipeng Huang , Shaobin Zhuang , Canmiao Fu , Binxin Yang , Ying Zhang , Chong Sun , Zhizheng Zhang , Yali Wang , Chen Li , Zheng-Jun Zha

The increasing use of synthetic media, particularly deepfakes, is an emerging challenge for digital content verification. Although recent studies use both audio and visual information, most integrate these cues within a single model, which…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Sayeem Been Zaman , Wasimul Karim , Arefin Ittesafun Abian , Reem E. Mohamed , Md Rafiqul Islam , Asif Karim , Sami Azam

Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its…

人工智能 · 计算机科学 2023-03-09 Yihan Cao , Siyu Li , Yixin Liu , Zhiling Yan , Yutong Dai , Philip S. Yu , Lichao Sun

Climate science demands automated workflows to transform comprehensive questions into data-driven statements across massive, heterogeneous datasets. However, generic LLM agents and static scripting pipelines lack climate-specific context…

机器学习 · 计算机科学 2025-11-26 Hyeonjae Kim , Chenyue Li , Wen Deng , Mengxi Jin , Wen Huang , Mengqian Lu , Binhang Yuan

Document Question Answering (DocQA) is a very common task. Existing methods using Large Language Models (LLMs) or Large Vision Language Models (LVLMs) and Retrieval Augmented Generation (RAG) often prioritize information from a single…

机器学习 · 计算机科学 2025-03-19 Siwei Han , Peng Xia , Ruiyi Zhang , Tong Sun , Yun Li , Hongtu Zhu , Huaxiu Yao

Recently, Agentic AI has become an increasingly popular research field. However, we argue that current agent research practices lack standardization and scientific rigor, making it hard to conduct fair comparisons among methods. As a…

We present CreAgentive, an agent workflow driven multi-category creative generation engine that addresses four key limitations of contemporary large language models in writing stories, drama and other categories of creatives: restricted…

计算与语言 · 计算机科学 2025-10-01 Yuyang Cheng , Linyue Cai , Changwei Peng , Yumiao Xu , Rongfang Bie , Yong Zhao

Existing web agents typically initiate exploration from the root URL, which is inefficient for complex websites with deep hierarchical structures. Without a global view of the website's structure, agents frequently fall into navigation…

计算与语言 · 计算机科学 2026-04-24 Weixi Tong , Yifeng Di , Tianyi Zhang

Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, managing the heterogeneous information and high token costs associated with multimodal inputs…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yifan Du , Zikang Liu , Jinbiao Peng , Jie Wu , Junyi Li , Jinyang Li , Wayne Xin Zhao , Ji-Rong Wen

Contemporary multi-agent systems encounter persistent challenges in cross-platform interoperability, dynamic task scheduling, and efficient resource sharing. Agents with heterogeneous implementations often lack standardized interfaces;…

人工智能 · 计算机科学 2025-07-08 Yuyang Cheng , Yumiao Xu , Chaojia Yu , Yong Zhao

Automating the transformation of user interface (UI) designs into front-end code holds significant promise for accelerating software development and democratizing design workflows. While multimodal large language models (MLLMs) can…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Yilei Jiang , Yaozhi Zheng , Yuxuan Wan , Jiaming Han , Qunzhong Wang , Michael R. Lyu , Xiangyu Yue

Recent progress in multimodal graph neural networks has demonstrated that augmenting atomic XYZ geometries with textual chemical descriptors can enhance predictive accuracy across a range of electronic and thermodynamic properties. However,…

多智能体系统 · 计算机科学 2025-06-27 Can Polat , Mehmet Tuncel , Mustafa Kurban , Erchin Serpedin , Hasan Kurban

Automated content-aware layout generation -- the task of arranging visual elements such as text, logos, and underlays on a background canvas -- remains a fundamental yet under-explored problem in intelligent design systems. While recent…

‹ 上一页 1 8 9 10 下一页 ›