English
Related papers

Related papers: A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline…

200 papers

REST APIs play important roles in enriching the action space of web agents, yet most API-based agents rely on curated and uniform toolsets that do not reflect the complexity of real-world APIs. Building tool-using agents for arbitrary…

Computation and Language · Computer Science 2025-06-26 Xinyi Ni , Haonan Jian , Qiuyang Wang , Vedanshi Chetan Shah , Pengyu Hong

Imagine decision-makers uploading data and, within minutes, receiving clear, actionable insights delivered straight to their fingertips. That is the promise of the AI Data Scientist, an autonomous Agent powered by large language models…

Artificial Intelligence · Computer Science 2025-08-26 Farkhad Akimov , Munachiso Samuel Nwadike , Zangir Iklassov , Martin Takáč

The integration of generative artificial intelligence (AI) into architectural design has advanced significantly, enabling the generation of text, images, and 3D models. However, prior AI applications lack support for text-to-parametric…

Human-Computer Interaction · Computer Science 2025-05-20 Guangxi Feng , Wei Yan

Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing document agents mainly transform papers into static artifacts such as summaries,…

Computation and Language · Computer Science 2026-03-31 Dasen Dai , Biao Wu , Meng Fang , Wenhao Wang

The era of the Internet of Agents (IoA) is taking shape: LLM agents are expected to fulfill user goals by orchestrating fast-growing populations of Model Context Protocol (MCP) servers, Agent-to-Agent (A2A) endpoints, reusable skills, and…

Artificial Intelligence · Computer Science 2026-05-29 Wei Zheng , Yang Yan , Yiyang Shao , Jinyang Li , Zeze Chang , Yukuang Jia , Qiming Mao , Chihyung Wang , Jingbin Zhou

Synthesizing informative commercial reports from massive and noisy web sources is critical for high-stakes business decisions. Although current deep research agents achieve notable progress, their reports still remain limited in terms of…

Computation and Language · Computer Science 2026-01-09 Mingyue Cheng , Daoyu Wang , Qi Liu , Shuo Yu , Xiaoyu Tao , Yuqian Wang , Chengzhong Chu , Yu Duan , Mingkang Long , Enhong Chen

The transition from optical identification of 2D quantum materials to practical device fabrication requires dynamic reasoning beyond the detection accuracy. While recent domain-specific Multimodal Large Language Models (MLLMs) successfully…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Sankalp Pandey , Xuan-Bac Nguyen , Hoang-Quan Nguyen , Tim Faltermeier , Nicholas Borys , Hugh Churchill , Khoa Luu

Vision-and-language navigation requires an agent to navigate through a real 3D environment following natural language instructions. Despite significant advances, few previous works are able to fully utilize the strong correspondence between…

Computer Vision and Pattern Recognition · Computer Science 2020-10-06 Yicong Hong , Cristian Rodriguez-Opazo , Qi Wu , Stephen Gould

Meta-analysis is a systematic research methodology that synthesizes data from multiple existing studies to derive comprehensive conclusions. This approach not only mitigates limitations inherent in individual studies but also facilitates…

Artificial Intelligence · Computer Science 2026-01-22 Wanghan Xu , Wenlong Zhang , Fenghua Ling , Ben Fei , Yusong Hu , Runmin Ma , Bo Zhang , Fangxuan Ren , Jintai Lin , Wanli Ouyang , Lei Bai

LLMs are increasingly deployed as agents, systems capable of planning, reasoning, and dynamically calling external tools. However, in visual reasoning, prior approaches largely remain limited by predefined workflows and static toolsets. In…

Computation and Language · Computer Science 2025-08-28 Shitian Zhao , Haoquan Zhang , Shaoheng Lin , Ming Li , Qilong Wu , Kaipeng Zhang , Chen Wei

Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still struggle with spatial reasoning, persistent state, long-horizon consistency, and causal…

Data products enable end users to gain greater insights about their data by providing supporting assets, such as example question-SQL pairs which can be answered using the data or views over the database tables. However, producing useful…

Artificial Intelligence · Computer Science 2026-03-12 Priyadarshini Tamilselvan , Gregory Bramble , Sola Shirai , Ken C. L. Wong , Faisal Chowdhury , Horst Samulowitz

Existing work on vision and language navigation mainly relies on navigation-related losses to establish the connection between vision and language modalities, neglecting aspects of helping the navigation agent build a deep understanding of…

Computation and Language · Computer Science 2024-02-06 Yue Zhang , Quan Guo , Parisa Kordjamshidi

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework, where the input and output of the system are multimodal (i.e., audio and visual speech). With the proposed AV2AV, two key…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Jeongsoo Choi , Se Jin Park , Minsu Kim , Yong Man Ro

Computational thematic analysis is rapidly emerging as a method of using large text corpora to understand the lived experience of people across the continuum of health care: patients, practitioners, and everyone in between. However, many…

Human-Computer Interaction · Computer Science 2024-12-20 Luka Ugaya Mazza , Plinio Morita , James R. Wallace

Machine Learning (ML) research is spread through academic papers featuring rich multimodal content, including text, diagrams, and tabular results. However, translating these multimodal elements into executable code remains a challenging and…

Software Engineering · Computer Science 2025-05-27 Zijie Lin , Yiqing Shen , Qilin Cai , He Sun , Jinrui Zhou , Mingjun Xiao

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical…

Software Engineering · Computer Science 2026-04-02 Zehai He , Wenyi Hong , Zhen Yang , Ziyang Pan , Mingdao Liu , Xiaotao Gu , Jie Tang

There are different goals for literature research, from understanding an unfamiliar topic to generate hypothesis for the next research project. The nature of literature research also varies according to user's familiarity level of the…

Human-Computer Interaction · Computer Science 2026-03-25 Zefei Xie , Yuhan Guo , Kai Xu

Recent advances in AI enable the automatic generation of visualizations directly from textual prompts using agentic workflows. However, visualizations produced via one-shot generative methods often suffer from insufficient quality,…

Human-Computer Interaction · Computer Science 2026-03-19 Roxana Bujack , Li-Ta Lo , Ethan Stam , Ayan Biswas , David Rogers

This paper proposes a visual analytics framework that addresses the complex user interactions required through a command-line interface to run analyses in distributed data analysis systems. The visual analytics framework facilitates the…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-11-12 Abdullah-Al-Raihan Nayeem , Mohammed Elshambakey , Todd Dobbs , Huikyo Lee , Daniel Crichton , Yimin Zhu , Chanachok Chokwitthaya , William J. Tolone , Isaac Cho