English
Related papers

Related papers: OPGAgent: An Agent for Auditable Dental Panoramic …

200 papers

The unprecedented advancements in Multimodal Large Language Models (MLLMs) have demonstrated strong potential in interacting with humans through both language and visual inputs to perform downstream tasks such as visual question answering…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Wenjia Xu , Zijian Yu , Boyang Mu , Zhiwei Wei , Yuanben Zhang , Guangzuo Li , Jiuniu Wang , Mugen Peng

Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zhigang Wang , Yifei Su , Chenhui Li , Dong Wang , Yan Huang , Bin Zhao , Xuelong Li

Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the right of the fridge" requires grounding semantic references,…

Robotics · Computer Science 2026-03-20 Swagat Padhan , Lakshya Jain , Bhavya Minesh Shah , Omkar Patil , Thao Nguyen , Nakul Gopalan

A central challenge in explainable AI, particularly in the visual domain, is producing explanations grounded in human-understandable concepts. To tackle this, we introduce OCEAN (Object-Centric Explananda via Agent Negotiation), a novel,…

Artificial Intelligence · Computer Science 2025-09-30 Benjamin Teoh , Ben Glocker , Francesca Toni , Avinash Kori

Every day, countless surgeries are performed worldwide, each within the distinct settings of operating rooms (ORs) that vary not only in their setups but also in the personnel, tools, and equipment used. This inherent diversity poses a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Ege Özsoy , Chantal Pellegrini , Matthias Keicher , Nassir Navab

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to…

Multiagent Systems · Computer Science 2025-12-18 Philip R. Liu , Sparsh Bansal , Jimmy Dinh , Aditya Pawar , Ramani Satishkumar , Shail Desai , Neeraj Gupta , Xin Wang , Shu Hu

Large-scale single-cell and Perturb-seq investigations routinely involve clustering cells and subsequently annotating each cluster with Gene-Ontology (GO) terms to elucidate the underlying biological programs. However, both stages,…

Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However, current multi-modal LLMs are primarily confined to bi-modal…

Artificial Intelligence · Computer Science 2026-03-03 Xiaoxi Li , Wenxiang Jiao , Jiarui Jin , Shijian Wang , Guanting Dong , Jiajie Jin , Hao Wang , Yinuo Wang , Ji-Rong Wen , Yuan Lu , Zhicheng Dou

Extracting implicit knowledge and logical reasoning abilities from large language models (LLMs) has consistently been a significant challenge. The advancement of multi-agent systems has further en-hanced the capabilities of LLMs. Inspired…

Artificial Intelligence · Computer Science 2025-09-23 Hailong Yang , Mingxian Gu , Renhuo Zhao , Fuping Hu , Zhaohong Deng , Yitang Chen

We show that multi-agent systems guided by vision-language models (VLMs) improve end-to-end autonomous scientific discovery. By treating plots as verifiable checkpoints, a VLM-as-a-judge evaluates figures against dynamically generated…

Computation and Language · Computer Science 2025-11-19 Kahaan Gandhi , Boris Bolliet , Inigo Zubeldia

We introduce GenAgent, unifying visual understanding and generation through an agentic multimodal model. Unlike unified models that face expensive training costs and understanding-generation trade-offs, GenAgent decouples these capabilities…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Kaixun Jiang , Yuzheng Wang , Junjie Zhou , Pandeng Li , Zhihang Liu , Chen-Wei Xie , Zhaoyu Chen , Yun Zheng , Wenqiang Zhang

Optical coherence tomography (OCT) is a commonly-used method of extracting high resolution retinal information. Moreover there is an increasing demand for the automated retinal layer segmentation which facilitates the retinal disease…

Image and Video Processing · Electrical Eng. & Systems 2020-09-30 Zeyu Fu , Yang Sun , Xiangyu Zhang , Scott Stainton , Shaun Barney , Jeffry Hogg , William Innes , Satnam Dlay

In this paper, we propose a framework that incorporates experts diagnostics and insights into the analysis of Optical Coherence Tomography (OCT) using multi-modal learning. To demonstrate the effectiveness of this approach, we create a…

Image and Video Processing · Electrical Eng. & Systems 2022-03-22 Y. Logan , K. Kokilepersaud , G. Kwon , G. AlRegib , C. Wykoff , H. Yu

Recent advances in representation learning often rely on holistic embeddings that entangle multiple semantic components, limiting interpretability and generalization. These issues are especially critical in medical imaging, where downstream…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Sifan Song , Siyeop Yoon , Pengfei Jin , Sekeun Kim , Matthew Tivnan , Yujin Oh , Runqi Meng , Ling Chen , Zhiliang Lyu , Dufan Wu , Ning Guo , Xiang Li , Quanzheng Li

This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlike existing approaches that either build intricate workflows…

Artificial Intelligence · Computer Science 2026-05-04 Bin Lei , Weitai Kang , Zijian Zhang , Winson Chen , Xi Xie , Shan Zuo , Mimi Xie , Ali Payani , Mingyi Hong , Yan Yan , Caiwen Ding

Fast surrogate models for expensive simulations are now essential across the sciences, yet they typically operate as black boxes. We present \texttt{GWAgent}, a large language model (LLM)-based workflow that constructs interpretable…

General Relativity and Quantum Cosmology · Physics 2026-05-13 Tousif Islam , Digvijay Wadekar , Tejaswi Venumadhav , Matias Zaldarriaga , Ajit Kumar Mehta , Javier Roulet , Barak Zackay

Recent agentic systems demonstrate that large language models can generate scientific visualizations from natural language. However, reliability remains a major limitation: systems may execute invalid operations, introduce subtle but…

Human-Computer Interaction · Computer Science 2026-03-27 Nathaniel Gorski , Shusen Liu , Bei Wang

The advancement in large language models (LLMs) and large vision models has fueled the rapid progress in multi-modal vision-language reasoning capabilities. However, existing vision-language models (VLMs) remain challenged by compositional…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Yichang Xu , Gaowen Liu , Ramana Rao Kompella , Sihao Hu , Fatih Ilhan , Selim Furkan Tekin , Zachary Yahn , Ling Liu

Ultra-wide optical coherence tomography angiography (UW-OCTA) is an emerging imaging technique that offers significant advantages over traditional OCTA by providing an exceptionally wide scanning range of up to 24 x 20 $mm^{2}$, covering…

Image and Video Processing · Electrical Eng. & Systems 2023-11-20 Hao Wei , Peilun Shi , Guitao Bai , Minqing Zhang , Shuangle Li , Wu Yuan

Patients take care of what their teeth will be like after the orthodontics. Orthodontists usually describe the expectation movement based on the original smile images, which is unconvincing. The growth of deep-learning generative models…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Feihong Shen , JIngjing Liu , Jianwen Lou , Haizhen Li , Bing Fang , Chenglong Ma , Jin Hao , Yang Feng , Youyi Zheng