中文
相关论文

相关论文: OPGAgent: An Agent for Auditable Dental Panoramic …

200 篇论文

The unprecedented advancements in Multimodal Large Language Models (MLLMs) have demonstrated strong potential in interacting with humans through both language and visual inputs to perform downstream tasks such as visual question answering…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Wenjia Xu , Zijian Yu , Boyang Mu , Zhiwei Wei , Yuanben Zhang , Guangzuo Li , Jiuniu Wang , Mugen Peng

Open-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Zhigang Wang , Yifei Su , Chenhui Li , Dong Wang , Yan Huang , Bin Zhao , Xuelong Li

Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the right of the fridge" requires grounding semantic references,…

机器人学 · 计算机科学 2026-03-20 Swagat Padhan , Lakshya Jain , Bhavya Minesh Shah , Omkar Patil , Thao Nguyen , Nakul Gopalan

A central challenge in explainable AI, particularly in the visual domain, is producing explanations grounded in human-understandable concepts. To tackle this, we introduce OCEAN (Object-Centric Explananda via Agent Negotiation), a novel,…

人工智能 · 计算机科学 2025-09-30 Benjamin Teoh , Ben Glocker , Francesca Toni , Avinash Kori

Every day, countless surgeries are performed worldwide, each within the distinct settings of operating rooms (ORs) that vary not only in their setups but also in the personnel, tools, and equipment used. This inherent diversity poses a…

计算机视觉与模式识别 · 计算机科学 2024-04-11 Ege Özsoy , Chantal Pellegrini , Matthias Keicher , Nassir Navab

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to…

多智能体系统 · 计算机科学 2025-12-18 Philip R. Liu , Sparsh Bansal , Jimmy Dinh , Aditya Pawar , Ramani Satishkumar , Shail Desai , Neeraj Gupta , Xin Wang , Shu Hu

Large-scale single-cell and Perturb-seq investigations routinely involve clustering cells and subsequently annotating each cluster with Gene-Ontology (GO) terms to elucidate the underlying biological programs. However, both stages,…

Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world. However, current multi-modal LLMs are primarily confined to bi-modal…

Extracting implicit knowledge and logical reasoning abilities from large language models (LLMs) has consistently been a significant challenge. The advancement of multi-agent systems has further en-hanced the capabilities of LLMs. Inspired…

人工智能 · 计算机科学 2025-09-23 Hailong Yang , Mingxian Gu , Renhuo Zhao , Fuping Hu , Zhaohong Deng , Yitang Chen

We show that multi-agent systems guided by vision-language models (VLMs) improve end-to-end autonomous scientific discovery. By treating plots as verifiable checkpoints, a VLM-as-a-judge evaluates figures against dynamically generated…

计算与语言 · 计算机科学 2025-11-19 Kahaan Gandhi , Boris Bolliet , Inigo Zubeldia

We introduce GenAgent, unifying visual understanding and generation through an agentic multimodal model. Unlike unified models that face expensive training costs and understanding-generation trade-offs, GenAgent decouples these capabilities…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Kaixun Jiang , Yuzheng Wang , Junjie Zhou , Pandeng Li , Zhihang Liu , Chen-Wei Xie , Zhaoyu Chen , Yun Zheng , Wenqiang Zhang

Optical coherence tomography (OCT) is a commonly-used method of extracting high resolution retinal information. Moreover there is an increasing demand for the automated retinal layer segmentation which facilitates the retinal disease…

图像与视频处理 · 电气工程与系统科学 2020-09-30 Zeyu Fu , Yang Sun , Xiangyu Zhang , Scott Stainton , Shaun Barney , Jeffry Hogg , William Innes , Satnam Dlay

In this paper, we propose a framework that incorporates experts diagnostics and insights into the analysis of Optical Coherence Tomography (OCT) using multi-modal learning. To demonstrate the effectiveness of this approach, we create a…

图像与视频处理 · 电气工程与系统科学 2022-03-22 Y. Logan , K. Kokilepersaud , G. Kwon , G. AlRegib , C. Wykoff , H. Yu

Recent advances in representation learning often rely on holistic embeddings that entangle multiple semantic components, limiting interpretability and generalization. These issues are especially critical in medical imaging, where downstream…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Sifan Song , Siyeop Yoon , Pengfei Jin , Sekeun Kim , Matthew Tivnan , Yujin Oh , Runqi Meng , Ling Chen , Zhiliang Lyu , Dufan Wu , Ning Guo , Xiang Li , Quanzheng Li

This paper introduces \textsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, images, audio, and video. Unlike existing approaches that either build intricate workflows…

人工智能 · 计算机科学 2026-05-04 Bin Lei , Weitai Kang , Zijian Zhang , Winson Chen , Xi Xie , Shan Zuo , Mimi Xie , Ali Payani , Mingyi Hong , Yan Yan , Caiwen Ding

Fast surrogate models for expensive simulations are now essential across the sciences, yet they typically operate as black boxes. We present \texttt{GWAgent}, a large language model (LLM)-based workflow that constructs interpretable…

广义相对论与量子宇宙学 · 物理学 2026-05-13 Tousif Islam , Digvijay Wadekar , Tejaswi Venumadhav , Matias Zaldarriaga , Ajit Kumar Mehta , Javier Roulet , Barak Zackay

Recent agentic systems demonstrate that large language models can generate scientific visualizations from natural language. However, reliability remains a major limitation: systems may execute invalid operations, introduce subtle but…

人机交互 · 计算机科学 2026-03-27 Nathaniel Gorski , Shusen Liu , Bei Wang

The advancement in large language models (LLMs) and large vision models has fueled the rapid progress in multi-modal vision-language reasoning capabilities. However, existing vision-language models (VLMs) remain challenged by compositional…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Yichang Xu , Gaowen Liu , Ramana Rao Kompella , Sihao Hu , Fatih Ilhan , Selim Furkan Tekin , Zachary Yahn , Ling Liu

Ultra-wide optical coherence tomography angiography (UW-OCTA) is an emerging imaging technique that offers significant advantages over traditional OCTA by providing an exceptionally wide scanning range of up to 24 x 20 $mm^{2}$, covering…

图像与视频处理 · 电气工程与系统科学 2023-11-20 Hao Wei , Peilun Shi , Guitao Bai , Minqing Zhang , Shuangle Li , Wu Yuan

Patients take care of what their teeth will be like after the orthodontics. Orthodontists usually describe the expectation movement based on the original smile images, which is unconvincing. The growth of deep-learning generative models…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Feihong Shen , JIngjing Liu , Jianwen Lou , Haizhen Li , Bing Fang , Chenglong Ma , Jin Hao , Yang Feng , Youyi Zheng