中文
相关论文

相关论文: OralGPT-Plus: Learning to Use Visual Tools via Rei…

200 篇论文

Gastrointestinal diseases impose a growing global health burden, and endoscopy is a primary tool for early diagnosis. However, routine endoscopic image interpretation still suffers from missed lesions and limited efficiency. Although…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Peixi Peng , Housheng Xie , Yanling Wei , Guangcong Ruan , Xiaoyang Zou , Qian Cao , Yongjian Nian , Guoyan Zheng

Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely unexamined. Current benchmarks, mostly single-modality data, can't evaluate progressive…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Gui Wang , Zehao Zhong , YongSong Zhou , Yudong Li , Ende Wu , Wooi Ping Cheah , Rong Qu , Jianfeng Ren , Linlin Shen

Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (MLLMs) facilitate interactive medical image analysis, their application in dermatology…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yize Liu , Siyuan Yan , Ming Hu , Lie Ju , Xieji Li , Feilong Tang , Wei Feng , Zongyuan Ge

Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However, current reward models often fail to generalize, and…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Alexander Gambashidze , Li Pengyi , Matvey Skripkin , Andrey Galichin , Anton Gusarov , Konstantin Sobolev , Andrey Kuznetsov , Ivan Oseledets

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Sheng Lu , Hao Chen , Rui Yin , Juyan Ba , Yu Zhang , Yuanzhe Li

Medical image analysis increasingly relies on large vision-language models (VLMs), yet most systems remain single-pass black boxes that offer limited control over reasoning, safety, and spatial grounding. We propose R^4, an agentic…

Dental panoramic X-ray imaging is a popular diagnostic method owing to its very small dose of radiation. For an automated computer-aided diagnosis system in dental clinics, automatic detection and identification of individual teeth from…

计算机视觉与模式识别 · 计算机科学 2021-01-26 Minyoung Chung , Jusang Lee , Sanguk Park , Minkyung Lee , Chae Eun Lee , Jeongjin Lee , Yeong-Gil Shin

We present a framework for training large language models (LLMs) as diagnostic agents with reinforcement learning, enabling them to manage multi-turn interactive diagnostic processes, adaptively select examinations, and commit to final…

Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained…

图像与视频处理 · 电气工程与系统科学 2026-05-11 Xiang Li , Xueheng Li , Yu Wang , Xuanhua He , Zhangchi Hu , Weiwei Yu , Chengjun Xie

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although several multimodal reasoning models have been explored in the medical domain, most of them…

人工智能 · 计算机科学 2025-09-11 Ruiqi Wu , Yuang Yao , Tengfei Ma , Chenran Zhang , Na Su , Tao Zhou , Geng Chen , Wen Fan , Yi Zhou

Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are processed once to…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Shengyi Hua , Jianfeng Wu , Tianle Shen , Kangzhe Hu , Zhongzhen Huang , Shujuan Ni , Zhihong Zhang , Yuan Li , Zhe Wang , Xiaofan Zhang

Vision-Language Models (VLMs) have significantly advanced medical visual question answering, yet their performance in ultrasound remains suboptimal. In clinical practice, sonographers explicitly focus on lesion regions to formulate reports,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Yue Zhou , Erxuan Wu , Yikang Sun , Hongjoo Lee , Yuan Bi , Huixiong Xu , Nassir Navab , Zhongliang Jiang

Effectively retrieving, reasoning, and understanding multimodal information remains a critical challenge for agentic systems. Traditional Retrieval-augmented Generation (RAG) methods rely on linear interaction histories, which struggle to…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Qiuchen Wang , Shihang Wang , Yu Zeng , Qiang Zhang , Fanrui Zhang , Zhuoning Guo , Bosi Zhang , Wenxuan Huang , Lin Chen , Zehui Chen , Pengjun Xie , Ruixue Ding

Automated interpretation of medical images demands robust modeling of complex visual-semantic relationships while addressing annotation scarcity, label imbalance, and clinical plausibility constraints. We introduce MIRNet (Medical Image…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Shufeng Kong , Zijie Wang , Nuan Cui , Hao Tang , Yihan Meng , Yuanyuan Wei , Feifan Chen , Yingheng Wang , Zhuo Cai , Yaonan Wang , Yulong Zhang , Yuzheng Li , Zibin Zheng , Caihua Liu , Hao Liang

Although numerous strategies have recently been proposed to enhance the autonomous interaction capabilities of multimodal agents in graphical user interface (GUI), their reliability remains limited when faced with complex or out-of-domain…

计算与语言 · 计算机科学 2025-10-06 Pengzhou Cheng , Lingzhong Dong , Zeng Wu , Zongru Wu , Xiangru Tang , Chengwei Qin , Zhuosheng Zhang , Gongshen Liu

Recent advances in reasoning with large language models (LLMs)has shown remarkable reasoning capabilities in domains such as mathematics and coding, yet their application to clinical diagnosis remains underexplored. Here, we introduce…

计算与语言 · 计算机科学 2025-04-16 Wuyang Lan , Wenzheng Wang , Changwei Ji , Guoxing Yang , Yongbo Zhang , Xiaohong Liu , Song Wu , Guangyu Wang

In this work, we focused on deep learning image processing in the context of oral rare diseases, which pose challenges due to limited data availability. A crucial step involves teeth detection, segmentation and numbering in panoramic…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Hocine Kadi , Théo Sourget , Marzena Kawczynski , Sara Bendjama , Bruno Grollemund , Agnès Bloch-Zupan

Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidence. However, existing methods rely on exact-match binary rewards, which in clinical…

人工智能 · 计算机科学 2026-05-28 Yuwei Miao , Gen Li , Yunsheng Zeng , Xiandong Li , Yujin Wang , Siyu Chen , Luning Wang , Yunhao Qiao , Junfeng Wang , Jianwei Lv , Bo Yuan

To advance biomedical vison-language model capabilities through scaling up, fine-tuning, and instruction tuning, develop vision-language models with improved performance in handling long text, explore strategies to efficiently adopt vision…

人工智能 · 计算机科学 2025-05-26 Cheng Peng , Kai Zhang , Mengxian Lyu , Hongfang Liu , Lichao Sun , Yonghui Wu

With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achieve open-world visual perception remains an open question. In…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Chris Kelly , Luhui Hu , Bang Yang , Yu Tian , Deshun Yang , Cindy Yang , Zaoshan Huang , Zihao Li , Jiayin Hu , Yuexian Zou