中文
相关论文

相关论文: OralGPT-Plus: Learning to Use Visual Tools via Rei…

200 篇论文

In recent years, symbolic regression has been of wide interest to provide an interpretable symbolic representation of potentially large data relationships. Initially circled to genetic algorithms, symbolic regression methods now include a…

机器学习 · 计算机科学 2022-02-10 Laure Crochepierre , Lydia Boudjeloud-Assala , Vincent Barbesant

We propose MARL-Rad, a multi-modal multi-agent reinforcement learning framework for radiology report generation that trains the entire agentic system on policy within its deployed radiology workflow. MARL-Rad addresses the limitation of…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Kaito Baba , Risa Kishikawa , Satoshi Kodera

Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techniques improve overall…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Anees Ur Rehman Hashmi , Numan Saeed , Christoph Lippert

Can reinforcement learning with hard, verifiable rewards teach a compact language model to reason about physics, or does it primarily learn to pattern-match toward correct answers? We study this question by training a 1.5B-parameter…

人工智能 · 计算机科学 2026-03-05 Tarjei Paule Hage , Markus J. Buehler

Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning through iterative visual interaction. Medical VLMs often rely…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Meng Lu , Yuxing Lu , Yuchen Zhuang , Megan Mullins , Yang Xie , Guanghua Xiao , Charles Fleming , Wenqi Shi , Xuan Wang

Medical image segmentation driven by free-text clinical instructions is a critical frontier in computer-aided diagnosis. However, existing multimodal and foundation models struggle with the semantic ambiguity of clinical reports and fail to…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Chenyu Xue , Yiran Liu , Mian Zhou , Jionglong Su , Zhixiang Lu

Open-Vocabulary Multimodal Emotion Recognition (OV-MER) aims to predict emotions without being constrained by predefined label spaces, thereby enabling fine-grained emotion understanding. Unlike traditional discriminative methods, OV-MER…

人机交互 · 计算机科学 2026-05-08 Zheng Lian , Fan Zhang , Lan Chen , Yazhou Zhang , Rui Liu , Jinyang Wu , Haoyu Chen , Xiaobai Li , Xiaojiang Peng , Bin He , Jianhua Tao

Recent self-supervised contrastive learning methods greatly benefit from the Siamese structure that aims to minimizing distances between positive pairs. These methods usually apply random data augmentation to input images, expecting the…

计算机视觉与模式识别 · 计算机科学 2023-05-16 Sheng Wang , Zixu Zhuang , Xi Ouyang , Lichi Zhang , Zheren Li , Chong Ma , Tianming Liu , Dinggang Shen , Qian Wang

Large language models (LLMs) constitute a breakthrough state-of-the-art Artificial Intelligence technology which is rapidly evolving and promises to aid in medical diagnosis. However, the correctness and the accuracy of their returns has…

计算与语言 · 计算机科学 2024-02-07 Dimitrios P. Panagoulias , Maria Virvou , George A. Tsihrintzis

Vision-language models have been extensively explored across a wide range of tasks, achieving satisfactory performance; however, their application in medical imaging remains underexplored. In this work, we propose a unified framework -…

图像与视频处理 · 电气工程与系统科学 2024-07-18 Khai Le-Duc , Ryan Zhang , Ngoc Son Nguyen , Tan-Hanh Pham , Anh Dao , Ba Hung Ngo , Anh Totti Nguyen , Truong-Son Hy

Learning general-purpose reasoning capabilities has long been a challenging problem in AI. Recent research in large language models (LLMs), such as DeepSeek-R1, has shown that reinforcement learning techniques like GRPO can enable…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Jiaer Xia , Yuhang Zang , Peng Gao , Sharon Li , Kaiyang Zhou

Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limited spatial awareness…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Qiushan Guo , Shalini De Mello , Hongxu Yin , Wonmin Byeon , Ka Chun Cheung , Yizhou Yu , Ping Luo , Sifei Liu

Metal artifacts in Dental CBCT severely obscure anatomical structures, hindering diagnosis. Current deep learning for Metal Artifact Reduction (MAR) faces limitations: supervised methods suffer from spectral blurring due to…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Zhi Li , Yaqi Wang , Bingtao Ma , Yifan Zhang , Huiyu Zhou , Shuai Wang

Retrieval-Augmented Language Models (RALMs) represent a classic paradigm where models enhance generative capabilities using external knowledge retrieved via a specialized module. Recent advancements in Agent techniques enable Large Language…

计算与语言 · 计算机科学 2025-05-28 Weiqi Wu , Xin Guan , Shen Huang , Yong Jiang , Pengjun Xie , Fei Huang , Jiuxin Cao , Hai Zhao , Jingren Zhou

Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation arises from training VLMs on static internet-scale…

机器人学 · 计算机科学 2026-04-01 Haowen Liu , Shaoxiong Yao , Haonan Chen , Jiawei Gao , Jiayuan Mao , Jia-Bin Huang , Yilun Du

The computer-assisted radiologic informative report is currently emerging in dental practice to facilitate dental care and reduce time consumption in manual panoramic radiographic interpretation. However, the amount of dental radiographs…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Amani Almalki , Longin Jan Latecki

Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promising direction is to…

计算机视觉与模式识别 · 计算机科学 2026-02-20 Qiucheng Wu , Jing Shi , Simon Jenni , Kushal Kafle , Tianyu Wang , Shiyu Chang , Handong Zhao

The impression is crucial for the referring physicians to grasp key information since it is concluded from the findings and reasoning of radiologists. To alleviate the workload of radiologists and reduce repetitive human labor in impression…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Jinpeng Hu , Zhihong Chen , Yang Liu , Xiang Wan , Tsung-Hui Chang

While reinforcement learning for large language model alignment has progressed rapidly in recent years, transferring these paradigms to high-stakes medical question answering reveals a fundamental paradigm mismatch. Reinforcement Learning…

Unified multimodal models (UMMs) have emerged as a powerful paradigm for seamlessly unifying text and image understanding and generation. However, prevailing evaluations treat these abilities in isolation, such that tasks with multimodal…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Yongyuan Liang , Wei Chow , Feng Li , Ziqiao Ma , Xiyao Wang , Jiageng Mao , Jiuhai Chen , Jiatao Gu , Yue Wang , Furong Huang