中文
相关论文

相关论文: VoxLogicA UI: Supporting Declarative Medical Image…

200 篇论文

The Medico 2025 challenge addresses Visual Question Answering (VQA) for Gastrointestinal (GI) imaging, organized as part of the MediaEval task series. The challenge focuses on developing Explainable Artificial Intelligence (XAI) models that…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Sushant Gautam , Vajira Thambawita , Michael Riegler , Pål Halvorsen , Steven Hicks

The main aim of the work presented here is to contribute to computer science advances in the multimodal usability area, in-as-much as it addresses one of the major issues relating to the generation of effective oral system messages: how to…

人机交互 · 计算机科学 2007-08-28 Suzanne Kieffer , Noëlle Carbonell

This paper presents preliminary results in the definition of a comprehensive benchmark framework designed to systematically evaluate spatial reasoning capabilities in neural networks, with a particular focus on morphological properties such…

机器学习 · 计算机科学 2025-08-19 Manuela Imbriani , Gina Belmonte , Mieke Massink , Alessandro Tofani , Vincenzo Ciancia

Interpretability of modern visual models is crucial, particularly in high-stakes applications. However, existing interpretability methods typically suffer from either reliance on white-box model access or insufficient quantitative rigor. To…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Chenchen Zhao , Muxi Chen , Qiang Xu

In order to be able to use artificial intelligence (AI) in medicine without scepticism and to recognise and assess its growing potential, a basic understanding of this topic is necessary among current and future medical staff. Under the…

图像与视频处理 · 电气工程与系统科学 2022-08-16 Hanna Siebert , Marian Himstedt , Mattias Heinrich

Document Visual Question Answering (DocVQA) requires models to jointly understand textual semantics, spatial layout, and visual features. Current methods struggle with explicit spatial relationship modeling, inefficiency with…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Ahmad Mohammadshirazi , Pinaki Prasad Guha Neogi , Dheeraj Kulshrestha , Rajiv Ramnath

Virtual reality simulation has become a popular approach for training and assessing medical students. It offers diverse scenarios, realistic visuals, and quantitative performance metrics for objective evaluation. However, creating these…

软件工程 · 计算机科学 2023-11-27 Vladimir Poliakov , Dzmitry Tsetserukou , Emmanuel Vander Poorten

Artificial Intelligence (AI) is one of the major technological advancements of this century, bearing incredible potential for users through AI-powered applications and tools in numerous domains. Being often black-box (i.e., its…

Remarkable success of modern image-based AI methods and the resulting interest in their applications in critical decision-making processes has led to a surge in efforts to make such intelligent systems transparent and explainable. The need…

人工智能 · 计算机科学 2020-11-30 Adriano Lucieri , Muhammad Naseer Bajwa , Andreas Dengel , Sheraz Ahmed

Recent advancements in multimodal large language models have driven breakthroughs in visual question answering. Yet, a critical gap persists, `conceptualization'-the ability to recognize and reason about the same concept despite variations…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Zahra Babaiee , Peyman M. Kiasari , Daniela Rus , Radu Grosu

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

Advances in data collection in radiation therapy have led to an abundance of opportunities for applying data mining and machine learning techniques to promote new data-driven insights. In light of these advances, supporting collaboration…

This paper introduces an innovative software system for fundus image analysis that deliberately diverges from the conventional screening approach, opting not to predict specific diagnoses. Instead, our methodology mimics the diagnostic…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Dmitry Ryabtsev , Boris Vasilyev , Sergey Shershakov

Artificial intelligence has made significant strides in medical visual question answering (Med-VQA), yet prevalent studies often interpret images holistically, overlooking the visual regions of interest that may contain crucial information,…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Xupeng Chen , Zhixin Lai , Kangrui Ruan , Shichu Chen , Jiaxiang Liu , Zuozhu Liu

Recent advances in Generative AI have transformed how users interact with data analysis through natural language interfaces. However, many systems rely too heavily on LLMs, creating risks of hallucination, opaque reasoning, and reduced user…

人机交互 · 计算机科学 2025-09-04 Ratanond Koonchanok , Alex Kale , Khairi Reda

Graphical user interface (GUI) has become integral to modern society, making it crucial to be understood for human-centric systems. However, unlike natural images or documents, GUIs comprise artificially designed graphical elements arranged…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Ziwei Wang , Weizhi Chen , Leyang Yang , Sheng Zhou , Shengchu Zhao , Hanbei Zhan , Jiongchao Jin , Liangcheng Li , Zirui Shao , Jiajun Bu

When embodied AI is expanding from traditional object detection and recognition to more advanced tasks of robot manipulation and actuation planning, visual spatial reasoning from the video inputs is necessary to perceive the spatial…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Haoming Wang , Qiyao Xue , Weichen Liu , Wei Gao

The use of deep learning in computer vision tasks such as image classification has led to a rapid increase in the performance of such systems. Due to this substantial increment in the utility of these systems, the use of artificial…

图像与视频处理 · 电气工程与系统科学 2023-04-05 Vinay Jogani , Joy Purohit , Ishaan Shivhare , Seema C Shrawne

The increasing demand for transparent and reliable models, particularly in high-stakes decision-making areas such as medical image analysis, has led to the emergence of eXplainable Artificial Intelligence (XAI). Post-hoc XAI techniques,…

计算机视觉与模式识别 · 计算机科学 2024-11-18 Junlin Hou , Sicen Liu , Yequan Bie , Hongmei Wang , Andong Tan , Luyang Luo , Hao Chen

Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as image captioning, visual question answering, and image-text retrieval. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Ilias Stogiannidis , Steven McDonagh , Sotirios A. Tsaftaris