English
Related papers

Related papers: VoxLogicA UI: Supporting Declarative Medical Image…

200 papers

Explainable AI (XAI) interfaces seek to make large language models more transparent, yet explanation alone does not produce understanding. Explaining a system's behavior is not the same as being able to engage with it, to probe and…

Human-Computer Interaction · Computer Science 2026-03-18 Gabrielle Benabdallah

Large Vision-Language Models offer a new paradigm for AI-driven image understanding, enabling models to perform tasks without task-specific training. This flexibility holds particular promise across medicine, where expert-annotated data is…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Anita Rau , Mark Endo , Josiah Aklilu , Jaewoo Heo , Khaled Saab , Alberto Paderno , Jeffrey Jopling , F. Christopher Holsinger , Serena Yeung-Levy

Neuronal network models and corresponding computer simulations are invaluable tools to aid the interpretation of the relationship between neuron properties, connectivity and measured activity in cortical tissue. Spatiotemporal patterns of…

Neurons and Cognition · Quantitative Biology 2022-09-16 Johanna Senk , Corto Carde , Espen Hagen , Torsten W. Kuhlen , Markus Diesmann , Benjamin Weyers

Medical Visual Language Models have shown great potential in various healthcare applications, including medical image captioning and diagnostic assistance. However, most existing models rely on text-based instructions, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Tan-Hanh Pham , Chris Ngo , Trong-Duong Bui , Minh Luu Quang , Tan-Huong Pham , Truong-Son Hy

Evaluating UX in the context of AI's complexity, unpredictability, and generative nature presents unique challenges. How can we support HCI researchers to create comprehensive UX evaluation plans? In this paper, we introduce EvAlignUX, a…

Human-Computer Interaction · Computer Science 2025-07-09 Qingxiao Zheng , Minrui Chen , Pranav Sharma , Yiliu Tang , Mehul Oswal , Yiren Liu , Yun Huang

We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs' spatial reasoning. By combining minimal manual supervision with large-scale Visual Question-Answering (VQA) pairs…

Robotics · Computer Science 2026-05-05 Yangzhe Kong , Daeun Song , Jing Liang , Dinesh Manocha , Ziyu Yao , Xuesu Xiao

Accurate diagnosis of ophthalmic diseases relies heavily on the interpretation of multimodal ophthalmic images, a process often time-consuming and expertise-dependent. Visual Question Answering (VQA) presents a potential interdisciplinary…

Image and Video Processing · Electrical Eng. & Systems 2024-10-23 Xiaolan Chen , Ruoyu Chen , Pusheng Xu , Weiyi Zhang , Xianwen Shang , Mingguang He , Danli Shi

Recent advancements in model checking have demonstrated significant potential across diverse applications, particularly in signal and image analysis. Medical imaging stands out as a critical domain where model checking can be effectively…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Elhoucine Elfatimi , Lahcen El fatimi

Recent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology,…

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant)-322K, a diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Qi Feng

Modern Vision-Language Models (VLMs) exhibit unprecedented capabilities in cross-modal semantic understanding between visual and textual modalities. Given the intrinsic need for multi-modal integration in clinical applications, VLMs have…

Image and Video Processing · Electrical Eng. & Systems 2025-06-24 Haoneng Lin , Cheng Xu , Jing Qin

Deep learning has made a remarkable impact in the field of natural image processing over the past decade. Consequently, there is a great deal of interest in replicating this success across unsolved tasks in related domains, such as medical…

Image and Video Processing · Electrical Eng. & Systems 2021-05-14 Teofilo E. Zosa

Artificial intelligence (AI) has shown great promise for diagnostic imaging assessments. However, the application of AI to support medical diagnostics in clinical routine comes with many challenges. The algorithms should have high…

Computer Vision and Pattern Recognition · Computer Science 2020-08-17 Milda Pocevičiūtė , Gabriel Eilertsen , Claes Lundström

Visual Question Answering (VQA) is an evolving research field aimed at enabling machines to answer questions about visual content by integrating image and language processing techniques such as feature extraction, object detection, text…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Ngoc Dung Huynh , Mohamed Reda Bouadjenek , Sunil Aryal , Imran Razzak , Hakim Hacid

Vision language models (VLMs) are designed to extract relevant visuospatial information from images. Some research suggests that VLMs can exhibit humanlike scene understanding, while other investigations reveal difficulties in their ability…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Sangeet Khemlani , Tyler Tran , Nathaniel Gyory , Anthony M. Harrison , Wallace E. Lawson , Ravenna Thielstrom , Hunter Thompson , Taaren Singh , J. Gregory Trafton

3D modeling is becoming a well-developed field of medicine, but its applicability can be limited due to the lack of software allowing for easy utilizations of generated 3D visualizations. By leveraging recent advances in virtual reality, we…

Human-Computer Interaction · Computer Science 2020-12-07 Alex J. Deakyne , Erik N. Gaasedelen , Tinen L. Iles , Paul A. Iaizzo

Design for Voice User Interfaces (VUIs) has become more relevant in recent years due to the enormous advances of speech technologies and their growing presence in our everyday lives. Although modern VUIs still present interaction issues,…

Human-Computer Interaction · Computer Science 2019-04-15 Gisela Reyes-Cruz , Joel Fischer , Stuart Reeves

Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms that will enable them. With these goals in mind we have…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Aneeq Zia , Max Berniker , Rogerio Garcia Nespolo , Xiaorui Zhang , Conor Perreault , Kiran Bhattacharyya , Xi Liu , Ziheng Wang , Satoshi Kondo , Satoshi Kasai , Kousuke Hirasawa , Bo Liu , David Austin , Yiheng Wang , Michal Futrega , Jean-Francois Puget , Zhenqiang Li , Yoichi Sato , Ryo Fujii , Ryo Hachiuma , Mana Masuda , Hideo Saito , An Wang , Mengya Xu , Mobarakol Islam , Long Bai , Winnie Pang , Hongliang Ren , Chinedu Nwoye , Luca Sestini , Nicolas Padoy , Maximilian Nielsen , Samuel Schüttler , Thilo Sentker , Hümeyra Husseini , Ivo Baltruschat , Rüdiger Schmitz , René Werner , Aleksandr Matsun , Mugariya Farooq , Numan Saaed , Jose Renato Restom Viera , Mohammad Yaqub , Neil Getty , Fangfang Xia , Zixuan Zhao , Xiaotian Duan , Xing Yao , Ange Lou , Hao Yang , Jintong Han , Jack Noble , Jie Ying Wu , Tamer Abdulbaki Alshirbaji , Nour Aldeen Jalal , Herag Arabian , Ning Ding , Knut Moeller , Weiliang Chen , Quan He , Muhammad Bilal , Taofeek Akinosho , Adnan Qayyum , Massimo Caputo , Hunaid Vohra , Michael Loizou , Anuoluwapo Ajayi , Ilhem Berrou , Faatihah Niyi-Odumosu , Charlie Budd , Oluwatosin Alabi , Tom Vercauteren , Ruoxi Zhao , Ayberk Acar , John Han , Jumanh Atoum , Yinhong Qin , Surong Hua , Lu Ping , Wenming Wu , Rongfeng Wei , Jinlin Wu , You Pang , Zhen Chen , Tim Jaspers , Amine Yamlahi , Piotr Kalinowski , Dominik Michael , Tim Rädsch , Marco Hübner , Danail Stoyanov , Stefanie Speidel , Lena Maier-Hein , Jie Tian , Ruxin Zhang , Khang Hoang Nguyen , Anh Quoc Nguyen , Tam Minh Nguyen , Khoi Dinh Tran , Minh Nguyen Dang Nhat , Trinh Thi Doan Pham , Linh Van Nguyen , Chunyang Jiang , Dewei Yang , Haitao Li , Yannick Prudent , Thibaut Boissin , Mahmood Alam , Shazad Ashraf , Andrew D. Beggs , Lukman Akanbi , Manuel D. Delgado , Narain Gupta , Amir M. Hajiyavand , Iqbal Qasim , Hafiz A. Alaka , Junaid Qadir , Shu Yang , Yihui Wang , Hao Chen , Shin Paul , Yosuke Yamagishi , Zhang Dong , Hongyun Li , Hongyu Gu , Xiaoliu Ding , Xiaoyao Liu , Xingyu Zhao , Mariana Ribeiro , Tiago Jesus , André Ferreira , Guilherme Barbosa , João Carvalho , Leonardo Barroso , Nuno Gomes , Rafael Peixoto , Rodrigo Ralha , Victor Alves , Stephanie , Nattapat Ittikosil , Achita Chitrapan , Quan Huu Cap , Jiayuan Huang , Shreyas C Dhake , Sergi Kavtaradze , Mobarak I Hoque , Ka Young Kim , Su Yong Yun , Young Tae Kim , Hyeon Bae Kim , Seong Tae Kim , Zuxing Deng , Ling Li , Jieyu Zheng , Xiaojian Li , Anthony Jarc

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Haoyu Zhang , Meng Liu , Zaijing Li , Haokun Wen , Weili Guan , Yaowei Wang , Liqiang Nie

Visual analytics (VA) is typically applied to complex data, thus requiring complex tools. While visual analytics empowers analysts in data analysis, analysts may get lost in the complexity occasionally. This highlights the need for…

Human-Computer Interaction · Computer Science 2025-07-25 Yuheng Zhao , Xueli Shu , Liwen Fan , Lin Gao , Yu Zhang , Siming Chen