English
Related papers

Related papers: Specialized Foundation Models for Intelligent Oper…

200 papers

Rapid advancements in foundation models, including Large Language Models, Vision-Language Models, Multimodal Large Language Models, and Vision-Language-Action Models, have opened new avenues for embodied AI in mobile service robotics. By…

Robotics · Computer Science 2026-03-11 Matthew Lisondra , Beno Benhabib , Goldie Nejat

The simulation of complex systems increasingly relies on sophisticated but fundamentally opaque computational black-box simulators. Surrogate models play a central role in reducing the computational cost of complex systems simulations…

Deep learning has revolutionized medical image registration by achieving unprecedented speeds, yet its clinical application is hindered by a limited ability to generalize beyond the training domain, a critical weakness given the typically…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Fengting Zhang , Yue He , Qinghao Liu , Yaonan Wang , Xiang Chen , Hang Zhang

Recent advances in spatial omics technologies have revolutionized our ability to study biological systems with unprecedented resolution. By preserving the spatial context of molecular measurements, these methods enable comprehensive mapping…

Quantitative Methods · Quantitative Biology 2025-09-18 Zhiwei Fan , Tiangang Wang , Kexin Huang , Binwu Ying , Xiaobo Zhou

Optical systems are becoming increasingly important by resolving many bottlenecks in today's communication, electronics, and biomedical systems. However, given the continuous nature of optics, the inability to efficiently analyze optical…

Logic in Computer Science · Computer Science 2014-03-13 Sanaz Khan-Afshar , Umair Siddique , Mohamed Yousri Mahmoud , Vincent Aravantinos , Ons Seddiki , Osman Hasan , Sofiene Tahar

Despite significant advances in artificial intelligence (AI) for computer vision, its application in medical imaging has been limited by the burden and limits of expert-generated labels. We used images from optical coherence tomography…

Computer Vision and Pattern Recognition · Computer Science 2018-02-27 Cecilia S. Lee , Ariel J. Tyring , Yue Wu , Sa Xiao , Ariel S. Rokem , Nicolaas P. Deruyter , Qinqin Zhang , Adnan Tufail , Ruikang K. Wang , Aaron Y. Lee

Surgical Video Question Answering (VideoQA) provides a promising paradigm for dynamic intraoperative interpretation, enabling real-time decision support and context-aware retrieval in clinical environments. Nevertheless, existing approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Diandian Guo , Xikai Yang , Ruiyang Li , Jialun Pei , Pheng-Ann Heng

Large visual language models (VLMs) have shown strong multi-modal medical reasoning ability, but most operate as end-to-end black boxes, diverging from clinicians' evidence-based, staged workflows and hindering clinical accountability.…

Artificial Intelligence · Computer Science 2026-03-12 Yuexi Du , Jinglu Wang , Shujie Liu , Nicha C. Dvornek , Yan Lu

The human ability to seamlessly perform multimodal reasoning and physical interaction in the open world is a core goal for general purpose embodied intelligent systems. Recent vision-language-action (VLA) models, which are co-trained on…

Full-Reference image quality assessment (FR IQA) is important for image compression, restoration and generative modeling, yet current neural metrics remain slow and vulnerable to adversarial perturbations. We present BiRQA, a compact FR IQA…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Aleksandr Gushchin , Dmitriy S. Vatolin , Anastasia Antsiferova

Comprehensively understanding surgical scenes in Surgical Visual Question Answering (Surgical VQA) requires reasoning over multiple objects. Previous approaches address this task using cross-modal fusion strategies to enhance reasoning…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Wenjun Hou , Yi Cheng , Kaishuai Xu , Yan Hu , Wenjie Li , Jiang Liu

Foundation models refer to artificial intelligence (AI) models that are trained on massive amounts of data and demonstrate broad generalizability across various tasks with high accuracy. These models offer versatile, one-for-many or…

Image and Video Processing · Electrical Eng. & Systems 2024-11-06 Rina Bao , Erfan Darzi , Sheng He , Chuan-Heng Hsiao , Mohammad Arafat Hussain , Jingpeng Li , Atle Bjornerud , Ellen Grant , Yangming Ou

Multimodal information is frequently available in medical tasks. By combining information from multiple sources, clinicians are able to make more accurate judgments. In recent years, multiple imaging techniques have been used in clinical…

Optical character recognition (OCR) and multilingual text understanding remain major failure modes of multimodal large language models (MLLMs), particularly in real-world images containing cluttered layouts, small fonts, blur, occlusion,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Qinwu Xu , Yifan Jiang , Haoyu Ren

In this position paper we describe a general framework for applying machine learning and pattern recognition techniques in healthcare. In particular, we are interested in providing an automated tool for monitoring and incrementing the level…

Computers and Society · Computer Science 2015-12-01 Filippo Maria Bianchi , Enrico De Santis , Hedieh Montazeri , Parisa Naraei , Alireza Sadeghian

Quantum computing is rapidly emerging as a new computing paradigm with the potential to improve decision-making, optimization, and simulation across industries. For industrial engineering (IE) and operations research (OR), this shift…

Quantum Physics · Physics 2025-10-24 Emily L. Tucker , Mohammadhossein Mohammadisiahroudi

Optimization modeling plays a critical role in the application of Operations Research (OR) tools to address real-world problems, yet they pose challenges and require extensive expertise from OR experts. With the advent of large language…

Computation and Language · Computer Science 2025-07-30 Chenyu Huang , Zhengyang Tang , Shixi Hu , Ruoqing Jiang , Xin Zheng , Dongdong Ge , Benyou Wang , Zizhuo Wang

Recent advancements in machine learning (ML) and deep learning (DL), particularly through the introduction of Foundation Models (FMs), have significantly enhanced surgical scene understanding within minimally invasive surgery (MIS). This…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Ufaq Khan , Umair Nawaz , Adnan Qayyum , Shazad Ashraf , Yutong Xie , Muhammad Haris Khan , Muhammad Bilal , Junaid Qadir

Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent representations frequently entangle physiologic severity, intervention intensity, observational…

Machine Learning · Computer Science 2026-05-19 Yuanyun Zhang , Shi Li

The rapid advancement of artificial intelligence (AI) techniques has opened up new opportunities to revolutionize various fields, including operations research (OR). This survey paper explores the integration of AI within the OR process…

Optimization and Control · Mathematics 2024-03-28 Zhenan Fan , Bissan Ghaddar , Xinglu Wang , Linzi Xing , Yong Zhang , Zirui Zhou