English
Related papers

Related papers: Surgical Instruction Generation with Transformers

200 papers

Computer-assisted surgical (CAS) systems enhance surgical execution and outcomes by providing advanced support to surgeons. These systems often rely on deep learning models trained on complex, challenging-to-annotate data. While synthetic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Sabina Martyniak , Joanna Kaleta , Diego Dall'Alba , Michał Naskręt , Szymon Płotka , Przemysław Korzeniowski

The emergence of multi-modal deep learning models has made significant impacts on clinical applications in the last decade. However, the majority of models are limited to single-tasking, without considering disease diagnosis is indeed a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Lijian Xu , Ziyu Ni , Xinglong Liu , Xiaosong Wang , Hongsheng Li , Shaoting Zhang

Procedural videos, exemplified by recipe demonstrations, are instrumental in conveying step-by-step instructions. However, understanding such videos is challenging as it involves the precise localization of steps and the generation of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Anil Batra , Davide Moltisanti , Laura Sevilla-Lara , Marcus Rohrbach , Frank Keller

Surgical data science is a new research field that aims to observe all aspects of the patient treatment process in order to provide the right assistance at the right time. Due to the breakthrough successes of deep learning-based solutions…

Recent language models have achieved impressive performance in natural language tasks by incorporating instructions with task input during fine-tuning. Since all samples in the same natural language task can be explained with the same task…

Computation and Language · Computer Science 2023-11-14 Jin Myung Kwak , Minseon Kim , Sung Ju Hwang

Recent advances in synthetic imaging open up opportunities for obtaining additional data in the field of surgical imaging. This data can provide reliable supplements supporting surgical applications and decision-making through computer…

Image and Video Processing · Electrical Eng. & Systems 2023-12-07 Simeon Allmendinger , Patrick Hemmer , Moritz Queisner , Igor Sauer , Leopold Müller , Johannes Jakubik , Michael Vössing , Niklas Kühl

Autonomous navigation in underwater environments presents challenges due to factors such as light absorption and water turbidity, limiting the effectiveness of optical sensors. Sonar systems are commonly used for perception in underwater…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Ivano Donadi , Emilio Olivastri , Daniel Fusaro , Wanmeng Li , Daniele Evangelista , Alberto Pretto

Videos are prominent learning materials to prepare surgical trainees before they enter the operating room (OR). In this work, we explore techniques to enrich the video-based surgery learning experience. We propose Surgment, a system that…

Human-Computer Interaction · Computer Science 2024-06-27 Jingying Wang , Haoran Tang , Taylor Kantor , Tandis Soltani , Vitaliy Popov , Xu Wang

Recipe generation from food images and ingredients is a challenging task, which requires the interpretation of the information from another modality. Different from the image captioning task, where the captions usually have one sentence,…

Computer Vision and Pattern Recognition · Computer Science 2022-02-17 Hao Wang , Guosheng Lin , Steven C. H. Hoi , Chunyan Miao

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Nikita Dvornik , Isma Hadji , Ran Zhang , Konstantinos G. Derpanis , Animesh Garg , Richard P. Wildes , Allan D. Jepson

Recently, diffusion transformers have gained wide attention with its excellent performance in text-to-image and text-to-vidoe models, emphasizing the need for transformers as backbone for diffusion models. Transformer-based models have…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Nithin Gopalakrishnan Nair , Jeya Maria Jose Valanarasu , Vishal M. Patel

Transformer has achieved impressive successes for various computer vision tasks. However, most of existing studies require to pretrain the Transformer backbone on a large-scale labeled dataset (e.g., ImageNet) for achieving satisfactory…

Computer Vision and Pattern Recognition · Computer Science 2023-01-04 Yuexiang Li , Yawen Huang , Nanjun He , Kai Ma , Yefeng Zheng

Deep learning has led to state-of-the-art results for many medical imaging tasks, such as segmentation of different anatomical structures. With the increased numbers of deep learning publications and openly available code, the approach to…

Image and Video Processing · Electrical Eng. & Systems 2020-05-19 Tom van Sonsbeek , Veronika Cheplygina

An increasing number of colonoscopic guidance and assistance systems rely on machine learning algorithms which require a large amount of high-quality training data. In order to ensure high performance, the latter has to resemble a…

Image and Video Processing · Electrical Eng. & Systems 2022-05-24 Abhishek Dinkar Jagtap , Mattias Heinrich , Marian Himstedt

The Impression section of a radiology report summarizes crucial radiology findings in natural language and plays a central role in communicating these findings to physicians. However, the process of generating impressions by summarizing…

Computation and Language · Computer Science 2018-10-10 Yuhao Zhang , Daisy Yi Ding , Tianpei Qian , Christopher D. Manning , Curtis P. Langlotz

We present a novel masked image modeling (MIM) approach, context autoencoder (CAE), for self-supervised representation pretraining. We pretrain an encoder by making predictions in the encoded representation space. The pretraining tasks…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Xiaokang Chen , Mingyu Ding , Xiaodi Wang , Ying Xin , Shentong Mo , Yunhao Wang , Shumin Han , Ping Luo , Gang Zeng , Jingdong Wang

Models trained for classification often assume that all testing classes are known while training. As a result, when presented with an unknown class during testing, such closed-set assumption forces the model to classify it as one of the…

Computer Vision and Pattern Recognition · Computer Science 2019-04-03 Poojan Oza , Vishal M Patel

Transformer architecture has widespread applications, particularly in Natural Language Processing and computer vision. Recently Transformers have been employed in various aspects of time-series analysis. This tutorial provides an overview…

Machine Learning · Computer Science 2023-07-27 Sabeen Ahmed , Ian E. Nielsen , Aakash Tripathi , Shamoon Siddiqui , Ghulam Rasool , Ravi P. Ramachandran

Surgical simulation offers a promising addition to conventional surgical training. However, available simulation tools lack photorealism and rely on hardcoded behaviour. Denoising Diffusion Models are a promising alternative for…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Yannik Frisch , Ssharvien Kumar Sivakumar , Çağhan Köksal , Elsa Böhm , Felix Wagner , Adrian Gericke , Ghazal Ghazaei , Anirban Mukhopadhyay

Recent advances in deep learning and natural language generation have significantly improved image captioning, enabling automated, human-like descriptions for visual content. In this work, we apply these captioning techniques to generate…

Computation and Language · Computer Science 2024-12-06 Amnon Bleich , Antje Linnemann , Bjoern H. Diem , Tim OF Conrad