English
Related papers

Related papers: PoCaP Corpus: A Multimodal Dataset for Smart Opera…

200 papers

To meet the growing demand for systematic surgical training, wet-lab environments have become indispensable platforms for hands-on practice in ophthalmology. Yet, traditional wet-lab training depends heavily on manual performance…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Negin Ghamsarian , Raphael Sznitman , Klaus Schoeffmann , Jens Kowal

Vision-language pre-training (VLP) models have been demonstrated to be effective in many computer vision applications. In this paper, we consider developing a VLP model in the medical domain for making computer-aided diagnoses (CAD) based…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Qiuhui Chen , Xinyue Hu , Zirui Wang , Yi Hong

Understanding the workflow of surgical procedures in complex operating rooms requires a deep understanding of the interactions between clinicians and their environment. Surgical activity recognition (SAR) is a key computer vision task that…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Idris Hamoud , Vinkle Srivastav , Muhammad Abdullah Jamal , Didier Mutter , Omid Mohareri , Nicolas Padoy

This contribution introduces a dataset of 7th-order Ambisonic Room Impulse Responses (HOA-RIRs), created using the Image Source Method. By employing higher-order Ambisonics, our dataset enables precise spatial audio reproduction, a critical…

Sound · Computer Science 2025-06-02 Shivam Saini , Jürgen Peissig

Nowadays, research in speech technologies has gotten a lot out thanks to recently created public domain corpora that contain thousands of recording hours. These large amounts of data are very helpful for training the new complex models…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-12 Guillermo Cámbara , Alex Peiró-Lilja , Mireia Farrús , Jordi Luque

The availability of realistic simulated corpora is of key importance for the future progress of distant speech recognition technology. The reliability, flexibility and low computational cost of a data simulation process may ultimately allow…

Audio and Speech Processing · Electrical Eng. & Systems 2017-11-28 Mirco Ravanelli , Piergiorgio Svaizer , Maurizio Omologo

Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora,…

Sound · Computer Science 2024-12-31 Shreya G. Upadhyay , Ali N. Salman , Carlos Busso , Chi-Chun Lee

This dissertation covers a single-processor approach to the speech processing pipeline of bilateral Cochlear Implants (CIs). The use of only a single processor to provide binaural stimulation signals overcomes the synchronization problem,…

Sound · Computer Science 2014-09-24 Taher Shahbazi Mirzahasanloo

Cardiac pulsation is a physiological confound of functional magnetic resonance imaging (fMRI) time-series that introduces spurious signal fluctuations in proximity to blood vessels. fMRI alone is not sufficiently fast to resolve cardiac…

Emotion recognition is a topic of significant interest in assistive robotics due to the need to equip robots with the ability to comprehend human behavior, facilitating their effective interaction in our society. Consequently, efficient and…

Human-Computer Interaction · Computer Science 2023-12-05 Rutherford Agbeshi Patamia , Paulo E. Santos , Kingsley Nketia Acheampong , Favour Ekong , Kwabena Sarpong , She Kun

While natural language processing (NLP) of unstructured clinical narratives holds the potential for patient care and clinical research, portability of NLP approaches across multiple sites remains a major challenge. This study investigated…

Many recent studies leverage the pre-trained CLIP for text-video cross-modal retrieval by tuning the backbone with additional heavy modules, which not only brings huge computational burdens with much more parameters, but also leads to the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Siteng Huang , Biao Gong , Yulin Pan , Jianwen Jiang , Yiliang Lv , Yuyuan Li , Donglin Wang

The majority of current research in deep learning based image registration addresses inter-patient brain registration with moderate deformation magnitudes. The recent Learn2Reg medical registration benchmark has demonstrated that…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Mattias P. Heinrich , Lasse Hansen

We define a representation framework for extracting spatial information from radiology reports (Rad-SpRL). We annotated a total of 2000 chest X-ray reports with 4 spatial roles corresponding to the common radiology entities. Our focus is on…

Computation and Language · Computer Science 2019-08-14 Surabhi Datta , Yuqi Si , Laritza Rodriguez , Sonya E Shooshan , Dina Demner-Fushman , Kirk Roberts

The multi-modality imaging system offers optimal fused images for safe and precise interventions in modern clinical practices, such as computed tomography - ultrasound (CT-US) guidance for needle insertion. However, the limited dexterity…

Robotics · Computer Science 2025-02-18 Feng Li , Yuan Bi , Dianye Huang , Zhongliang Jiang , Nassir Navab

We propose MORAL (a multimodal reinforcement learning framework for decision making in autonomous laboratories) that enhances sequential decision-making in autonomous robotic laboratories through the integration of visual and textual…

Machine Learning · Computer Science 2025-04-07 Natalie Tirabassi , Sathish A. P. Kumar , Sumit Jha , Arvind Ramanathan

Although fully autonomous systems still face challenges due to patients' anatomical variability, teleoperated systems appear to be more practical in current healthcare settings. This paper presents an anatomy-aware control framework for…

Robotics · Computer Science 2026-02-13 Davide Nardi , Edoardo Lamon , Daniele Fontanelli , Matteo Saveriano , Luigi Palopoli

Transcription of broadcast news is an interesting and challenging application for large-vocabulary continuous speech recognition (LVCSR). We present in detail the structure of a manually segmented and annotated corpus including over 160…

Computation and Language · Computer Science 2014-12-16 Felix Weninger , Björn Schuller , Florian Eyben , Martin Wöllmer , Gerhard Rigoll

Localisation of surgical tools constitutes a foundational building block for computer-assisted interventional technologies. Works in this field typically focus on training deep learning models to perform segmentation tasks. Performance of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Zhe Han , Charlie Budd , Gongyu Zhang , Huanyu Tian , Christos Bergeles , Tom Vercauteren

While recent automatic speech recognition systems achieve remarkable performance when large amounts of adequate, high quality annotated speech data is used for training, the same systems often only achieve an unsatisfactory result for tasks…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-19 Michael Gref , Oliver Walter , Christoph Schmidt , Sven Behnke , Joachim Köhler