English
Related papers

Related papers: Deep Multimodal Fusion for Surgical Feedback Class…

200 papers

The way of understanding online higher education has greatly changed due to the worldwide pandemic situation. Teaching is undertaken remotely, and the faculty incorporate lecture audio recordings as part of the teaching material. This new…

Computation and Language · Computer Science 2024-01-01 Oscar Sapena , Eva Onaindia

Balancing dialogue, music, and sound effects with accompanying video is crucial for immersive storytelling, yet current audio mixing workflows remain largely manual and labor-intensive. While recent advancements have introduced the visually…

Sound · Computer Science 2026-01-15 Junhua Huang , Chao Huang , Chenliang Xu

Semi-supervised learning addresses the issue of limited annotations in medical images effectively, but its performance is often inadequate for complex backgrounds and challenging tasks. Multi-modal fusion methods can significantly improve…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Dongdong Meng , Sheng Li , Hao Wu , Guoping Wang , Xueqing Yan

We learn about the world from a diverse range of sensory information. Automated systems lack this ability as investigation has centred on processing information presented in a single form. Adapting architectures to learn from multiple…

Machine Learning · Computer Science 2020-10-27 Jason Armitage , Shramana Thakur , Rishi Tripathi , Jens Lehmann , Maria Maleshkova

The interaction of Older adults with robots requires effective feedback to keep them aware of the state of the interaction for optimum interaction quality. This study examines the effect of different feedback modalities in a table setting…

Robotics · Computer Science 2021-03-16 Noa Markfeld , Samuel Olatunji , Dana Gutman , Shay Givati , Vardit Sarne-Fleischmann , Yael Edan

Multimodal surface material classification plays a critical role in advancing tactile perception for robotic manipulation and interaction. In this paper, we present Surformer v2, an enhanced multi-modal classification architecture designed…

Robotics · Computer Science 2025-09-08 Manish Kansana , Sindhuja Penchala , Shahram Rahimi , Noorbakhsh Amiri Golilarz

Learning to reliably perceive and understand the scene is an integral enabler for robots to operate in the real-world. This problem is inherently challenging due to the multitude of object types as well as appearance changes caused by…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Abhinav Valada , Rohit Mohan , Wolfram Burgard

Recent advances in machine learning models allowed robots to identify objects on a perceptual nonsymbolic level (e.g., through sensor fusion and natural language understanding). However, these primarily black-box learning models still lack…

Robotics · Computer Science 2023-07-11 Amr Gomaa , Bilal Mahdy , Niko Kleer , Michael Feld , Frank Kirchner , Antonio Krüger

Machine learning models are widely used to support stealth assessment in digital learning environments. Existing approaches typically rely on abstracted gameplay log data, which may overlook subtle behavioral cues linked to learners'…

Machine Learning · Computer Science 2025-07-31 Clemens Witt , Thiemo Leonhardt , Nadine Bergner , Mareen Grillenberger

The healthcare domain is characterized by heterogeneous data modalities, such as imaging and physiological data. In practice, the variety of medical data assists clinicians in decision-making. However, most of the current state-of-the-art…

Image and Video Processing · Electrical Eng. & Systems 2021-11-05 Nasir Hayat , Krzysztof J. Geras , Farah E. Shamout

In this work we explore the application of AI to robotic welding. Robotic welding is a widely used technology in many industries, but robots currently do not have the capability to detect welding defects which get introduced due to various…

Interactive segmentation is a promising strategy for building robust, generalisable algorithms for volumetric medical image segmentation. However, inconsistent and clinically unrealistic evaluation hinders fair comparison and misrepresents…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Parhom Esmaeili , Virginia Fernandez , Pedro Borges , Eli Gibson , Sebastien Ourselin , M. Jorge Cardoso

Audio-based disease prediction is emerging as a promising supplement to traditional medical diagnosis methods, facilitating early, convenient, and non-invasive disease detection and prevention. Multimodal fusion, which integrates features…

Nowadays, pre-trained encoders are widely used in medical image segmentation due to their strong capability in extracting rich and generalized feature representations. However, existing methods often fail to fully leverage these features,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Xiaolin Gou , Chuanlin Liao , Jizhe Zhou , Fengshuo Ye , Yi Lin

Leveraging information across diverse modalities is known to enhance performance on multimodal segmentation tasks. However, effectively fusing information from different modalities remains challenging due to the unique characteristics of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Md Kaykobad Reza , Ashley Prater-Bennette , M. Salman Asif

Real-time prediction of technical errors from cataract surgical videos can be highly beneficial, particularly for telementoring, which involves remote guidance and mentoring through digital platforms. However, the rarity of surgical errors…

Image and Video Processing · Electrical Eng. & Systems 2025-03-31 Maxime Faure , Pierre-Henri Conze , Béatrice Cochener , Anas-Alexis Benyoussef , Mathieu Lamard , Gwenolé Quellec

Audio-Video Emotion Recognition is now attacked with Deep Neural Network modeling tools. In published papers, as a rule, the authors show only cases of the superiority in multi-modality over audio-only or video-only modality. However, there…

Signal Processing · Electrical Eng. & Systems 2021-08-02 Xin Chang , Władysław Skarbek

Automatic speaker naming is the problem of localizing as well as identifying each speaking character in a TV/movie/live show video. This is a challenging problem mainly attributes to its multimodal nature, namely face cue alone is…

Computer Vision and Pattern Recognition · Computer Science 2015-07-20 Yongtao Hu , Jimmy Ren , Jingwen Dai , Chang Yuan , Li Xu , Wenping Wang

Humans do not understand individual events in isolation; rather, they generalize concepts within classes and compare them to others. Existing audio-video pre-training paradigms only focus on the alignment of the overall audio-video…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Kaixuan Cong , Yifan Wang , Rongkun Xue , Yuyang Jiang , Yiming Feng , Jing Yang

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

Computation and Language · Computer Science 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter
‹ Prev 1 8 9 10 Next ›