English
Related papers

Related papers: CathAction: A Benchmark for Endovascular Intervent…

200 papers

Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic understanding of what is happening within the scene for safe interaction with other road…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Salman Khan , Izzeddin Teeti , Reza Javanmard Alitappeh , Mihaela C. Stoian , Eleonora Giunchiglia , Gurkirt Singh , Andrew Bradley , Fabio Cuzzolin

Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field has achieved remarkable progress in recent years, further…

Learning from (procedural) videos has increasingly served as a pathway for embodied agents to acquire skills from human demonstrations. To do this, video understanding models must be able to obtain structured understandings, such as the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zitian Tang , Rohan Myer Krishnan , Zhiqiu Yu , Chen Sun

The healthcare landscape is evolving, with patients seeking reliable information about their health conditions and available treatment options. Despite the abundance of information sources, the digital age overwhelms individuals with…

Human-Computer Interaction · Computer Science 2025-04-16 Pragnya Ramjee , Bhuvan Sachdeva , Satvik Golechha , Shreyas Kulkarni , Geeta Fulari , Kaushik Murali , Mohit Jain

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Fadime Sener , Dibyadip Chatterjee , Daniel Shelepov , Kun He , Dipika Singhania , Robert Wang , Angela Yao

Causality helps people reason about and understand complex systems, particularly through what-if analyses that explore how interventions might alter outcomes. Although existing methods embrace causal reasoning using interventions and…

Human-Computer Interaction · Computer Science 2025-07-22 Yanming Zhang , Krishnakumar Hegde , Klaus Mueller

Out of all existing frameworks for surgical workflow analysis in endoscopic videos, action triplet recognition stands out as the only one aiming to provide truly fine-grained and comprehensive information on surgical activities. This…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Chinedu Innocent Nwoye , Tong Yu , Cristians Gonzalez , Barbara Seeliger , Pietro Mascagni , Didier Mutter , Jacques Marescaux , Nicolas Padoy

Humans commonly work with multiple objects in daily life and can intuitively transfer manipulation skills to novel objects by understanding object functional regularities. However, existing technical approaches for analyzing and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Yun Liu , Haolin Yang , Xu Si , Ling Liu , Zipeng Li , Yuxiang Zhang , Yebin Liu , Li Yi

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Purpose: Echocardiographic interpretation requires video-level reasoning and guideline-based measurement analysis, which current deep learning models for cardiac ultrasound do not support. We present EchoAgent, a framework that enables…

Echocardiography plays a critical role in the diagnosis and monitoring of cardiovascular diseases as a non-invasive real-time assessment of cardiac structure and function. However, the growing scale of echocardiographic video data presents…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zhe Li , Hadrien Reynaud , Alberto Gomez , Bernhard Kainz

Perceiving and autonomously navigating through work zones is a challenging and underexplored problem. Open datasets for this long-tailed scenario are scarce. We propose the ROADWork dataset to learn to recognize, observe, analyze, and drive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Anurag Ghosh , Shen Zheng , Robert Tamburo , Khiem Vuong , Juan Alvarez-Padilla , Hailiang Zhu , Michael Cardei , Nicholas Dunn , Christoph Mertz , Srinivasa G. Narasimhan

Objective: Interventional devices, catheters and insertable imaging devices such as transesophageal echo (TOE) probes are routinely used in minimally invasive cardiovascular procedures. Detecting their positions and orientations in X-ray…

Image and Video Processing · Electrical Eng. & Systems 2025-03-11 YingLiang Ma , Sandra Howell , Aldo Rinaldi , Tarv Dhanjal , Kawal S. Rhode

This paper presents VDAct, a dataset for a Video-grounded Dialogue on Event-driven Activities, alongside VDEval, a session-based context evaluation metric specially designed for the task. Unlike existing datasets, VDAct includes longer and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Wiradee Imrattanatrai , Masaki Asada , Kimihiro Hasegawa , Zhi-Qi Cheng , Ken Fukuda , Teruko Mitamura

High-quality, large-scale data is essential for robust deep learning models in medical applications, particularly ultrasound image analysis. Diffusion models facilitate high-fidelity medical image generation, reducing the costs associated…

Image and Video Processing · Electrical Eng. & Systems 2024-04-01 Pooria Ashrafian , Milad Yazdani , Moein Heidari , Dena Shahriari , Ilker Hacihaliloglu

Missions to small celestial bodies rely heavily on optical feature tracking for characterization of and relative navigation around the target body. While deep learning has led to great advancements in feature detection and description,…

Instrumentation and Methods for Astrophysics · Physics 2023-01-16 Travis Driver , Katherine Skinner , Mehregan Dor , Panagiotis Tsiotras

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Linfeng Dong , Yuchen Yang , Hao Wu , Wei Wang , Yuenan Hou , Zhihang Zhong , Xiao Sun

Optical coherence tomography angiography (OCTA) performs non-invasive visualization and characterization of microvasculature in research and clinical applications mainly in ophthalmology and dermatology. A wide variety of instruments,…

Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT). Yet, existing methods largely relegate clinicians to passive observers of final…

Cataract surgery is a sight saving surgery that is performed over 10 million times each year around the world. With such a large demand, the ability to organize surgical wards and operating rooms efficiently is critical to delivery this…

Image and Video Processing · Electrical Eng. & Systems 2021-06-22 Andrés Marafioti , Michel Hayoz , Mathias Gallardo , Pablo Márquez Neila , Sebastian Wolf , Martin Zinkernagel , Raphael Sznitman
‹ Prev 1 4 5 6 7 8 10 Next ›