English
Related papers

Related papers: SurgOnAir: Hierarchy-Aware Real-Time Surgical Vide…

200 papers

This work introduces the first framework for reconstructing surgical dialogue from unstructured real-world recordings, which is crucial for characterizing teaching tasks. In surgical training, the formative verbal feedback that trainers…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Firdavs Nasriddinov , Rafal Kocielnik , Arushi Gupta , Cherine Yang , Elyssa Wong , Anima Anandkumar , Andrew Hung

Unlike offline processing, streaming video vision-language models face two fundamental constraints: causality and accumulation. Causality prevents access to future frames that offline methods exploit, while accumulation causes tokens to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xueyi Chen , Keda Tao , Kele Shao , Huan Wang

From a computer science viewpoint, a surgical domain model needs to be a conceptual one incorporating both behavior and data. It should therefore model actors, devices, tools, their complex interactions and data flow. To capture and model…

Computer Vision and Pattern Recognition · Computer Science 2021-06-30 Ege Özsoy , Evin Pınar Örnek , Ulrich Eck , Federico Tombari , Nassir Navab

Conversation agents powered by large language models are revolutionizing the way we interact with visual data. Recently, large vision-language models (LVLMs) have been extensively studied for both images and videos. However, these studies…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Juseong Jin , Chang Wook Jeong

Recently video generation has achieved substantial progress with realistic results. Nevertheless, existing AI-generated videos are usually very short clips ("shot-level") depicting a single scene. To deliver a coherent long video…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Xinyuan Chen , Yaohui Wang , Lingjun Zhang , Shaobin Zhuang , Xin Ma , Jiashuo Yu , Yali Wang , Dahua Lin , Yu Qiao , Ziwei Liu

Surgical robotics holds much promise for improving patient safety and clinician experience in the Operating Room (OR). However, it also comes with new challenges, requiring strong team coordination and effective OR management. Automatic…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Idris Hamoud , Muhammad Abdullah Jamal , Vinkle Srivastav , Didier Mutter , Nicolas Padoy , Omid Mohareri

Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot architecture that enables…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yawen Luo , Xiaoyu Shi , Junhao Zhuang , Yutian Chen , Quande Liu , Xintao Wang , Pengfei Wan , Tianfan Xue

Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to…

Sound · Computer Science 2026-02-13 Yufan Wen , Zhaocheng Liu , YeGuo Hua , Ziyi Guo , Lihua Zhang , Chun Yuan , Jian Wu

Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on…

Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not designed for streaming…

Computation and Language · Computer Science 2026-04-07 Tomer Krichli , Bhiksha Raj , Joseph Keshet

Stream-based runtime monitors are used in safety-critical applications such as Unmanned Aerial Systems (UAS) to compute comprehensive statistics and logical assessments of system health that provide the human operator with critical…

Formal Languages and Automata Theory · Computer Science 2022-05-26 Jan Baumeister , Bernd Finkbeiner , Stefan Gumhold , Malte Schledjewski

Intensive Care Units (ICUs) are critical environments characterized by high-stakes monitoring and complex data management. However, current practices often rely on manual data transcription and fragmented information systems, introducing…

Human-Computer Interaction · Computer Science 2025-12-11 Yibowen Zhao , Yiming Cao , Zhiqi Shen , Juan Du , Yonghui Xu , Lizhen Cui , Cyril Leung

Real-time streaming video understanding in domains such as autonomous driving and intelligent surveillance poses challenges beyond conventional offline video processing, requiring continuous perception, proactive decision making, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Haolin Yang , Feilong Tang , Lingxiao Zhao , Xinlin Zhuang , Yifan Lu , Xiang An , Ming Hu , Xiaofeng Zhang , Abdalla Swikir , Junjun He , Zongyuan Ge , Muhammad Haris Khan , Imran Razzak

Unpaired video-to-video translation aims to translate videos between a source and a target domain without the need of paired training data, making it more feasible for real applications. Unfortunately, the translated videos generally suffer…

Computer Vision and Pattern Recognition · Computer Science 2022-12-22 Kaihong Wang , Kumar Akash , Teruhisa Misu

Sparsity and low-rank models have been popular for reconstructing images and videos from limited or corrupted measurements. Dictionary or transform learning methods are useful in applications such as denoising, inpainting, and medical image…

Machine Learning · Statistics 2019-07-23 Brian E. Moore , Saiprasad Ravishankar , Raj Rao Nadakuditi , Jeffrey A. Fessler

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with…

In this paper, we tackle the task of online video temporal grounding (OnVTG), which requires the model to locate events related to a given text query within a video stream. Unlike regular video temporal grounding, OnVTG requires the model…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Minghang Zheng , Yuxin Peng , Benyuan Sun , Yi Yang , Yang Liu

One of the potential solutions for model interpretation is to train a surrogate model: a more transparent model that approximates the behavior of the model to be explained. Typically, classification rules or decision trees are used due to…

Human-Computer Interaction · Computer Science 2022-01-20 Jun Yuan , Brian Barr , Kyle Overton , Enrico Bertini

Storyboarding is an established method for designing user experiences. Generative AI can support this process by helping designers quickly create visual narratives. However, existing tools only focus on accurate text-to-image generation.…

Human-Computer Interaction · Computer Science 2024-07-11 Zhaohui Liang , Xiaoyu Zhang , Kevin Ma , Zhao Liu , Xipei Ren , Kosa Goucher-Lambert , Can Liu

Computer-assisted surgery (CAS) aims to provide the surgeon with the right type of assistance at the right moment. Such assistance systems are especially relevant in laparoscopic surgery, where CAS can alleviate some of the drawbacks that…

Computer Vision and Pattern Recognition · Computer Science 2017-02-14 Sebastian Bodenstedt , Martin Wagner , Darko Katić , Patrick Mietkowski , Benjamin Mayer , Hannes Kenngott , Beat Müller-Stich , Rüdiger Dillmann , Stefanie Speidel