中文
相关论文

相关论文: ARST: Auto-Regressive Surgical Transformer for Pha…

200 篇论文

Automatic surgical workflow recognition is a key component for developing context-aware computer-assisted systems in the operating theatre. Previous works either jointly modeled the spatial features with short fixed-range temporal…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Yueming Jin , Yonghao Long , Cheng Chen , Zixu Zhao , Qi Dou , Pheng-Ann Heng

Deep learning has facilitated the automation of radiotherapy by predicting accurate dose distribution maps. However, existing methods fail to derive the desirable radiotherapy parameters that can be directly input into the treatment…

计算机视觉与模式识别 · 计算机科学 2024-03-01 Jiaqi Cui , Yuanyuan Xu , Jianghong Xiao , Yuchen Fei , Jiliu Zhou , Xingcheng Peng , Yan Wang

Autonomous laparoscopic camera control must maintain a stable and safe surgical view under rapid tool-tissue interactions while remaining interpretable to surgeons. We present a strategy-grounded framework that couples high-level…

机器人学 · 计算机科学 2026-02-25 Keyu Zhou , Peisen Xu , Yahao Wu , Jiming Chen , Gaofeng Li , Shunlei Li

Driven by the continuous development of models such as Multi-Layer Perceptron, Convolutional Neural Network (CNN), and Transformer, deep learning has made breakthrough progress in fields such as computer vision and natural language…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Shuang Liu , Lina Zhao , Tian Wang , Huaqing Wang

Recently, foundation models based on Vision Transformers (ViTs) have become widely available. However, their fine-tuning process is highly resource-intensive, and it hinders their adoption in several edge or low-energy applications. To this…

计算机视觉与模式识别 · 计算机科学 2024-08-19 Alessio Devoto , Federico Alvetreti , Jary Pomponi , Paolo Di Lorenzo , Pasquale Minervini , Simone Scardapane

Audio-visual automatic speech recognition (AV-ASR) extends speech recognition by introducing the video modality as an additional source of information. In this work, the information contained in the motion of the speaker's mouth is used to…

计算机视觉与模式识别 · 计算机科学 2022-11-02 Dmitriy Serdyuk , Otavio Braga , Olivier Siohan

Video generation, while capable of generating realistic videos, is computationally expensive and slow, prohibiting real-time applications. In this paper, we observe that video latents encoded via an autoencoder under the Latent Diffusion…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Dennis Menn , Chih-Hsien Chou

Following the technological advancements in medicine, the operation rooms are evolving into intelligent environments. The context-aware systems (CAS) can comprehensively interpret the surgical state, enable real-time warning, and support…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Negin Ghamsarian

Surgical phase recognition plays a crucial role in surgical workflow analysis, enabling various applications such as surgical monitoring, skill assessment, and workflow optimization. Despite significant advancements in deep learning-based…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Ka Young Kim , Hyeon Bae Kim , Seong Tae Kim

We present ASSET, a neural architecture for automatically modifying an input high-resolution image according to a user's edits on its semantic segmentation map. Our architecture is based on a transformer with a novel attention mechanism.…

计算机视觉与模式识别 · 计算机科学 2022-05-25 Difan Liu , Sandesh Shetty , Tobias Hinz , Matthew Fisher , Richard Zhang , Taesung Park , Evangelos Kalogerakis

Anomaly detection in time series data is crucial across various domains. The scarcity of labeled data for such tasks has increased the attention towards unsupervised learning methods. These approaches, often relying solely on reconstruction…

机器学习 · 计算机科学 2024-05-14 Ramin Ghorbani , Marcel J. T. Reinders , David M. J. Tax

Video activity recognition has become increasingly important in robots and embodied AI. Recognizing continuous video activities poses considerable challenges due to the fast expansion of streaming video, which contains multi-scale and…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Hao Wu , Donglin Bai , Shiqi Jiang , Qianxi Zhang , Yifan Yang , Xin Ding , Ting Cao , Yunxin Liu , Fengyuan Xu

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and…

We propose a novel attentive sequence to sequence translator (ASST) for clip localization in videos by natural language descriptions. We make two contributions. First, we propose a bi-directional Recurrent Neural Network (RNN) with a finely…

计算机视觉与模式识别 · 计算机科学 2018-08-28 Ke Ning , Linchao Zhu , Ming Cai , Yi Yang , Di Xie , Fei Wu

Online egocentric gaze estimation predicts where a camera wearer is looking from first-person video using only past and current frames, a task essential for augmented reality and assistive technologies. Unlike third-person gaze estimation,…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Jia Li , Wenjie Zhao , Shijian Deng , Bolin Lai , Yuheng Wu , RUijia Chen , Jon E. Froehlich , Yuhang Zhao , Yapeng Tian

Recurrent Neural Networks were, until recently, one of the best ways to capture the timely dependencies in sequences. However, with the introduction of the Transformer, it has been proven that an architecture with only attention-mechanisms…

机器学习 · 计算机科学 2021-08-19 Radostin Cholakov , Todor Kolev

Action recognition is an open and challenging problem in computer vision. While current state-of-the-art models offer excellent recognition results, their computational expense limits their impact for many real-world applications. In this…

计算机视觉与模式识别 · 计算机科学 2020-08-03 Yue Meng , Chung-Ching Lin , Rameswar Panda , Prasanna Sattigeri , Leonid Karlinsky , Aude Oliva , Kate Saenko , Rogerio Feris

Autoregressive video models offer distinct advantages over bidirectional diffusion models in creating interactive video content and supporting streaming applications with arbitrary duration. In this work, we present Next-Frame Diffusion…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Xinle Cheng , Tianyu He , Jiayi Xu , Junliang Guo , Di He , Jiang Bian

Surgical gesture recognition is important for surgical data science and computer-aided intervention. Even with robotic kinematic information, automatically segmenting surgical steps presents numerous challenges because surgical…

计算机视觉与模式识别 · 计算机科学 2020-03-11 Beatrice van Amsterdam , Matthew J. Clarkson , Danail Stoyanov

Intra-operative anticipation of instrument usage is a necessary component for context-aware assistance in surgery, e.g. for instrument preparation or semi-automation of robotic tasks. However, the sparsity of instrument occurrences in long…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Dominik Rivoir , Sebastian Bodenstedt , Isabel Funke , Felix von Bechtolsheim , Marius Distler , Jürgen Weitz , Stefanie Speidel
‹ 上一页 1 8 9 10 下一页 ›