中文
相关论文

相关论文: SurgSora: Object-Aware Diffusion Model for Control…

200 篇论文

Research in ultrasound imaging is limited in reproducibility by two factors: First, many existing ultrasound pipelines are protected by intellectual property, rendering exchange of code difficult. Second, most pipelines are implemented in…

计算机视觉与模式识别 · 计算机科学 2018-05-11 Rüdiger Göbl , Nassir Navab , Christoph Hennersperger

Recent advancements in diffusion techniques have propelled image and video generation to unprecedented levels of quality, significantly accelerating the deployment and application of generative AI. However, 3D shape generation technology…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Yangguang Li , Zi-Xin Zou , Zexiang Liu , Dehu Wang , Yuan Liang , Zhipeng Yu , Xingchao Liu , Yuan-Chen Guo , Ding Liang , Wanli Ouyang , Yan-Pei Cao

Text-guided generative diffusion models unlock powerful image creation and editing tools. While these have been extended to video generation, current approaches that edit the content of existing footage while retaining structure require…

计算机视觉与模式识别 · 计算机科学 2023-02-07 Patrick Esser , Johnathan Chiu , Parmida Atighehchian , Jonathan Granskog , Anastasis Germanidis

Modern video generation models like Sora have achieved remarkable success in producing high-quality videos. However, a significant limitation is their inability to offer interactive control to users, a feature that promises to open up…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Yash Jain , Anshul Nasery , Vibhav Vineet , Harkirat Behl

Diffusion-based video super-resolution (VSR) methods deliver strong perceptual quality but are often unsuitable for latency-sensitive scenarios due to reliance on future frames and expensive multi-step denoising. We propose Stream-DiffVSR,…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Hau-Shiang Shiu , Chin-Yang Lin , Zhixiang Wang , Chi-Wei Hsiao , Po-Fan Yu , Yu-Chih Chen , Yu-Lun Liu

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Korrawe Karunratanakul , Konpat Preechakul , Supasorn Suwajanakorn , Siyu Tang

While recent video generation models have achieved significant visual fidelity, they often suffer from the lack of explicit physical controllability and plausibility. To address this, some recent studies attempted to guide the video…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Haoze Zhang , Tianyu Huang , Zichen Wan , Xiaowei Jin , Hongzhi Zhang , Hui Li , Wangmeng Zuo

Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains.However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Tao Wu , Xuewei Li , Zhongang Qi , Di Hu , Xintao Wang , Ying Shan , Xi Li

Recent works on text-to-3d generation show that using only 2D diffusion supervision for 3D generation tends to produce results with inconsistent appearances (e.g., faces on the back view) and inaccurate shapes (e.g., animals with extra…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Cheng Chen , Xiaofeng Yang , Fan Yang , Chengzeng Feng , Zhoujie Fu , Chuan-Sheng Foo , Guosheng Lin , Fayao Liu

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Meng Wei , Kun Yuan , Shi Li , Yue Zhou , Long Bai , Nassir Navab , Hongliang Ren , Hong Joo Lee , Tom Vercauteren , Nicolas Padoy

Existing segmentation models trained on a single medical imaging dataset often lack robustness when encountering unseen organs or tumors. Developing a robust model capable of identifying rare or novel tumor categories not present during…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Rong Wu , Ziqi Chen , Liming Zhong , Heng Li , Hai Shu

Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions, including in surgical videos. While conventional…

图像与视频处理 · 电气工程与系统科学 2026-03-04 Guiqiu Liao , Matjaz Jogan , Marcel Hussing , Kenta Nakahashi , Kazuhiro Yasufuku , Amin Madani , Eric Eaton , Daniel A. Hashimoto

Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions…

计算机视觉与模式识别 · 计算机科学 2022-09-01 Mingyuan Zhang , Zhongang Cai , Liang Pan , Fangzhou Hong , Xinying Guo , Lei Yang , Ziwei Liu

In computer-assisted surgery, automatically recognizing anatomical organs is crucial for understanding the surgical scene and providing intraoperative assistance. While machine learning models can identify such structures, their deployment…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Danush Kumar Venkatesh , Dominik Rivoir , Micha Pfeiffer , Fiona Kolbinger , Stefanie Speidel

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Xingyuan Li , Zirui Wang , Yang Zou , Zhixin Chen , Jun Ma , Zhiying Jiang , Long Ma , Jinyuan Liu

Recent progress in 3D scene understanding enables scalable learning of representations across large datasets of diverse scenes. As a consequence, generalization to unseen scenes and objects, rendering novel views from just a single or a…

计算机视觉与模式识别 · 计算机科学 2024-05-06 Allan Jabri , Sjoerd van Steenkiste , Emiel Hoogeboom , Mehdi S. M. Sajjadi , Thomas Kipf

The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Zihan Su , Xuerui Qiu , Hongbin Xu , Tangyu Jiang , Junhao Zhuang , Chun Yuan , Ming Li , Shengfeng He , Fei Richard Yu

Foundation models in video generation are demonstrating remarkable capabilities as potential world models for simulating the physical world. However, their application in high-stakes domains like surgery, which demand deep, specialized…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Zhen Chen , Qing Xu , Jinlin Wu , Biao Yang , Yuhao Zhai , Geng Guo , Jing Zhang , Yinlu Ding , Nassir Navab , Jiebo Luo

Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Delong Liu , Zhaohui Hou , Mingjie Zhan , Shihao Han , Zhicheng Zhao , Fei Su

Surgery is a high-stakes domain where surgeons must navigate critical anatomical structures and actively avoid potential complications while achieving the main task at hand. Such surgical activity has been shown to affect long-term patient…