中文
相关论文

相关论文: DepthPilot: From Controllability to Interpretabili…

200 篇论文

Fetal ultrasound (US) examinations require the acquisition of multiple planes, each providing unique diagnostic information to evaluate fetal development and screening for congenital anomalies. However, obtaining a comprehensive,…

Many robotic tasks involving some form of 3D visual perception greatly benefit from a complete knowledge of the working environment. However, robots often have to tackle unstructured environments and their onboard visual sensors can only…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Andrea Rosasco , Stefano Berti , Fabrizio Bottarel , Michele Colledanchise , Lorenzo Natale

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip. In this paper, for the sake of both furthering this…

计算机视觉与模式识别 · 计算机科学 2018-08-10 Lijie Fan , Wenbing Huang , Chuang Gan , Junzhou Huang , Boqing Gong

Achieving precise camera control in video generation remains challenging, as existing methods often rely on camera pose annotations that are difficult to scale to large and dynamic datasets and are frequently inconsistent with depth…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Zelin Zhao , Xinyu Gong , Bangya Liu , Ziyang Song , Jun Zhang , Suhui Wu , Yongxin Chen , Hao Zhang

High-fidelity reconstruction of deformable tissues from endoscopic videos remains challenging due to the limitations of existing methods in capturing subtle color variations and modeling global deformations. While 3D Gaussian Splatting…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Qun Ji , Peng Li , Mingqiang Wei

We introduce ColonSLAM, a system that combines classical multiple-map metric SLAM with deep features and topological priors to create topological maps of the whole colon. The SLAM pipeline by itself is able to create disconnected individual…

计算机视觉与模式识别 · 计算机科学 2024-09-26 Javier Morlana , Juan D. Tardós , José M. M. Montiel

Foundation models in video generation are demonstrating remarkable capabilities as potential world models for simulating the physical world. However, their application in high-stakes domains like surgery, which demand deep, specialized…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Zhen Chen , Qing Xu , Jinlin Wu , Biao Yang , Yuhao Zhai , Geng Guo , Jing Zhang , Yinlu Ding , Nassir Navab , Jiebo Luo

In recent years, convolutional neural networks have demonstrated promising performance in a variety of medical image segmentation tasks. However, when a trained segmentation model is deployed into the real clinical world, the model may not…

图像与视频处理 · 电气工程与系统科学 2020-12-24 Shuo Wang , Giacomo Tarroni , Chen Qin , Yuanhan Mo , Chengliang Dai , Chen Chen , Ben Glocker , Yike Guo , Daniel Rueckert , Wenjia Bai

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

机器人学 · 计算机科学 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

Colorectal cancer remains one of the deadliest cancers in the world. In recent years computer-aided methods have aimed to enhance cancer screening and improve the quality and availability of colonoscopies by automatizing sub-tasks. One such…

计算机视觉与模式识别 · 计算机科学 2023-10-04 Anita Rau , Binod Bhattarai , Lourdes Agapito , Danail Stoyanov

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Purpose: Surgical scene understanding plays a critical role in the technology stack of tomorrow's intervention-assisting systems in endoscopic surgeries. For this, tracking the endoscope pose is a key component, but remains challenging due…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Michel Hayoz , Christopher Hahne , Mathias Gallardo , Daniel Candinas , Thomas Kurmann , Maximilian Allan , Raphael Sznitman

Automatic colorectal polyp detection in colonoscopy video is a fundamental task, which has received a lot of attention. Manually annotating polyp region in a large scale video dataset is time-consuming and expensive, which limits the…

图像与视频处理 · 电气工程与系统科学 2021-01-01 Zhi-Qin Zhan , Huazhu Fu , Yan-Yao Yang , Jingjing Chen , Jie Liu , Yu-Gang Jiang

Capsule endoscopy is an evolutional technique for examining and diagnosing intractable gastrointestinal diseases. Because of the huge amount of data, analyzing capsule endoscope videos is very time-consuming and labor-intensive for…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Xinkai Zhao , Chaowei Fang , Feng Gao , De-Jun Fan , Xutao Lin , Guanbin Li

Echocardiography interpretation requires integrating multi-view temporal evidence with quantitative measurements and guideline-grounded reasoning, yet existing foundation-model pipelines largely solve isolated subtasks and fail when tool…

Endoscopic procedures such as esophagogastroduodenoscopy (EGD) and colonoscopy play a critical role in diagnosing and managing gastrointestinal (GI) disorders. However, the documentation burden associated with these procedures place…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Evandros Kaklamanos , Kristjana Kristinsdottir , Jonathan Huang , Dustin Carlson , Rajesh Keswani , John Pandolfino , Mozziyar Etemadi

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incoherent content that…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhexiao Xiong , Yizhi Song , Liu He , Wei Xiong , Yu Yuan , Feng Qiao , Nathan Jacobs

Detection and diagnosis of colon polyps are key to preventing colorectal cancer. Recent evidence suggests that AI-based computer-aided detection (CADe) and computer-aided diagnosis (CADx) systems can enhance endoscopists' performance and…

图像与视频处理 · 电气工程与系统科学 2024-03-05 Carlo Biffi , Giulio Antonelli , Sebastian Bernhofer , Cesare Hassan , Daizen Hirata , Mineo Iwatate , Andreas Maieron , Pietro Salvagnini , Andrea Cherubini

Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision-Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge. Existing 3D…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Haijier Chen , Bo Xu , Shoujian Zhang , Haoze Liu , Jiaxuan Lin , Jingrong Wang

Robotic and autonomous systems need dense spatial cues, but many monocular depth models are heavy, task-specific, or hard to attach to an existing multimodal stack. CLIP offers strong semantic representations, yet most CLIP-based depth…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Taewan Cho , Taeryang Kim , Andrew Jaeyong Choi