中文
相关论文

相关论文: Buffer Anytime: Zero-Shot Video Depth and Normal f…

200 篇论文

Optical flow estimation is a crucial subfield of computer vision, serving as a foundation for video tasks. However, the real-world robustness is limited by animated synthetic datasets for training. This introduces domain gaps when applied…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Yingping Liang , Ying Fu , Yutao Hu , Wenqi Shao , Jiaming Liu , Debing Zhang

Generic Boundary Detection (GBD) aims at locating the general boundaries that divide videos into semantically coherent and taxonomy-free units, and could serve as an important pre-processing step for long-form video understanding. Previous…

计算机视觉与模式识别 · 计算机科学 2022-10-06 Jing Tan , Yuhong Wang , Gangshan Wu , Limin Wang

Segment Anything (SAM), an advanced universal image segmentation model trained on an expansive visual dataset, has set a new benchmark in image segmentation and computer vision. However, it faced challenges when it came to distinguishing…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Xiao Feng Zhang , Tian Yi Song , Jia Wei Yao

One of the solutions of depth imaging of moving scene is to project a static pattern on the object and use just a single image for reconstruction. However, if the motion of the object is too fast with respect to the exposure time of the…

计算机视觉与模式识别 · 计算机科学 2017-10-03 Yuki Shiba , Satoshi Ono , Ryo Furukawa , Shinsaku Hiura , Hiroshi Kawasaki

Recent developments in generative diffusion models have turned many dreams into realities. For video object insertion, existing methods typically require additional information, such as a reference video or a 3D asset of the object, to…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Qi Zhao , Zhan Ma , Pan Zhou

Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand massive labeled datasets to resolve semantic ambiguities.…

Recent advancements in text-to-image generation have enabled significant progress in zero-shot 3D shape generation. This is achieved by score distillation, a methodology that uses pre-trained text-to-image diffusion models to optimize the…

计算机视觉与模式识别 · 计算机科学 2023-05-29 Zhenzhen Weng , Zeyu Wang , Serena Yeung

Stereo matching provides depth estimation from binocular images for downstream applications. These applications mostly take video streams as input and require temporally consistent depth maps. However, existing methods mainly focus on the…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Jiaxi Zeng , Chengtang Yao , Yuwei Wu , Yunde Jia

Operating effectively in novel real-world environments requires robotic systems to estimate and interact with previously unseen objects. Current state-of-the-art models address this challenge by using large amounts of training data and…

机器人学 · 计算机科学 2026-02-06 Octavio Arriaga , Proneet Sharma , Jichen Guo , Marc Otto , Siddhant Kadwe , Rebecca Adam

In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are…

计算机视觉与模式识别 · 计算机科学 2015-05-05 Vignesh Ramanathan , Kevin Tang , Greg Mori , Li Fei-Fei

By training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhibit a superior…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Xiang Zhang , Bingxin Ke , Hayko Riemenschneider , Nando Metzger , Anton Obukhov , Markus Gross , Konrad Schindler , Christopher Schroers

When trained on large-scale datasets, image captioning models can understand the content of images from a general domain but often fail to generate accurate, detailed captions. To improve performance, pretraining-and-finetuning has been a…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Taehoon Kim , Mark Marsden , Pyunghwan Ahn , Sangyun Kim , Sihaeng Lee , Alessandra Sala , Seung Hwan Kim

Efficiently predicting motion plans directly from vision remains a fundamental challenge in robotics, where planning typically requires explicit goal specification and task-specific design. Recent vision-language-action (VLA) models infer…

Multi-Camera arrays are increasingly employed in both consumer and industrial applications, and various passive techniques are documented to estimate depth from such camera arrays. Current depth estimation methods provide useful estimations…

计算机视觉与模式识别 · 计算机科学 2018-06-22 Hossein Javidnia , Peter Corcoran

In this paper, we present a novel zero-shot camera calibration method that estimates camera parameters with no calibration image. It is common sense that we need at least one or more pattern images for camera calibration. However, the…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Jae-Yeong Lee

Tracking cells and detecting mitotic events in time-lapse microscopy image sequences is a crucial task in biomedical research. However, it remains highly challenging due to dividing objects, low signal-tonoise ratios, indistinct boundaries,…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Zhu Chen , Mert Edgü , Er Jin , Johannes Stegmaier

This paper aims at demystifying a single motion-blurred image with events and revealing temporally continuous scene dynamics encrypted behind motion blurs. To achieve this end, an Implicit Video Function (IVF) is learned to represent a…

计算机视觉与模式识别 · 计算机科学 2023-04-07 Zhangyi Cheng , Xiang Zhang , Lei Yu , Jianzhuang Liu , Wen Yang , Gui-Song Xia

Composed Video Retrieval (CoVR) retrieves a target video given a query video and a modification text describing the intended change. Existing CoVR benchmarks emphasize appearance shifts or coarse event changes and therefore do not test the…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Animesh Gupta , Jay Parmar , Ishan Rajendrakumar Dave , Mubarak Shah

Recovering the camera motion and scene geometry from visual data is a fundamental problem in the field of computer vision. Its success in standard vision is attributed to the maturity of feature extraction, data association and multi-view…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Zhongyang Ren , Bangyan Liao , Delei Kong , Jinghang Li , Peidong Liu , Laurent Kneip , Guillermo Gallego , Yi Zhou

Acquiring accurate depth information of transparent objects using off-the-shelf RGB-D cameras is a well-known challenge in Computer Vision and Robotics. Depth estimation/completion methods are typically employed and trained on datasets with…

机器人学 · 计算机科学 2024-03-29 Avinash Ummadisingu , Jongkeum Choi , Koki Yamane , Shimpei Masuda , Naoki Fukaya , Kuniyuki Takahashi
‹ 上一页 1 8 9 10 下一页 ›