中文
相关论文

相关论文: L4P: Towards Unified Low-Level 4D Vision Perceptio…

200 篇论文

We address the unsupervised learning of several interconnected problems in low-level vision: single view depth prediction, camera motion estimation, optical flow, and segmentation of a video into the static scene and moving regions. Our key…

计算机视觉与模式识别 · 计算机科学 2019-03-13 Anurag Ranjan , Varun Jampani , Lukas Balles , Kihwan Kim , Deqing Sun , Jonas Wulff , Michael J. Black

Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained representations. Typically, existing methods embed 3D positions into…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Hanyu Zhou , Chuanhao Ma , Gim Hee Lee

Dense 4D reconstruction from unposed images remains a critical challenge, with current methods relying on slow test-time optimization or fragmented, task-specific feedforward models. We introduce UFO-4D, a unified feedforward framework to…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Junhwa Hur , Charles Herrmann , Songyou Peng , Philipp Henzler , Zeyu Ma , Todd Zickler , Deqing Sun

Dynamic urban environments are often captured by cameras placed at spatially separated locations with little or no view overlap. However, most existing 4D reconstruction methods assume densely overlapping views. When applied to such sparse…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Hina Kogure , Kei Katsumata , Taiki Miyanishi , Komei Sugiura

Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Chunyang Cheng , Tianyang Xu , Zhenhua Feng , Xiaojun Wu , ZhangyongTang , Hui Li , Zeyang Zhang , Sara Atito , Muhammad Awais , Josef Kittler

The landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to bridge understanding and generation tasks. Recent advances,…

Autonomous driving systems require a comprehensive understanding of the environment, achieved by extracting visual features essential for perception, planning, and control. However, models trained solely on single-task objectives or generic…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Huy-Dung Nguyen , Anass Bairouk , Mirjana Maras , Wei Xiao , Tsun-Hsuan Wang , Patrick Chareyre , Ramin Hasani , Marc Blanchon , Daniela Rus

There has been significant progresses for image object detection in recent years. Nevertheless, video object detection has received little attention, although it is more challenging and more important in practical scenarios. Built upon the…

计算机视觉与模式识别 · 计算机科学 2017-12-01 Xizhou Zhu , Jifeng Dai , Lu Yuan , Yichen Wei

Feature matching plays a fundamental role in many computer vision tasks, yet existing methods heavily rely on scarce and clean multi-view image collections, which constrains their generalization to diverse and challenging scenarios.…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yingping Liang , Yutao Hu , Wenqi Shao , Ying Fu

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, generalization remains a persistent challenge, particularly when…

机器人学 · 计算机科学 2025-02-27 Zerui Li , Gengze Zhou , Haodong Hong , Yanyan Shao , Wenqi Lyu , Yanyuan Qiao , Qi Wu

Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is challenging to design a deep learning-based perception…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Meiqi Sun , Zhonghan Zhao , Wenhao Chai , Hanjun Luo , Shidong Cao , Yanting Zhang , Jenq-Neng Hwang , Gaoang Wang

In human-centered environments such as restaurants, homes, and warehouses, robots often face challenges in accurately recognizing 3D objects. These challenges stem from the complexity and variability of these environments, including diverse…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Songsong Xiong , Hamidreza Kasaei

We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Kuan Heng Lin , Zhizheng Liu , Pablo Salamanca , Yash Kant , Ryan Burgert , Yuancheng Xu , Koichi Namekata , Yiwei Zhao , Bolei Zhou , Micah Goldblum , Paul Debevec , Ning Yu

Vision-Language Pre-training (VLP) has achieved impressive performance on various cross-modal downstream tasks. However, most existing methods can only learn from aligned image-caption data and rely heavily on expensive regional features,…

计算机视觉与模式识别 · 计算机科学 2022-03-18 Wei Li , Can Gao , Guocheng Niu , Xinyan Xiao , Hao Liu , Jiachen Liu , Hua Wu , Haifeng Wang

We pose a new problem, In-2-4D, for generative 4D (i.e., 3D + motion) inbetweening to interpolate two single-view images. In contrast to video/4D generation from only text or a single image, our interpolative task can leverage more precise…

图形学 · 计算机科学 2025-09-30 Sauradip Nag , Daniel Cohen-Or , Hao Zhang , Ali Mahdavi-Amiri

Amidst the rapid advancement of camera-based autonomous driving technology, effectiveness is often prioritized with limited attention to computational efficiency. To address this issue, this paper introduces LRHPerception, a real-time…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Haixi Zhang , Aiyinsi Zuo , Zirui Li , Chunshu Wu , Tong Geng , Zhiyao Duan

Video analysis tasks rely heavily on identifying the pixels from different frames that correspond to the same visual target. To tackle this problem, recent studies have advocated feature learning methods that aim to learn distinctive…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Rui Li , Shenglong Zhou , Dong Liu

Multi-view image compression plays a critical role in 3D-related applications. Existing methods adopt a predictive coding architecture, which requires joint encoding to compress the corresponding disparity as well as residual information.…

图像与视频处理 · 电气工程与系统科学 2023-04-13 Xinjie Zhang , Jiawei Shao , Jun Zhang

Accurate reconstruction and tracking of dynamic human faces from image sequences is challenging because non-rigid deformations, expression changes, and viewpoint variations occur simultaneously, creating significant ambiguity in geometry…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Umut Kocasari , Simon Giebenhain , Richard Shaw , Matthias Nießner

Reconstructing 4D dynamic scenes from casually captured monocular videos is valuable but highly challenging, as each timestamp is observed from a single viewpoint. We introduce Vivid4D, a novel approach that enhances 4D monocular video…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Jiaxin Huang , Sheng Miao , BangBang Yang , Yuewen Ma , Yiyi Liao