中文
相关论文

相关论文: MosaicMem: Hybrid Spatial Memory for Controllable …

200 篇论文

Pre-trained conditional diffusion models have demonstrated remarkable potential in image editing. However, they often face challenges with temporal consistency, particularly in the talking head domain, where continuous changes in facial…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

In this paper we start with a simple question, how is it possible that humans can recognize different movements over skin with only a prior visual experience of them? Or in general, what is the representation of spatial sequences that are…

人工智能 · 计算机科学 2023-11-14 Viacheslav M. Osaulenko

Generating high-quality videos that synthesize desired realistic content is a challenging task due to their intricate high-dimensionality and complexity of videos. Several recent diffusion-based methods have shown comparable performance by…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Kihong Kim , Haneol Lee , Jihye Park , Seyeon Kim , Kwanghee Lee , Seungryong Kim , Jaejun Yoo

Joint camera pose and dense geometry estimation from a set of images or a monocular video remains a challenging problem due to its computational complexity and inherent visual ambiguities. Most dense incremental reconstruction systems…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Kirill Mazur , Gwangbin Bae , Andrew J. Davison

Video decomposition is very important to extract moving foreground objects from complex backgrounds in computer vision, machine learning, and medical imaging, e.g., extracting moving contrast-filled vessels from the complex and noisy…

计算机视觉与模式识别 · 计算机科学 2022-05-09 Binjie Qin , Haohao Mao , Ruipeng Zhang , Yueqi Zhu , Song Ding , Xu Chen

Current video representations heavily rely on unstable and over-grained priors for motion and appearance modelling, \emph{i.e.}, pixel-level matching and tracking. A tracking error of just a few pixels would lead to the collapse of the…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Ye Chen , Liming Tan , Yupeng Zhu , Yuanbin Wang , Bingbing Ni

Structure-from-Motion (SfM), a task aiming at jointly recovering camera poses and 3D geometry of a scene given a set of images, remains a hard problem with still many open challenges despite decades of significant progress. The traditional…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Bardienus Duisterhof , Lojze Zust , Philippe Weinzaepfel , Vincent Leroy , Yohann Cabon , Jerome Revaud

Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as iterative shot…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Kaiwen Zhang , Liming Jiang , Angtian Wang , Jacob Zhiyuan Fang , Tiancheng Zhi , Qing Yan , Hao Kang , Xin Lu , Xingang Pan

Legged robots have the potential to expand the reach of autonomy beyond paved roads. In this work, we consider the difficult problem of locomotion on challenging terrains using a single forward-facing depth camera. Due to the partial…

机器人学 · 计算机科学 2023-04-04 Ruihan Yang , Ge Yang , Xiaolong Wang

In video analysis, background models have many applications such as background/foreground separation, change detection, anomaly detection, tracking, and more. However, while learning such a model in a video captured by a static camera is a…

计算机视觉与模式识别 · 计算机科学 2022-09-19 Guy Erez , Ron Shapira Weber , Oren Freifeld

Temporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersion of users. Despite its importance, this task remains…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Yun Wang , Junjie Hu , Qiaole Dong , Yongjian Zhang , Yanwei Fu , Tin Lun Lam , Dapeng Wu

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain unambiguous…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Tanzila Rahman , Hsin-Ying Lee , Jian Ren , Sergey Tulyakov , Shweta Mahajan , Leonid Sigal

Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simultaneous (i) near-perfect preservation of each item's identity,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Gemma Canet Tarrés , Manel Baradad , Francesc Moreno-Noguer , Yumeng Li

Masked image modeling (MIM) has emerged as a promising approach for pre-training Vision Transformers (ViTs). MIMs predict masked tokens token-wise to recover target signals that are tokenized from images or generated by pre-trained models…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Taekyung Kim , Byeongho Heo , Dongyoon Han

Sustaining high fidelity and high throughput of perception tasks over vision sensor streams on edge devices remains a formidable challenge, especially given the continuing increase in image sizes (e.g., generated by 4K cameras) and…

多媒体 · 计算机科学 2023-05-08 Ila Gokarn , Hemanth Sabella , Yigong Hu , Tarek Abdelzaher , Archan Misra

We propose POse-guided SElective Fusion (POSEFusion), a single-view human volumetric capture method that leverages tracking-based methods and tracking-free inference to achieve high-fidelity and dynamic 3D reconstruction. By contributing a…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Zhe Li , Tao Yu , Zerong Zheng , Kaiwen Guo , Yebin Liu

One key aspect of spatially aware robots is the ability to "find their bearings", ie. to correctly situate themselves in previously seen spaces. In this work, we focus on this particular scenario of continuous robotics operations, where…

机器人学 · 计算机科学 2025-10-24 Mert Bulent Sariyildiz , Philippe Weinzaepfel , Guillaume Bono , Gianluca Monaci , Christian Wolf

Video world models have shown immense promise for interactive simulation and entertainment, but current systems still struggle with two important aspects of interactivity: user control over the environment for reproducible, editable…

人工智能 · 计算机科学 2026-04-01 Ryan Po , David Junhao Zhang , Amir Hertz , Gordon Wetzstein , Neal Wadhwa , Nataniel Ruiz

In monocular videos that capture dynamic scenes, estimating the 3D geometry of video contents has been a fundamental challenge in computer vision. Specifically, the task is significantly challenged by the object motion, where existing…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Seong Hyeon Park , Jinwoo Shin