English
Related papers

Related papers: BundleMoCap: Efficient, Robust and Smooth Motion C…

200 papers

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Zhengqi Li , Richard Tucker , Forrester Cole , Qianqian Wang , Linyi Jin , Vickie Ye , Angjoo Kanazawa , Aleksander Holynski , Noah Snavely

We introduce VideoMamba, a novel adaptation of the pure Mamba architecture, specifically designed for video recognition. Unlike transformers that rely on self-attention mechanisms leading to high computational costs by quadratic complexity,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Jinyoung Park , Hee-Seon Kim , Kangwook Ko , Minbeom Kim , Changick Kim

This paper proposes a new method for live free-viewpoint human performance capture with dynamic details (e.g., cloth wrinkles) using a single RGBD camera. Our main contributions are: (i) a multi-layer representation of garments and body,…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Tao Yu , Zerong Zheng , Yuan Zhong , Jianhui Zhao , Qionghai Dai , Gerard Pons-Moll , Yebin Liu

Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning to inject certain…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Pengyang Ling , Jiazi Bu , Pan Zhang , Xiaoyi Dong , Yuhang Zang , Tong Wu , Huaian Chen , Jiaqi Wang , Yi Jin

We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a setting that arises naturally when multiple observers capture a shared event. Prior approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Shuo Sun , Unal Artan , Malcolm Mielle , Achim J. Lilienthaland , Martin Magnusson

Video temporal grounding aims to pinpoint a video segment that matches the query description. Despite the recent advance in short-form videos (\textit{e.g.}, in minutes), temporal grounding in long videos (\textit{e.g.}, in hours) is still…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yulin Pan , Xiangteng He , Biao Gong , Yiliang Lv , Yujun Shen , Yuxin Peng , Deli Zhao

Despite significant progress, there remain three limitations to the previous multi-view clustering algorithms. First, they often suffer from high computational complexity, restricting their feasibility for large-scale datasets. Second, they…

Machine Learning · Computer Science 2023-01-25 Dong Huang , Chang-Dong Wang , Jian-Huang Lai

The idea of video super resolution is to use different view points of a single scene to enhance the overall resolution and quality. Classical energy minimization approaches first establish a correspondence of the current frame to all its…

Computer Vision and Pattern Recognition · Computer Science 2017-12-05 Jonas Geiping , Hendrik Dirks , Daniel Cremers , Michael Moeller

Recent studies on motion estimation have advocated an optimized motion representation that is globally consistent across the entire video, preferably for every pixel. This is challenging as a uniform representation may not account for the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Rui Li , Dong Liu

We present Capturing the Unseen (CAPUS), a novel facial motion capture (MoCap) technique that operates without visual signals. CAPUS leverages miniaturized Inertial Measurement Units (IMUs) as a new sensing modality for facial motion…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Youjia Wang , Yiwen Wu , Hengan Zhou , Hongyang Lin , Xingyue Peng , Jingyan Zhang , Yingsheng Zhu , Yingwenqi Jiang , Yatu Zhang , Lan Xu , Jingya Wang , Jingyi Yu

Diffusion models have made significant advances in generating high-quality images, but their application to video generation has remained challenging due to the complexity of temporal motion. Zero-shot video editing offers a solution by…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Xirui Li , Chao Ma , Xiaokang Yang , Ming-Hsuan Yang

Recently, large-scale pre-training methods like CLIP have made great progress in multi-modal research such as text-video retrieval. In CLIP, transformers are vital for modeling complex multi-modal relations. However, in the vision…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Shuai Zhao , Linchao Zhu , Xiaohan Wang , Yi Yang

As an essential part of structure from motion (SfM) and Simultaneous Localization and Mapping (SLAM) systems, motion averaging has been extensively studied in the past years and continues to attract surging research attention. While…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Xinyi Li , Lin Yuan , Longin Jan Latecki , Haibin Ling

Localization and mapping with heterogeneous multi-sensor fusion have been prevalent in recent years. To adequately fuse multi-modal sensor measurements received at different time instants and different frequencies, we estimate the…

Robotics · Computer Science 2023-02-16 Jiajun Lv , Xiaolei Lang , Jinhong Xu , Mengmeng Wang , Yong Liu , Xingxing Zuo

The exponential growth of video content has created an urgent need for efficient multimodal moment retrieval systems. However, existing approaches face three critical challenges: (1) fixed-weight fusion strategies fail across cross modal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Toan Le Ngo Thanh , Phat Ha Huu , Tan Nguyen Dang Duy , Thong Nguyen Le Minh , Anh Nguyen Nhu Tinh

Motion plays a crucial role in understanding videos and most state-of-the-art neural models for video classification incorporate motion information typically using optical flows extracted by a separate off-the-shelf method. As the…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Heeseung Kwon , Manjin Kim , Suha Kwak , Minsu Cho

This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task. The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying…

Computation and Language · Computer Science 2023-12-05 Keito Kudo , Haruki Nagasawa , Jun Suzuki , Nobuyuki Shimizu

We propose a real time deep learning framework for video-based facial expression capture. Our process uses a high-end facial capture pipeline based on FACEGOOD to capture facial expression. We train a convolutional neural network to produce…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Hongwei Xu , Leijia Dai , Jianxing Fu , Xiangyuan Wang , Quanwei Wang

We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Jiahui Lei , Yijia Weng , Adam Harley , Leonidas Guibas , Kostas Daniilidis

Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose…

Computer Vision and Pattern Recognition · Computer Science 2022-11-03 Yixuan Pei , Zhiwu Qing , Jun Cen , Xiang Wang , Shiwei Zhang , Yaxiong Wang , Mingqian Tang , Nong Sang , Xueming Qian