English
Related papers

Related papers: RAFT-MSF++: Temporal Geometry-Motion Feature Fusio…

200 papers

Earthquake hazard analysis and design of spatially distributed infrastructure, such as power grids and energy pipeline networks, require scenario-specific ground-motion time histories with realistic frequency content and spatiotemporal…

World models that forecast scene evolution by generating future video frames devote the bulk of their capacity to photometric details, yet the resulting predictions often remain geometrically inconsistent. We present VGGT-World, a geometry…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Xiangyu Sun , Shijie Wang , Fengyi Zhang , Lin Liu , Caiyan Jia , Ziying Song , Zi Huang , Yadan Luo

Recent advances in end-to-end unsupervised learning has significantly improved the performance of monocular depth prediction and alleviated the requirement of ground truth depth. Although a plethora of work has been done in enforcing…

Computer Vision and Pattern Recognition · Computer Science 2020-05-19 Vinay Kaushik , Brejesh Lall

Event cameras have shown promise in vision applications like optical flow estimation and stereo matching, with many specialized architectures leveraging the asynchronous and sparse nature of event data. However, existing works only focus…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Pengjie Zhang , Lin Zhu , Xiao Wang , Lizhi Wang , Wanxuan Lu , Hua Huang

For visual estimation of optical flow, a crucial function for many vision tasks, unsupervised learning, using the supervision of view synthesis has emerged as a promising alternative to supervised methods, since ground-truth flow is not…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Zitang Sun , Shin'ya Nishida , Zhengbo Luo

The precise fusion of computational fluid dynamic (CFD) data, wind tunnel tests data, and flight tests data in aerodynamic area is essential for obtaining comprehensive knowledge of both localized flow structures and global aerodynamic…

Machine Learning · Computer Science 2026-04-01 Qinye Zhu , Yu Xiang , Jun Zhang , Wenyong Wang

High-quality 4D reconstruction of human performance with complex interactions to various objects is essential in real-world scenarios, which enables numerous immersive VR/AR applications. However, recent advances still fail to provide…

Computer Vision and Pattern Recognition · Computer Science 2021-05-03 Zhuo Su , Lan Xu , Dawei Zhong , Zhong Li , Fan Deng , Shuxue Quan , Lu Fang

In multi-view medical diagnosis, deep learning-based models often fuse information from different imaging perspectives to improve diagnostic performance. However, existing approaches are prone to overfitting and rely heavily on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Jingyu Guo , Christos Matsoukas , Fredrik Strand , Kevin Smith

Point cloud registration has seen significant advancements with the application of deep learning techniques. However, existing approaches often overlook the potential of integrating radiometric information from RGB images. This limitation…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Zhaoyi Wang , Shengyu Huang , Jemil Avers Butt , Yuanzhou Cai , Matej Varga , Andreas Wieser

Radiance of real-world scenes typically spans a much wider dynamic range than what standard cameras can capture. While conventional HDR methods merge alternating-exposure frames, these approaches are inherently constrained to 2D pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Shin Dong-Yeon , Kim Jun-Seong , Kwon Byung-Ki , Tae-Hyun Oh

Image fusion is famous as an alternative solution to generate one high-quality image from multiple images in addition to image restoration from a single degraded image. The essence of image fusion is to integrate complementary information…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Pengwei Liang , Junjun Jiang , Qing Ma , Xianming Liu , Jiayi Ma

Multi-object tracking (MOT) in monocular videos is fundamentally challenged by occlusions and depth ambiguity, issues that conventional tracking-by-detection (TBD) methods struggle to resolve owing to a lack of geometric awareness. To…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xudong Han , Pengcheng Fang , Yueying Tian , Jianhui Yu , Xiaohao Cai , Daniel Roggen , Philip Birch

Multimodal large language models (MLLMs) are typically trained in multiple stages, with video-based supervised fine-tuning (Video-SFT) serving as a key step for improving visual understanding. Yet its effect on the fine-grained evolution of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Linghao Zhang , Jungang Li , Yonghua Hei , Sicheng Tao , Song Dai , Yibo Yan , Zihao Dongfang , Weiting Liu , Chenxi Qin , Hanqian Li , Xin Zou , Jiahao Zhang , Shuhang Xun , Haiyun Jiang , Xuming Hu

Generating high-dimensional visual modalities is a computationally intensive task. A common solution is progressive generation, where the outputs are synthesized in a coarse-to-fine spectral autoregressive manner. While diffusion models…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Moayed Haji-Ali , Willi Menapace , Ivan Skorokhodov , Arpit Sahni , Sergey Tulyakov , Vicente Ordonez , Aliaksandr Siarohin

Inferring spatial transcriptomics (ST) from histology enables scalable histogenomic profiling, yet current methods are largely restricted to single-tissue models. This fragmentation fails to leverage biological principles shared across…

Machine Learning · Computer Science 2026-05-13 Susu Hu , Stefanie Speidel

The goal of our work is to generate high-quality novel views from monocular videos of complex and dynamic scenes. Prior methods, such as DynamicNeRF, have shown impressive performance by leveraging time-varying dynamic radiation fields.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Xingyu Miao , Yang Bai , Haoran Duan , Yawen Huang , Fan Wan , Yang Long , Yefeng Zheng

Video-language models (VLMs) face rapid inference costs as visual token counts scale with video length. For example, 32 frames at $448{\times}448$ resolution already yield >8,000 visual tokens in Qwen3-VL, making LLM prefill the dominant…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Simin Huo , Ning LI

Crowd counting presents enormous challenges in the form of large variation in scales within images and across the dataset. These issues are further exacerbated in highly congested scenes. Approaches based on straightforward fusion of…

Computer Vision and Pattern Recognition · Computer Science 2019-08-30 Vishwanath A Sindagi , Vishal M. Patel

A number of deep learning based algorithms have been proposed to recover high-quality videos from low-quality compressed ones. Among them, some restore the missing details of each frame via exploring the spatiotemporal information of…

Image and Video Processing · Electrical Eng. & Systems 2021-08-13 Minyi Zhao , Yi Xu , Shuigeng Zhou

In this paper, we present a new self-supervised scene flow estimation approach for a pair of consecutive point clouds. The key idea of our approach is to represent discrete point clouds as continuous probability density functions using…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Pan He , Patrick Emami , Sanjay Ranka , Anand Rangarajan
‹ Prev 1 8 9 10 Next ›