English
Related papers

Related papers: Neural Multisensory Scene Inference

200 papers

Given large amount of real photos for training, Convolutional neural network shows excellent performance on object recognition tasks. However, the process of collecting data is so tedious and the background are also limited which makes it…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Yida Wang , Weihong Deng

Generalizable person re-identification (Re-ID) is a very hot research topic in machine learning and computer vision, which plays a significant role in realistic scenarios due to its various applications in public security and video…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Suncheng Xiang , Jingsheng Gao , Mengyuan Guan , Jiacheng Ruan , Chengfeng Zhou , Ting Liu , Dahong Qian , Yuzhuo Fu

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Wanhee Lee , Klemen Kotar , Rahul Mysore Venkatesh , Jared Watrous , Honglin Chen , Khai Loong Aw , Daniel L. K. Yamins

We revisit human motion synthesis, a task useful in various real world applications, in this paper. Whereas a number of methods have been developed previously for this task, they are often limited in two aspects: focusing on the poses while…

Computer Vision and Pattern Recognition · Computer Science 2021-06-01 Jingbo Wang , Sijie Yan , Bo Dai , Dahua LIn

Novel view synthesis and 3D modeling using implicit neural field representation are shown to be very effective for calibrated multi-view cameras. Such representations are known to benefit from additional geometric and semantic supervision.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Nikola Popovic , Danda Pani Paudel , Luc Van Gool

Recent advancements in 3D object generation using diffusion models have achieved remarkable success, but generating realistic 3D urban scenes remains challenging. Existing methods relying solely on 3D diffusion models tend to suffer a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Hanlei Guo , Jiahao Shao , Xinya Chen , Xiyang Tan , Sheng Miao , Yujun Shen , Yiyi Liao

This paper discusses the benefits of incorporating multimodal data for improving latent emotion recognition accuracy, focusing on micro-expression (ME) and physiological signals (PS). The proposed approach presents a novel multimodal…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Liangfei Zhang , Yifei Qian , Ognjen Arandjelovic , Anthony Zhu

We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception. By integrating 3D Gaussian Splatting techniques with…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Xueyan Zou , Yuchen Song , Ri-Zhao Qiu , Xuanbin Peng , Jianglong Ye , Sifei Liu , Xiaolong Wang

Recently neural scene representations have provided very impressive results for representing 3D scenes visually, however, their study and progress have mainly been limited to visualization of virtual models in computer graphics or scene…

Computer Vision and Pattern Recognition · Computer Science 2022-09-26 Yassine Ahmine , Arnab Dey , Andrew I. Comport

3D reconstruction and simulation, although interrelated, have distinct objectives: reconstruction requires a flexible 3D representation that can adapt to diverse scenes, while simulation needs a structured representation to model motion…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Shaojie Ma , Yawei Luo , Wei Yang , Yi Yang

Multimodal representation learning has shown promising improvements on various vision-language tasks. Most existing methods excel at building global-level alignment between vision and language while lacking effective fine-grained image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Zijia Zhao , Longteng Guo , Xingjian He , Shuai Shao , Zehuan Yuan , Jing Liu

In this paper, we present a method to reconstruct the world and multiple dynamic humans in 3D from a monocular video input. As a key idea, we represent both the world and multiple humans via the recently emerging 3D Gaussian Splatting…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Inhee Lee , Byungjun Kim , Hanbyul Joo

Realizing dexterous embodied manipulation necessitates the deep integration of heterogeneous multimodal sensory inputs. However, current vision-centric paradigms often overlook the critical force and geometric feedback essential for complex…

Robotics · Computer Science 2026-02-24 Yirui Sun , Guangyu Zhuge , Keliang Liu , Jie Gu , Zhihao xia , Qionglin Ren , Chunxu tian , Zhongxue Ga

3D Gaussian splatting (3DGS) has demonstrated exceptional performance in image-based 3D reconstruction and real-time rendering. However, regions with complex textures require numerous Gaussians to capture significant color variations…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Binxiao Huang , Zhihao Li , Shiyong Liu , Xiao Tang , Jiajun Tang , Jiaqi Lin , Yuxin Cheng , Zhenyu Chen , Xiaofei Wu , Ngai Wong

Masked Autoencoders (MAE) play a pivotal role in learning potent representations, delivering outstanding results across various 3D perception tasks essential for autonomous driving. In real-world driving scenarios, it's commonplace to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Jian Zou , Tianyu Huang , Guanglei Yang , Zhenhua Guo , Tao Luo , Chun-Mei Feng , Wangmeng Zuo

Understanding the geometric relationships between objects in a scene is a core capability in enabling both humans and autonomous agents to navigate in new environments. A sparse, unified representation of the scene topology will allow…

Computer Vision and Pattern Recognition · Computer Science 2022-05-18 Zachary Seymour , Niluthpol Chowdhury Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Learning generative models that span multiple data modalities, such as vision and language, is often motivated by the desire to learn more useful, generalisable representations that faithfully capture common underlying factors between the…

Machine Learning · Statistics 2019-11-11 Yuge Shi , N. Siddharth , Brooks Paige , Philip H. S. Torr

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Tianchen Deng , Xuefeng Chen , Yi Chen , Qu Chen , Yuyao Xu , Lijin Yang , Le Xu , Yu Zhang , Bo Zhang , Wuxiong Huang , Hesheng Wang

In this work, we present SceneDreamer, an unconditional generative model for unbounded 3D scenes, which synthesizes large-scale 3D landscapes from random noise. Our framework is learned from in-the-wild 2D image collections only, without…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Zhaoxi Chen , Guangcong Wang , Ziwei Liu

This position paper argues for the use of \emph{structured generative models} (SGMs) for the understanding of static scenes. This requires the reconstruction of a 3D scene from an input image (or a set of multi-view images), whereby the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Christopher K. I. Williams