English
Related papers

Related papers: MemoryDiorama: Generating Dynamic 3D Diorama from …

200 papers

Reconstructing dense geometry for dynamic scenes from a monocular video is a critical yet challenging task. Recent memory-based methods enable efficient online reconstruction, but they fundamentally suffer from a Memory Demand Dilemma: The…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Xudong Cai , Shuo Wang , Peng Wang , Yongcai Wang , Zhaoxin Fan , Wanting Li , Tianbao Zhang , Jianrong Tao , Yeying Jin , Deying Li

End-to-end learning of robot control policies, structured as neural networks, has emerged as a promising approach to robotic manipulation. To complete many common tasks, relevant objects are required to pass in and out of a robot's field of…

Deep neural networks have shown superior performance in many regimes to remember familiar patterns with large amounts of data. However, the standard supervised deep learning paradigm is still limited when facing the need to learn new…

Machine Learning · Computer Science 2018-11-16 Jing Shi , Jiaming Xu , Yiqun Yao , Bo Xu

3D object detection and occupancy prediction are critical tasks in autonomous driving, attracting significant attention. Despite the potential of recent vision-based methods, they encounter challenges under adverse conditions. Thus,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Lianqing Zheng , Jianan Liu , Runwei Guan , Long Yang , Shouyi Lu , Yuanzhe Li , Xiaokai Bai , Jie Bai , Zhixiong Ma , Hui-Liang Shen , Xichan Zhu

Agent-assisted memory recall is one critical research problem in the field of human-computer interaction. In conventional methods, the agent can retrieve information from its equipped memory module to help the person recall incomplete or…

Artificial Intelligence · Computer Science 2025-08-01 Qian Zhao , Zhuo Sun , Bin Guo , Zhiwen Yu

Generating realistic 3D worlds occupied by moving humans has many applications in games, architecture, and synthetic data creation. But generating such scenes is expensive and labor intensive. Recent work generates human poses and motions…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Hongwei Yi , Chun-Hao P. Huang , Shashank Tripathi , Lea Hering , Justus Thies , Michael J. Black

This paper presents BioNeRF, a biologically plausible architecture that models scenes in a 3D representation and synthesizes new views through radiance fields. Since NeRF relies on the network weights to store the scene's 3-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Leandro A. Passos , Douglas Rodrigues , Danilo Jodas , Kelton A. P. Costa , Ahsan Adeel , João Paulo Papa

Recent works have shown that it is possible to automatically predict intrinsic image properties like memorability. In this paper, we take a step forward addressing the question: "Can we make an image more memorable?". Methods for…

Computer Vision and Pattern Recognition · Computer Science 2017-04-07 Aliaksandr Siarohin , Gloria Zen , Cveta Majtanovic , Xavier Alameda-Pineda , Elisa Ricci , Nicu Sebe

In recent years, 3D visual foundation models pioneered by pointmap-based approaches such as DUSt3R have attracted a lot of interest, achieving impressive accuracy and strong generalization across diverse scenes. However, these methods are…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Shuang Guo , Filbert Febryanto , Lei Sun , Guillermo Gallego

While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space remains limited. To address this gap, we introduce geometrically referenced 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jiangye Yuan , Gowri Kumar , Baoyuan Wang

Realistic object interactions are crucial for creating immersive virtual experiences, yet synthesizing realistic 3D object dynamics in response to novel interactions remains a significant challenge. Unlike unconditional or text-conditioned…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Tianyuan Zhang , Hong-Xing Yu , Rundi Wu , Brandon Y. Feng , Changxi Zheng , Noah Snavely , Jiajun Wu , William T. Freeman

Humans and traditional computer vision methods rely on a diverse set of monocular cues to infer 3D structure from a single image, such as shading, texture, silhouette, etc. While recent deep generative models have dramatically advanced…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Xiang Li , Zirui Wang , Zixuan Huang , James M. Rehg

Image view synthesis has seen great success in reconstructing photorealistic visuals, thanks to deep learning and various novel representations. The next key step in immersive virtual experiences is view synthesis of dynamic scenes.…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Kai-En Lin , Guowei Yang , Lei Xiao , Feng Liu , Ravi Ramamoorthi

Recent studies on image memorability have shed light on the visual features that make generic images, object images or face photographs memorable. However, a clear understanding and reliable estimation of natural scene memorability remain…

Computer Vision and Pattern Recognition · Computer Science 2019-08-16 Jiaxin Lu , Mai Xu , Ren Yang , Zulin Wang

Current geometry-based monocular 3D object detection models can efficiently detect objects by leveraging perspective geometry, but their performance is limited due to the absence of accurate depth information. Though this issue can be…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Chenhang He , Jianqiang Huang , Xian-Sheng Hua , Lei Zhang

Light field cameras have been proved to be powerful tools for 3D reconstruction and virtual reality applications. However, the limited resolution of light field images brings a lot of difficulties for further information display and…

Image and Video Processing · Electrical Eng. & Systems 2020-08-27 Qingyan Sun , Shuo Zhang , Song Chang , Lixi Zhu , Youfang Lin

While computer vision models have made incredible strides in static image recognition, they still do not match human performance in tasks that require the understanding of complex, dynamic motion. This is notably true for real-world…

Neurons and Cognition · Quantitative Biology 2025-04-09 Jacob Yeung , Andrew F. Luo , Gabriel Sarch , Margaret M. Henderson , Deva Ramanan , Michael J. Tarr

The creation of accurate virtual models of real-world objects is imperative to robotic simulations and applications such as computer vision, artificial intelligence, and machine learning. This paper documents the different methods employed…

Robotics · Computer Science 2024-02-20 Nillan Nimal , Wenbin Li , Ronald Clark , Sajad Saeedi

Recently, Multimodal Large Language Models (MLLMs) have demonstrated significant potential in complex visual tasks through the integration of Chain-of-Thought (CoT) reasoning. However, in Video Question Answering, extended thinking…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Xiaokun Sun , Yubo Wang , Haoyu Cao , Linli Xu

Large language models (LLMs) face inherent limitations in memory, including restricted context windows, long-term knowledge forgetting, redundant information accumulation, and hallucination generation. These issues severely constrain…

Artificial Intelligence · Computer Science 2026-03-20 Deliang Wen , Ke Sun