中文
相关论文

相关论文: Render, Don't Decode: Weight-Space World Models wi…

200 篇论文

Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Thomas Ressler-Antal , Frank Fundel , Malek Ben Alaya , Stefan Andreas Baumann , Felix Krause , Ming Gui , Björn Ommer

Self-supervised learning methods overcome the key bottleneck for building more capable AI: limited availability of labeled data. However, one of the drawbacks of self-supervised architectures is that the representations that they learn are…

机器学习 · 计算机科学 2022-07-08 Avi Ziskind , Sujeong Kim , Giedrius T. Burachas

CodeNeRF is an implicit 3D neural representation that learns the variation of object shapes and textures across a category and can be trained, from a set of posed images, to synthesize novel views of unseen objects. Unlike the original…

图形学 · 计算机科学 2021-09-07 Wonbong Jang , Lourdes Agapito

Self-supervised learning (SSL) has made rapid progress, yet learned features often over-rely on contextual shortcuts-background textures and co-occurrence statistics. While video provides rich temporal variation, dense in-the-wild streams…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Seul-Ki Yeom , Marcel Simon , Eunbin Lee , Tae-Ho Kim

Neural Representations for Videos(NeRV) have emerged as a promising paradigm for video compression by representing videos as compact neural networks with efficient decoding. Hybrid NeRV methods further improve reconstruction quality through…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yunjie Xu , Xiang Feng , Chengkai Wang , Alan Wee-Chung Liew , Xuefei Yin , Yanming Zhu

Generating realistic intermediate shapes between non-rigidly deformed shapes is a challenging task in computer vision, especially with unstructured data (e.g., point clouds) where temporal consistency across frames is lacking, and…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Lu Sang , Zehranaz Canfes , Dongliang Cao , Riccardo Marin , Florian Bernard , Daniel Cremers

Human perceive the 3D world through 2D observations from limited viewpoints. While recent feed-forward generalizable 3D reconstruction models excel at recovering 3D structures from sparse images, their representations are often confined to…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Mochu Xiang , Zhelun Shen , Xuesong Li , Jiahui Ren , Jing Zhang , Chen Zhao , Shanshan Liu , Haocheng Feng , Jingdong Wang , Yuchao Dai

Identifying systems with high-dimensional inputs and outputs, such as systems measured by video streams, is a challenging problem with numerous applications in robotics, autonomous vehicles and medical imaging. In this paper, we propose a…

系统与控制 · 电气工程与系统科学 2021-05-11 Gerben Izaak Beintema , Roland Toth , Maarten Schoukens

An implicit neural representation (INR) is a neural network that approximates a spatiotemporal function. Many memory-intensive visualization tasks, including modern 4D CT scanning methods, represent data natively as INRs. While INRs are…

机器学习 · 计算机科学 2025-12-03 Jennifer Zvonek , Andrew Gillette

We propose a framework for aligning and fusing multiple images into a single view using neural image representations (NIRs), also known as implicit or coordinate-based neural representations. Our framework targets burst images that exhibit…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Seonghyeon Nam , Marcus A. Brubaker , Michael S. Brown

We present a new data-driven reduced-order modeling approach to efficiently solve parametrized partial differential equations (PDEs) for many-query problems. This work is inspired by the concept of implicit neural representation (INR),…

数值分析 · 数学 2023-11-30 Tianshu Wen , Kookjin Lee , Youngsoo Choi

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Xinyu Zhang , Zhengtong Xu , Yutian Tao , Yeping Wang , Yu She , Abdeslam Boularias

Despite advancements in Neural Implicit models for 3D surface reconstruction, handling dynamic environments with interactions between arbitrary rigid, non-rigid, or deformable entities remains challenging. The generic reconstruction methods…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Sandika Biswas , Qianyi Wu , Biplab Banerjee , Hamid Rezatofighi

We present a simple neural rendering architecture that helps variational autoencoders (VAEs) learn disentangled representations. Instead of the deconvolutional network typically used in the decoder of VAEs, we tile (broadcast) the latent…

机器学习 · 计算机科学 2019-08-15 Nicholas Watters , Loic Matthey , Christopher P. Burgess , Alexander Lerchner

The recent success of implicit neural scene representations has presented a viable new method for how we capture and store 3D scenes. Unlike conventional 3D representations, such as point clouds, which explicitly store scene properties in…

计算机视觉与模式识别 · 计算机科学 2021-01-19 Amit Kohli , Vincent Sitzmann , Gordon Wetzstein

Rendering diffuse global illumination in real-time is often approximated by pre-computing and storing irradiance in a 3D grid of probes. As long as most of the scene remains static, probes approximate irradiance for all surfaces immersed in…

Infrared dim and small target detection presents a significant challenge due to dynamic multi-frame scenarios and weak target signatures in the infrared modality. Traditional low-rank plus sparse models often fail to capture dynamic…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Pei Liu , Yisi Luo , Wenzhen Wang , Xiangyong Cao

Implicit Neural Video Representation (INVR) has emerged as a novel approach for video representation and compression, using learnable grids and neural networks. Existing methods focus on developing new grid structures efficient for latent…

图像与视频处理 · 电气工程与系统科学 2026-04-09 Yiyang Li , Yanbo Gao , Shuai Li , Zhenyu Du , Jinglin Zhang , Hui Yuan , Mao Ye , Xingyu Gao

This work considers identifying parameters characterizing a physical system's dynamic motion directly from a video whose rendering configurations are inaccessible. Existing solutions require massive training data or lack generalizability to…

计算机视觉与模式识别 · 计算机科学 2022-05-12 Pingchuan Ma , Tao Du , Joshua B. Tenenbaum , Wojciech Matusik , Chuang Gan

Representations of the world environment play a crucial role in artificial intelligence. It is often inefficient to conduct reasoning and inference directly in the space of raw sensory representations, such as pixel values of images.…

机器学习 · 计算机科学 2022-04-12 Kenji Kawaguchi , Linjun Zhang , Zhun Deng