中文
相关论文

相关论文: Cosmos-Transfer1: Conditional World Generation wit…

200 篇论文

Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However, these models primarily animate static images without…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Luis Denninger , Sina Mokhtarzadeh Azar , Juergen Gall

Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Pengze Zhang , Yanze Wu , Mengtian Li , Xu Bai , Songtao Zhao , Fulong Ye , Chong Mou , Xinghui Li , Zhuowei Chen , Qian He , Mingyuan Gao

The capability of imagining internally with a mental model of the world is vitally important for human cognition. If a machine intelligent agent can learn a world model to create a "dream" environment, it can then internally ask what-if…

机器学习 · 计算机科学 2020-12-29 Minne Li , Mengyue Yang , Furui Liu , Xu Chen , Zhitang Chen , Jun Wang

A plant growth simulation can be characterized as a reconstructed visual representation of a plant or plant system. The phenotypic characteristics and plant structures are controlled by the scene environment and other contextual attributes.…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Mohamed Debbagh , Yixue Liu , Zhouzhou Zheng , Xintong Jiang , Shangpeng Sun , Mark Lefsrud

Novel-view synthesis techniques achieve impressive results for static scenes but struggle when faced with the inconsistencies inherent to casual capture settings: varying illumination, scene motion, and other unintended effects that are…

We present an overview and evaluation of a new, systematic approach for generation of highly realistic, annotated synthetic data for training of deep neural networks in computer vision tasks. The main contribution is a procedural world…

计算机视觉与模式识别 · 计算机科学 2017-10-19 Apostolia Tsirikoglou , Joel Kronander , Magnus Wrenninge , Jonas Unger

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the…

声音 · 计算机科学 2023-12-29 Zhifang Guo , Jianguo Mao , Rui Tao , Long Yan , Kazushige Ouchi , Hong Liu , Xiangdong Wang

Synthesizing controllable 6-DOF object manipulation trajectories in 3D environments is essential for enabling robots to interact with complex scenes, yet remains challenging due to the need for accurate spatial reasoning, physical…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Huajian Zeng , Abhishek Saroha , Daniel Cremers , Xi Wang

Conditional domain generation is a good way to interactively control sample generation process of deep generative models. However, once a conditional generative model has been created, it is often expensive to allow it to adapt to new…

机器学习 · 计算机科学 2018-05-28 Yingjing Lu

Illumination and texture editing are critical dimensions for world-to-world transfer, which is valuable for applications including sim2real and real2real visual data scaling up for embodied AI. Existing techniques generatively re-render the…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Yang Liu , Chuanchen Luo , Zimo Tang , Yingyan Li , Yuran Yang , Yuanyong Ning , Lue Fan , Junran Peng , Zhaoxiang Zhang

Integrated sensing and communications is a key enabler for the 6G wireless communication systems. The multiple sensing modalities will allow the base station to have a more accurate representation of the environment, leading to…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Mohammad Farzanullah , Han Zhang , Akram Bin Sediq , Ali Afana , Melike Erol-Kantarci

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a…

机器人学 · 计算机科学 2026-04-14 Liaoyuan Fan , Zetian Xu , Chen Cao , Wenyao Zhang , Mingqi Yuan , Jiayu Chen

Remote sensing world models aim to both explain observed changes and forecast plausible futures, two tasks that share spatiotemporal priors. Existing methods, however, typically address them separately, limiting cross-task transfer. We…

人工智能 · 计算机科学 2026-03-17 Linrui Xu , Zhongan Wang , Fei Shen , Gang Xu , Huiping Zhuang , Ming Li , Haifeng Li

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthesize only limited observations of a scene, leading to issues…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Yuheng Liu , Xin Lin , Xinke Li , Baihan Yang , Chen Wang , Kalyan Sunkavalli , Yannick Hold-Geoffroy , Hao Tan , Kai Zhang , Xiaohui Xie , Zifan Shi , Yiwei Hu

Human beings learn causal models and constantly use them to transfer knowledge between similar environments. We use this intuition to design a transfer-learning framework using object-oriented representations to learn the causal…

机器学习 · 计算机科学 2020-07-21 Purva Pruthi , Javier González , Xiaoyu Lu , Madalina Fiterau

We present Blocks2World, a novel method for 3D scene rendering and editing that leverages a two-step process: convex decomposition of images and conditioned synthesis. Our technique begins by extracting 3D parallelepipeds from various…

计算机视觉与模式识别 · 计算机科学 2023-07-14 Vaibhav Vavilala , Seemandhar Jain , Rahul Vasanth , Anand Bhattad , David Forsyth

We propose a new generative model of projected cosmic mass density maps inferred from weak gravitational lensing observations of distant galaxies (weak lensing mass maps). We construct the model based on a neural style transfer so that it…

宇宙学与河外天体物理 · 物理学 2024-05-24 Masato Shirasaki , Shiro Ikeda

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Runyi Yu , Chang Liu , Xiangyang Ji , Li Yuan , Jie Chen

Artificial intelligence is a key enabler for next-generation wireless communication and sensing. Yet, today's learning-based wireless techniques do not generalize well: most models are task-specific, environment-dependent, and limited to…

信号处理 · 电气工程与系统科学 2026-02-05 Vahid Yazdnian , Yasaman Ghasempour