中文
相关论文

相关论文: OmniWorld: A Multi-Domain and Multi-Modal Dataset …

200 篇论文

The field of generative AI has a transformative impact on various areas, including virtual reality, autonomous driving, the metaverse, gaming, and robotics. Among these applications, 3D object generation techniques are of utmost importance.…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Qinghong Sun , Yangguang Li , ZeXiang Liu , Xiaoshui Huang , Fenggang Liu , Xihui Liu , Wanli Ouyang , Jing Shao

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video generation with evolving…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Jiahao Wang , Luoxin Ye , TaiMing Lu , Junfei Xiao , Jiahan Zhang , Yuxiang Guo , Xijun Liu , Rama Chellappa , Cheng Peng , Alan Yuille , Jieneng Chen

Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Current fusion methods are typically evaluated on a small selection of public datasets, a limited…

机器学习 · 计算机科学 2026-05-07 Leyan Xue , Changqing Zhang , Kecheng Xue , Xiaohong Liu , Guangyu Wang , Zongbo Han

In the rapidly evolving landscape of autonomous driving, the capability to accurately predict future events and assess their implications is paramount for both safety and efficiency, critically aiding the decision-making process. World…

机器学习 · 计算机科学 2024-05-08 Yanchen Guan , Haicheng Liao , Zhenning Li , Jia Hu , Runze Yuan , Yunjian Li , Guohui Zhang , Chengzhong Xu

Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Pengze Zhang , Yanze Wu , Mengtian Li , Xu Bai , Songtao Zhao , Fulong Ye , Chong Mou , Xinghui Li , Zhuowei Chen , Qian He , Mingyuan Gao

Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically, existing 4D generative models directly embed macro scale…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Haonan Wang , Hanyu Zhou , Tao Gu , Luxin Yan

Recent advancements in unified multimodal understanding and visual generation (or multimodal generation) models have been hindered by their quadratic computational complexity and dependence on large-scale training data. We present…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

The emergence of Large Language Models (LLMs) has unified language generation tasks and revolutionized human-machine interaction. However, in the realm of image generation, a unified model capable of handling various tasks within a single…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Shitao Xiao , Yueze Wang , Junjie Zhou , Huaying Yuan , Xingrun Xing , Ruiran Yan , Chaofan Li , Shuting Wang , Tiejun Huang , Zheng Liu

Understanding the evolution of 3D scenes is important for effective autonomous driving. While conventional methods mode scene development with the motion of individual instances, world models emerge as a generative framework to describe the…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Lening Wang , Wenzhao Zheng , Yilong Ren , Han Jiang , Zhiyong Cui , Haiyang Yu , Jiwen Lu

4D human sensing and modeling are fundamental tasks in vision and graphics with numerous applications. With the advances of new sensors and algorithms, there is an increasing demand for more versatile datasets. In this work, we contribute…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Zhongang Cai , Daxuan Ren , Ailing Zeng , Zhengyu Lin , Tao Yu , Wenjia Wang , Xiangyu Fan , Yang Gao , Yifan Yu , Liang Pan , Fangzhou Hong , Mingyuan Zhang , Chen Change Loy , Lei Yang , Ziwei Liu

Generating multi-camera street-view videos is critical for augmenting autonomous driving datasets, addressing the urgent demand for extensive and varied data. Due to the limitations in diversity and challenges in handling lighting…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Jiachen Lu , Ze Huang , Zeyu Yang , Jiahui Zhang , Li Zhang

Multimodal Large Models (MLMs) are becoming a significant research focus, combining powerful large language models with multimodal learning to perform complex tasks across different data modalities. This review explores the latest…

机器学习 · 计算机科学 2024-07-02 Xinji Mai , Zeng Tao , Junxiong Lin , Haoran Wang , Yang Chang , Yanlan Kang , Yan Wang , Wenqiang Zhang

World modeling is a crucial task for enabling intelligent agents to effectively interact with humans and operate in dynamic environments. In this work, we propose MineWorld, a real-time interactive world model on Minecraft, an open-ended…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Junliang Guo , Yang Ye , Tianyu He , Haoyu Wu , Yushu Jiang , Tim Pearce , Jiang Bian

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Shihao Wang , Zhiding Yu , Xiaohui Jiang , Shiyi Lan , Min Shi , Nadine Chang , Jan Kautz , Ying Li , Jose M. Alvarez

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Shihao Wang , Zhiding Yu , Xiaohui Jiang , Shiyi Lan , Min Shi , Nadine Chang , Jan Kautz , Ying Li , Jose M. Alvarez

Accurate segmentation of longitudinal CT scans is important for monitoring tumor progression and evaluating treatment responses. However, existing 3D segmentation models solely focus on spatial information. To address this gap, we propose…

图像与视频处理 · 电气工程与系统科学 2025-04-25 Justin Namuk Kim , Yiqiao Liu , Rajath Soans , Keith Persson , Sarah Halek , Michal Tomaszewski , Jianda Yuan , Gregory Goldmacher , Antong Chen

The field of autonomous driving is experiencing a surge of interest in world models, which aim to predict potential future scenarios based on historical observations. In this paper, we introduce DFIT-OccWorld, an efficient 3D occupancy…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Haiming Zhang , Ying Xue , Xu Yan , Jiacheng Zhang , Weichao Qiu , Dongfeng Bai , Bingbing Liu , Shuguang Cui , Zhen Li

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Oskar Kristoffersen , Alba Reinders Sánchez , Morten Rieger Hannemose , Anders Bjorholm Dahl , Dim P. Papadopoulos

In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pinnacle in large language models (LLMs), while the computer…

计算机视觉与模式识别 · 计算机科学 2023-11-20 Chris Kelly , Luhui Hu , Cindy Yang , Yu Tian , Deshun Yang , Bang Yang , Zaoshan Huang , Zihao Li , Yuexian Zou

Multi-agent traffic simulation is central to developing and testing autonomous driving systems. Recent data-driven simulators have achieved promising results, but rely heavily on supervised learning from labeled trajectories or semantic…

机器人学 · 计算机科学 2026-04-01 Mozhgan Pourkeshavatz , Tianran Liu , Nicholas Rhinehart