中文
相关论文

相关论文: X-World: Controllable Ego-Centric Multi-Camera Wor…

200 篇论文

Recent advancements have exploited diffusion models for the synthesis of either LiDAR point clouds or camera image data in driving scenarios. Despite their success in modeling single-modality data marginal distribution, there is an…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Yichen Xie , Chenfeng Xu , Chensheng Peng , Shuqi Zhao , Nhat Ho , Alexander T. Pham , Mingyu Ding , Masayoshi Tomizuka , Wei Zhan

Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthesize only limited observations of a scene, leading to issues…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Yuheng Liu , Xin Lin , Xinke Li , Baihan Yang , Chen Wang , Kalyan Sunkavalli , Yannick Hold-Geoffroy , Hao Tan , Kai Zhang , Xiaohui Xie , Zifan Shi , Yiwei Hu

Training robot policies within a learned world model is trending due to the inefficiency of real-world interactions. The established image-based world models and policies have shown prior success, but lack robust geometric information that…

机器人学 · 计算机科学 2025-09-18 Guanxing Lu , Baoxiong Jia , Puhao Li , Yixin Chen , Ziwei Wang , Yansong Tang , Siyuan Huang

Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recent generative…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Kaicong Huang , Talha Azfar , Weisong Shi , Ruimin Ke

We provide a sober look at the application of Multimodal Large Language Models (MLLMs) in autonomous driving, challenging common assumptions about their ability to interpret dynamic driving scenarios. Despite advances in models like GPT-4o,…

机器人学 · 计算机科学 2024-10-29 Shiva Sreeram , Tsun-Hsuan Wang , Alaa Maalouf , Guy Rosman , Sertac Karaman , Daniela Rus

Learning natural, stable, and compositionally generalizable whole-body control policies for humanoid robots performing simultaneous locomotion and manipulation (loco-manipulation) remains a fundamental challenge in robotics. Existing…

机器人学 · 计算机科学 2026-03-10 Yutong Shen , Hangxu Liu , Penghui Liu , Jiashuo Luo , Yongkang Zhang , Rex Morvley , Chen Jiang , Jianwei Zhang , Lei Zhang

Modern computer vision algorithms typically require expensive data acquisition and accurate manual labeling. In this work, we instead leverage the recent progress in computer graphics to generate fully labeled, dynamic, and photo-realistic…

计算机视觉与模式识别 · 计算机科学 2016-05-23 Adrien Gaidon , Qiao Wang , Yohann Cabon , Eleonora Vig

Imitation learning has emerged as a promising approach towards building generalist robots. However, scaling imitation learning for large robot foundation models remains challenging due to its reliance on high-quality expert demonstrations.…

机器人学 · 计算机科学 2025-05-26 Chuning Zhu , Raymond Yu , Siyuan Feng , Benjamin Burchfiel , Paarth Shah , Abhishek Gupta

Emerging generative world models and vision-language-action (VLA) systems are rapidly reshaping automated driving by enabling scalable simulation, long-horizon forecasting, and capability-rich decision making. Across these directions,…

机器人学 · 计算机科学 2026-03-11 Rongxiang Zeng , Yongqi Dong

World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many generated environments require multi-agent interaction:…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Fangfu Liu , Kai He , Tianchang Shen , Tianshi Cao , Sanja Fidler , Yueqi Duan , Jun Gao , Igor Gilitschenski , Zian Wang , Xuanchi Ren

Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a single form of…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Boming Tan , Xiangdong Zhang , Ning Liao , Yuqing Zhang , Shaofeng Zhang , Xue Yang , Qi Fan , Yanyong Zhang

The development of software components for autonomous driving functions should always include an extensive and rigorous evaluation. Since real-world testing is expensive and safety-critical -- especially when facing dynamic racing scenarios…

机器人学 · 计算机科学 2020-06-18 Tim Stahl , Johannes Betz

Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and generating corresponding commands. However, current agents…

Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal consistency and physical constraints, leading to failures when mapped to executable actions.…

机器学习 · 计算机科学 2026-03-17 Christos Ziakas , Amir Bar , Alessandra Russo

Global registration of multi-view robot data is a challenging task. Appearance-based global localization approaches often fail under drastic view-point changes, as representations have limited view-point invariance. This work is based on…

机器人学 · 计算机科学 2018-02-19 Abel Gawel , Carlo Del Don , Roland Siegwart , Juan Nieto , Cesar Cadena

Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-distribution ones. However, such cases are underrepresented in real-world corpus collected by…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Haochen Tian , Tianyu Li , Haochen Liu , Jiazhi Yang , Yihang Qiu , Guang Li , Junli Wang , Yinfeng Gao , Zhang Zhang , Liang Wang , Hangjun Ye , Tieniu Tan , Long Chen , Hongyang Li

Camera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Jianhong Bai , Menghan Xia , Xiao Fu , Xintao Wang , Lianrui Mu , Jinwen Cao , Zuozhu Liu , Haoji Hu , Xiang Bai , Pengfei Wan , Di Zhang

Collaborative driving systems leverage vehicle-to-everything (V2X) communication for multi-agent collaborative perception to enhance driving safety, yet they remain constrained by scarce annotated real-world V2X driving datasets and limited…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Yihang Tao , Yu Guo , Senkang Hu , Yanan Ma , Zihan Fang , Sam Kwong , Yuguang Fang

With the rise and development of deep learning over the past decade, there has been a steady momentum of innovation and breakthroughs that convincingly push the state-of-the-art of cross-modal analytics between vision and language in…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Yehao Li , Yingwei Pan , Jingwen Chen , Ting Yao , Tao Mei

We propose Infinite-World, a robust interactive world model capable of maintaining coherent visual memory over 1000+ frames in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Ruiqi Wu , Xuanhua He , Meng Cheng , Tianyu Yang , Yong Zhang , Zhuoliang Kang , Xunliang Cai , Xiaoming Wei , Chunle Guo , Chongyi Li , Ming-Ming Cheng
‹ 上一页 1 8 9 10 下一页 ›