中文
相关论文

相关论文: FlowDreamer: A RGB-D World Model with Flow-based M…

200 篇论文

3D scene representation for robot manipulation should capture three key object properties: permanency -- objects that become occluded over time continue to exist; amodal completeness -- objects have 3D occupancy, even if only partial…

机器人学 · 计算机科学 2020-12-11 Zhenjia Xu , Zhanpeng He , Jiajun Wu , Shuran Song

Computational fluid dynamics (CFD) provides high-fidelity simulations of fluid flows but remains computationally expensive for many-query applications. In recent years deep learning (DL) has been used to construct data-driven fluid-dynamic…

机器学习 · 计算机科学 2026-04-13 David Ramos , Lucas Lacasa , Fermín Gutiérrez , Eusebio Valero , Gonzalo Rubio

Learned world models hold significant potential for robotic manipulation, as they can serve as simulator for real-world interactions. While extensive progress has been made in 2D video-based world models, these approaches often lack…

机器人学 · 计算机科学 2025-10-13 Chuanrui Zhang , Zhengxian Wu , Guanxing Lu , Yansong Tang , Ziwei Wang

We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models.…

Transparent objects are widely used in our daily lives, making it important to teach robots to interact with them. However, it's not easy because the reflective and refractive effects can make depth cameras fail to give accurate geometry…

计算机视觉与模式识别 · 计算机科学 2024-02-20 Tutian Tang , Jiyu Liu , Jieyi Zhang , Haoyuan Fu , Wenqiang Xu , Cewu Lu

Estimating the 3D motion of points in a scene, known as scene flow, is a core problem in computer vision. Traditional learning-based methods designed to learn end-to-end 3D flow often suffer from poor generalization. Here we present a…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Yair Kittenplon , Yonina C. Eldar , Dan Raviv

Though machine learning has achieved notable success in modeling sequential and spatial data for speech recognition and in computer vision, applications to remote sensing and climate science problems are seldom considered. In this paper, we…

机器学习 · 计算机科学 2019-10-22 Yimeng Min , S. Karthik Mukkavilli , Yoshua Bengio

Constructing simulation scenes that are both visually and physically realistic is a problem of practical interest in domains ranging from robotics to computer vision. This problem has become even more relevant as researchers wielding large…

机器人学 · 计算机科学 2024-06-03 Zoey Chen , Aaron Walsman , Marius Memmel , Kaichun Mo , Alex Fang , Karthikeya Vemuri , Alan Wu , Dieter Fox , Abhishek Gupta

We present High-Density Visual Particle Dynamics (HD-VPD), a learned world model that can emulate the physical dynamics of real scenes by processing massive latent point clouds containing 100K+ particles. To enable efficiency at this scale,…

机器学习 · 计算机科学 2024-07-01 William F. Whitney , Jacob Varley , Deepali Jain , Krzysztof Choromanski , Sumeet Singh , Vikas Sindhwani

Predicting pedestrian movements remains a complex and persistent challenge in robot navigation research. We must evaluate several factors to achieve accurate predictions, such as pedestrian interactions, the environment, crowd density, and…

机器人学 · 计算机科学 2023-09-20 Rashmi Bhaskara , Hrishikesh Viswanath , Aniket Bera

Robust perception and dynamics modeling are fundamental to real-world robotic policy learning. Recent methods employ video diffusion models (VDMs) to enhance robotic policies, improving their understanding and modeling of the physical…

机器人学 · 计算机科学 2026-03-25 Yueru Jia , Jiaming Liu , Shengbang Liu , Rui Zhou , Wanhe Yu , Yuyang Yan , Xiaowei Chi , Yandong Guo , Boxin Shi , Shanghang Zhang

We present an approach to modeling an image-space prior on scene motion. Our prior is learned from a collection of motion trajectories extracted from real video sequences depicting natural, oscillatory dynamics such as trees, flowers,…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Zhengqi Li , Richard Tucker , Noah Snavely , Aleksander Holynski

Current feed-forward 3D/4D reconstruction systems rely on dense geometry and pose supervision -- expensive to obtain at scale and particularly scarce for dynamic real-world scenes. We present Flow3r, a framework that augments visual…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Zhongxiao Cong , Qitao Zhao , Minsik Jeon , Shubham Tulsiani

Optical flow is a crucial component of the feature space for early visual processing of dynamic scenes especially in new applications such as self-driving vehicles, drones and autonomous robots. The dynamic vision sensors are well suited…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Himanshu Akolkar , SioHoi Ieng , Ryad Benosman

Dual-arm cooperative manipulation holds great promise for tackling complex real-world tasks that demand seamless coordination and adaptive dynamics. Despite substantial progress in learning-based motion planning, most approaches struggle to…

机器人学 · 计算机科学 2025-11-24 Jiaming Chen , Yiyu Jiang , Aoshen Huang , Yang Li , Wei Pan

Imitation learning is employed to learn sensorimotor coordination for steering angle prediction in an end-to-end fashion requires expert demonstrations. These expert demonstrations are paired with environmental perception and vehicle…

计算机视觉与模式识别 · 计算机科学 2022-02-14 Farzeen Munir , Shoaib Azam , Byung-Geun Lee , Moongu Jeon

World models that forecast scene evolution by generating future video frames devote the bulk of their capacity to photometric details, yet the resulting predictions often remain geometrically inconsistent. We present VGGT-World, a geometry…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Xiangyu Sun , Shijie Wang , Fengyi Zhang , Lin Liu , Caiyan Jia , Ziying Song , Zi Huang , Yadan Luo

Action recognition is a key problem in computer vision that labels videos with a set of predefined actions. Capturing both, semantic content and motion, along the video frames is key to achieve high accuracy performance on this task. Most…

计算机视觉与模式识别 · 计算机科学 2019-10-23 Xia Huang , Hossein Mousavi , Gemma Roig

Following its success in natural language processing and computer vision, foundation models that are pre-trained on large-scale multi-task datasets have also shown great potential in robotics. However, most existing robot foundation models…

机器人学 · 计算机科学 2025-03-13 Rujia Yang , Geng Chen , Chuan Wen , Yang Gao

Existing RGB-based imitation learning approaches typically employ traditional vision encoders such as ResNet or ViT, which lack explicit 3D reasoning capabilities. Recent geometry-grounded vision models, such as VGGT~\cite{wang2025vggt},…

机器人学 · 计算机科学 2025-09-22 An Dinh Vuong , Minh Nhat Vu , Ian Reid
‹ 上一页 1 8 9 10 下一页 ›