English
Related papers

Related papers: MotionAnymesh: Physics-Grounded Articulation for S…

200 papers

Simulation has become a key tool for training and evaluating home robots at scale, yet existing environments fail to capture the diversity and physical complexity of real indoor spaces. Current scene synthesis methods produce sparsely…

Robotics · Computer Science 2026-02-12 Nicholas Pfaff , Thomas Cohn , Sergey Zakharov , Rick Cory , Russ Tedrake

Creating accurate, physical simulations directly from real-world robot motion holds great value for safe, scalable, and affordable robot learning, yet remains exceptionally challenging. Real robot data suffers from occlusions, noisy camera…

Robotics · Computer Science 2025-06-10 Ben Moran , Mauro Comi , Arunkumar Byravan , Steven Bohez , Tom Erez , Zhibin Li , Leonard Hasenclever

Reconstructing articulated objects is essential for building digital twins of interactive environments. However, prior methods typically decouple geometry and motion by first reconstructing object shape in distinct states and then…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Licheng Shen , Saining Zhang , Honghan Li , Peilin Yang , Zihao Huang , Zongzheng Zhang , Hao Zhao

Reconstructing physically valid 3D scenes from single-view observations is a prerequisite for bridging the gap between visual perception and robotic control. However, in scenarios requiring precise contact reasoning, such as robotic…

Robotics · Computer Science 2026-05-19 Tianyi Xiang , Jiahang Cao , Sikai Guo , Guoyang Zhao , Andrew F. Luo , Jun Ma

We propose a Vision-Language Simulation Model (VLSM) that unifies visual and textual understanding to synthesize executable FlexScript from layout sketches and natural-language prompts, enabling cross-modal reasoning for industrial…

Artificial Intelligence · Computer Science 2026-01-14 YuChe Hsu , AnJui Wang , TsaiChing Ni , YuanFu Yang

While humans effortlessly discern intrinsic dynamics and adapt to new scenarios, modern AI systems often struggle. Current methods for visual grounding of dynamics either use pure neural-network-based simulators (black box), which may…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Junyi Cao , Shanyan Guan , Yanhao Ge , Wei Li , Xiaokang Yang , Chao Ma

Vision-language-action models have advanced robotic manipulation but remain constrained by reliance on the large, teleoperation-collected datasets dominated by the static, tabletop scenes. We propose a simulation-first framework to verify…

Robotics · Computer Science 2026-02-06 Wenbo Wang , Fangyun Wei , QiXiu Li , Xi Chen , Yaobo Liang , Chang Xu , Jiaolong Yang , Baining Guo

When embodied AI is expanding from traditional object detection and recognition to more advanced tasks of robot manipulation and actuation planning, visual spatial reasoning from the video inputs is necessary to perceive the spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Haoming Wang , Qiyao Xue , Weichen Liu , Wei Gao

We introduce Scan2Mesh, a novel data-driven generative approach which transforms an unstructured and potentially incomplete range scan into a structured 3D mesh representation. The main contribution of this work is a generative neural…

Computer Vision and Pattern Recognition · Computer Science 2019-04-03 Angela Dai , Matthias Nießner

Reconstructing simulation-ready deformable objects is important for vision, graphics, and robotics. Existing physics-driven methods can recover physical digital twins from videos, but they suffer from two fundamental limitations: they…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yang Yang , Yiyan Wang , Zheming Liu , Naoya Iwamoto

State-of-the-art text-to-motion generation models rely on the kinematic-aware, local-relative motion representation popularized by HumanML3D, which encodes motion relative to the pelvis and to the previous frame with built-in redundancy.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zichong Meng , Zeyu Han , Xiaogang Peng , Yiming Xie , Huaizu Jiang

Creating functional Digital Twins, simulatable 3D replicas of the real world, is a central challenge in computer vision. Current methods like NeRF produce visually rich but functionally incomplete twins. The key barrier is the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zhe Chen , Peilin Zheng , Wenshuo Chen , Xiucheng Wang , Yutao Yue , Nan Cheng

Recent advances in 3D content generation have amplified demand for dynamic models that are both visually realistic and physically consistent. However, state-of-the-art video diffusion models frequently produce implausible results such as…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Siwei Meng , Yawei Luo , Ping Liu

We present an inverse image-formation module that can enhance the robustness of existing visual SLAM pipelines for casually captured scenarios. Casual video captures often suffer from motion blur and varying appearances, which degrade the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Gwangtak Bae , Changwoon Choi , Hyeongjun Heo , Sang Min Kim , Young Min Kim

Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their…

Advances in 3D generative AI have enabled the creation of physical objects from text prompts, but challenges remain in creating objects involving multiple component types. We present a pipeline that integrates 3D generative AI with…

Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing approaches typically mitigate this limitation either…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Jiahua Chen , Qihong Tang , Weinong Wang , Qi Fan

Vision language models (VLMs) exhibit vast knowledge of the physical world, including intuition of physical and spatial properties, affordances, and motion. With fine-tuning, VLMs can also natively produce robot trajectories. We demonstrate…

Robotics · Computer Science 2025-05-16 William Xie , Max Conway , Yutong Zhang , Nikolaus Correll

Fast, accurate, and generalizable simulations are a key enabler of modern advances in robot design and control. However, existing simulation frameworks in robotics either model rigid environments and mechanisms only, or if they include…

Robotics · Computer Science 2024-02-21 Andrew Choi , Ran Jing , Andrew Sabelhaus , Mohammad Khalid Jawed

The fusion of vision and language has brought about a transformative shift in computer vision through the emergence of Vision-Language Models (VLMs). However, the resource-intensive nature of existing VLMs poses a significant challenge. We…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Jordan Shipard , Arnold Wiliem , Kien Nguyen Thanh , Wei Xiang , Clinton Fookes