English
Related papers

Related papers: SceneMI: Motion In-betweening for Modeling Human-S…

200 papers

Bimanual manipulation is a longstanding challenge in robotics due to the large number of degrees of freedom and the strict spatial and temporal synchronization required to generate meaningful behavior. Humans learn bimanual manipulation…

Robotics · Computer Science 2024-05-07 Arpit Bahety , Priyanka Mandikal , Ben Abbatematteo , Roberto Martín-Martín

Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent understanding of a visual scene. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Ritik Raina , Abe Leite , Alexandros Graikos , Seoyoung Ahn , Dimitris Samaras , Gregory J. Zelinsky

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Korrawe Karunratanakul , Konpat Preechakul , Supasorn Suwajanakorn , Siyu Tang

Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise models excel in 2D tasks and static scenarios, their ability to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Ziang Zhang , Zehan Wang , Guanghao Zhang , Weilong Dai , Yan Xia , Ziang Yan , Minjie Hong , Zhou Zhao

We introduce an approach to model surface properties governing bounces in everyday scenes. Our model learns end-to-end, starting from sensor inputs, to predict post-bounce trajectories and infer two underlying physical properties that…

Computer Vision and Pattern Recognition · Computer Science 2019-04-16 Senthil Purushwalkam , Abhinav Gupta , Danny M. Kaufman , Bryan Russell

We propose a two-stage framework for motion in-betweening that combines diffusion-based motion generation with physics-based character adaptation. In Stage 1, a character-agnostic diffusion model synthesizes transitions from sparse…

Graphics · Computer Science 2025-04-15 Jia Qin

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limits the deployment of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Manjunath Prasad Holenarasipura Rajiv , B. M. Vidyavathi

We present SceneNAT, a single-stage masked non-autoregressive Transformer that synthesizes complete 3D indoor scenes from natural language instructions through only a few parallel decoding passes, offering improved performance and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jeongjun Choi , Yeonsoo Park , H. Jin Kim

Scene-understanding is an important topic in the area of Computer Vision, and illustrates computational challenges with applications to a wide range of domains including remote sensing, surveillance, smart agriculture, robotics, autonomous…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Zachary A Daniels , Dimitris Metaxas

While diffusion models and large-scale motion datasets have advanced text-driven human motion synthesis, extending these advances to 4D human-object interaction (HOI) remains challenging, mainly due to the limited availability of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Shujia Li , Haiyu Zhang , Xinyuan Chen , Yaohui Wang , Yutong Ban

Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in privacy, safety, energy efficiency, and scalability. We explore an alternative: 4D perception…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Hao-Yu Hsu , Tianhang Cheng , Jing Wen , Alexander G. Schwing , Shenlong Wang

In recent years, there has been a significant amount of research on algorithms and control methods for distributed collaborative robots. However, the emergence of collective behavior in a swarm is still difficult to predict and control.…

Robotics · Computer Science 2024-08-21 Pengming Zhu , Zhiwen Zeng , Weijia Yao , Wei Dai , Huimin Lu , Zongtan Zhou

Styled motion in-betweening is crucial for computer animation and gaming. However, existing methods typically encode motion styles by modeling whole-body motions, often overlooking the representation of individual body parts. This…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Minyue Dai , Ke Fan , Bin Ji , Haoran Xu , Haoyu Zhao , Junting Dong , Jingbo Wang , Bo Dai

We investigate the problem of identifying objects that have been added, removed, or moved between a pair of captures (images or videos) of the same scene at different times. Accurately identifying verifiable changes is extremely challenging…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Yuqun Wu , Chih-hao Lin , Henry Che , Aditi Tiwari , Chuhang Zou , Shenlong Wang , Derek Hoiem

Dynamic Scene Graphs (DSGs) provide a structured representation of hierarchical, interconnected environments, but current approaches struggle to capture stochastic dynamics, partial observability, and multi-agent activity. These aspects are…

Robotics · Computer Science 2025-10-13 Lars Ohnemus , Nils Hantke , Max Weißer , Kai Furmans

Quantifying the gap between synthetic and real-world imagery is essential for improving both transformer-based models - that rely on large volumes of data - and datasets, especially in underexplored domains like aerial scene understanding…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Alina Marcu

We propose a new 3D holistic++ scene understanding problem, which jointly tackles two tasks from a single-view image: (i) holistic scene parsing and reconstruction---3D estimations of object bounding boxes, camera pose, and room layout, and…

Computer Vision and Pattern Recognition · Computer Science 2019-09-05 Yixin Chen , Siyuan Huang , Tao Yuan , Siyuan Qi , Yixin Zhu , Song-Chun Zhu

Diffusion models generate images with an unprecedented level of quality, but how can we freely rearrange image layouts? Recent works generate controllable scenes via learning spatially disentangled latent codes, but these methods do not…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Jiawei Ren , Mengmeng Xu , Jui-Chieh Wu , Ziwei Liu , Tao Xiang , Antoine Toisoul

We present "Humans and Structure from Motion" (HSfM), a method for jointly reconstructing multiple human meshes, scene point clouds, and camera parameters in a metric world coordinate system from a sparse set of uncalibrated multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Lea Müller , Hongsuk Choi , Anthony Zhang , Brent Yi , Jitendra Malik , Angjoo Kanazawa

High-dynamic scene optical flow is a challenging task, which suffers spatial blur and temporal discontinuous motion due to large displacement in frame imaging, thus deteriorating the spatiotemporal feature of optical flow. Typically,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Hanyu Zhou , Haonan Wang , Haoyue Liu , Yuxing Duan , Yi Chang , Luxin Yan
‹ Prev 1 8 9 10 Next ›