中文
相关论文

相关论文: Neural Implicit Event Generator for Motion Trackin…

200 篇论文

This work introduces a novel and adaptable architecture designed for real-time occupancy forecasting that outperforms existing state-of-the-art models on the Waymo Open Motion Dataset in Soft IOU. The proposed model uses recursive latent…

机器人学 · 计算机科学 2024-02-05 Bryce Ferenczi , Michael Burke , Tom Drummond

Motion capture using sparse inertial sensors has shown great promise due to its portability and lack of occlusion issues compared to camera-based tracking. Existing approaches typically assume that IMU sensors are tightly attached to the…

图形学 · 计算机科学 2025-08-14 Andela Ilic , Jiaxi Jiang , Paul Streli , Xintong Liu , Christian Holz

State-of-the-art frame interpolation methods generate intermediate frames by inferring object motions in the image from consecutive key-frames. In the absence of additional information, first-order approximations, i.e. optical flow, must be…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Stepan Tulyakov , Daniel Gehrig , Stamatios Georgoulis , Julius Erbach , Mathias Gehrig , Yuanyou Li , Davide Scaramuzza

The goal of imitation learning is to mimic expert behavior without access to an explicit reward signal. Expert demonstrations provided by humans, however, often show significant variability due to latent factors that are typically not…

机器学习 · 计算机科学 2017-11-16 Yunzhu Li , Jiaming Song , Stefano Ermon

We consider the problem of forecasting motion from a single image, i.e., predicting how objects in the world are likely to move, without the ability to observe other parameters such as the object velocities or the forces applied to them. We…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Gabrijel Boduljak , Laurynas Karazija , Iro Laina , Christian Rupprecht , Andrea Vedaldi

Camouflage poses challenges in distinguishing a static target, whereas any movement of the target can break this disguise. Existing video camouflaged object detection (VCOD) approaches take noisy motion estimation as input or model motion…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Xin Zhang , Tao Xiao , Gepeng Ji , Xuan Wu , Keren Fu , Qijun Zhao

Reinforcement learning (RL) for motion planning of multi-degree-of-freedom robots still suffers from low efficiency in terms of slow training speed and poor generalizability. In this paper, we propose a novel RL-based robot motion planning…

机器人学 · 计算机科学 2024-08-20 Zengjie Zhang , Jayden Hong , Amir Soufi Enayati , Homayoun Najjaran

In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state. A critical…

计算机视觉与模式识别 · 计算机科学 2023-08-21 An-Lan Wang , Kun-Yu Lin , Jia-Run Du , Jingke Meng , Wei-Shi Zheng

Multimodal semantic cues, such as textual descriptions, have shown strong potential in enhancing target perception for tracking. However, existing methods rely on static textual descriptions from large language models, which lack…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Yukuan Zhang , Jiarui Zhao , Shangqing Nie , Jin Kuang , Shengsheng Wang

Events across a timeline are a common data representation, seen in different temporal modalities. Individual atomic events can occur in a certain temporal ordering to compose higher level composite events. Examples of a composite event are…

机器学习 · 计算机科学 2022-02-14 Karan Samel , Zelin Zhao , Binghong Chen , Shuang Li , Dharmashankar Subramanian , Irfan Essa , Le Song

Event cameras are advantageous for tasks that require vision sensors with low-latency and sparse output responses. However, the development of deep network algorithms using event cameras has been slow because of the lack of large labelled…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Joachim Ott , Zuowen Wang , Shih-Chii Liu

Traditional RGB-based speech generation faces Temporal Granularity Mismatch since fixed camera exposure times inevitably blur the high-frequency articulatory transients essential for rendering emotional speech. To break this ceiling, we…

多媒体 · 计算机科学 2026-05-27 Jingping Fang , Lin Chen , Chenyang Xu , Tong Zhao , Weidong Cai , Xiaoming Chen

The tracking method based on the extreme learning machine (ELM) is efficient and effective. ELM randomly generates input weights and biases in the hidden layer, and then calculates and computes the output weights by reducing the iterative…

机器学习 · 计算机科学 2018-07-27 Jing Zhang , Huibing Wang , Yonggong Ren

The Global Event Processor (GEP) FPGA is an area-constrained, performance-critical element of the Large Hadron Collider's (LHC) ATLAS experiment. It needs to very quickly determine which small fraction of detected events should be retained…

Event-based cameras capture visual information as asynchronous streams of per-pixel brightness changes, generating sparse, temporally precise data. Compared to conventional frame-based sensors, they offer significant advantages in capturing…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Biswadeep Sen , Benoit R. Cottereau , Nicolas Cuperlier , Terence Sim

For compressive sensing of dynamic sparse signals, we develop an iterative pursuit algorithm. A dynamic sparse signal process is characterized by varying sparsity patterns over time/space. For such signals, the developed algorithm is able…

统计理论 · 数学 2012-10-15 Dave Zachariah , Saikat Chatterjee , Magnus Jansson

Recent advances in motion-aware large language models have shown remarkable promise for unifying motion understanding and generation tasks. However, these models typically treat understanding and generation separately, limiting the mutual…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Yuan-Ming Li , Qize Yang , Nan Lei , Shenghao Fu , Ling-An Zeng , Jian-Fang Hu , Xihan Wei , Wei-Shi Zheng

Accurately predicting how agents move in dynamic scenes is essential for safe autonomous driving. State-of-the-art motion forecasting models rely on datasets with manually annotated or post-processed trajectories. However, building these…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Yihong Xu , Yuan Yin , Éloi Zablocki , Tuan-Hung Vu , Alexandre Boulch , Matthieu Cord

We introduce Puppet-Master, an interactive video generator that captures the internal, part-level motion of objects, serving as a proxy for modeling object dynamics universally. Given an image of an object and a set of "drags" specifying…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Ruining Li , Chuanxia Zheng , Christian Rupprecht , Andrea Vedaldi

In this paper, we propose a new task of sub-event generation for an unseen process to evaluate the understanding of the coherence of sub-event actions and objects. To solve the problem, we design SubeventWriter, a sub-event sequence…

计算与语言 · 计算机科学 2022-10-20 Zhaowei Wang , Hongming Zhang , Tianqing Fang , Yangqiu Song , Ginny Y. Wong , Simon See