中文
相关论文

相关论文: Bi-modal Prediction and Transformation Coding for …

200 篇论文

This paper introduces a Multi-modal Diffusion model for Motion Prediction (MDMP) that integrates and synchronizes skeletal data and textual descriptions of actions to generate refined long-term motion predictions with quantifiable…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Leo Bringer , Joey Wilson , Kira Barton , Maani Ghaffari

Neural image compression (NIC) has outperformed traditional image codecs in rate-distortion (R-D) performance. However, it usually requires a dedicated encoder-decoder pair for each point on R-D curve, which greatly hinders its practical…

图像与视频处理 · 电气工程与系统科学 2022-09-21 Chenjian Gao , Tongda Xu , Dailan He , Hongwei Qin , Yan Wang

This paper investigates the problem of efficient computation of physically consistent multi-contact behaviors. Recent work showed that under mild assumptions, the problem could be decomposed into simpler kinematic and centroidal dynamic…

机器人学 · 计算机科学 2021-02-19 Brahayam Ponton , Majid Khadiv , Avadesh Meduri , Ludovic Righetti

Safe and efficient robot operation in complex human environments can benefit from good models of site-specific motion patterns. Maps of Dynamics (MoDs) provide such models by encoding statistical motion patterns in a map, but existing…

机器人学 · 计算机科学 2026-02-03 Yufei Zhu , Shih-Min Yang , Andrey Rudenko , Tomasz P. Kucner , Achim J. Lilienthal , Martin Magnusson

Almost all digital videos are coded into compact representations before being transmitted. Such compact representations need to be decoded back to pixels before being displayed to humans and - as usual - before being enhanced/analyzed by…

图像与视频处理 · 电气工程与系统科学 2023-11-03 Xihua Sheng , Li Li , Dong Liu , Houqiang Li

The aim of this paper is to perform a deeper geometric analysis of problems appearing in dynamics of affinely rigid bodies. First of all we present a geometric interpretation of the polar and two-polar decomposition of affine motion. Later…

数学物理 · 物理学 2016-02-18 Jan Jerzy Sławianowski , Barbara Gołubowska , Vasyl Kovalchuk

Multiple Description Coding (MDC) is a promising error-resilient source coding method that is particularly suitable for dynamic networks with multiple (yet noisy and unreliable) paths. However, conventional MDC video codecs suffer from…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Xinyue Hu , Wei Ye , Jiaxiang Tang , Eman Ramadan , Zhi-Li Zhang

Predicting and executing a sequence of actions without intermediate replanning, known as action chunking, is increasingly used in robot learning from human demonstrations. Yet, its effects on the learned policy remain inconsistent: some…

机器人学 · 计算机科学 2025-04-28 Yuejiang Liu , Jubayer Ibn Hamid , Annie Xie , Yoonho Lee , Maximilian Du , Chelsea Finn

This paper presents a novel learning-based clothing deformation method to generate rich and reasonable detailed deformations for garments worn by bodies of various shapes in various animations. In contrast to existing learning-based…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Tianxing Li , Rui Shi , Takashi Kanai

In physics-based cloth animation, rich folds and detailed wrinkles are achieved at the cost of expensive computational resources and huge labor tuning. Data-driven techniques make efforts to reduce the computation significantly by a…

计算机视觉与模式识别 · 计算机科学 2021-02-24 Lan Chen , Lin Gao , Jie Yang , Shibiao Xu , Juntao Ye , Xiaopeng Zhang , Yu-Kun Lai

Real-world data contains a vast amount of multimodal information, among which vision and language are the two most representative modalities. Moreover, increasingly heavier models, \textit{e}.\textit{g}., Transformers, have attracted the…

计算机视觉与模式识别 · 计算机科学 2023-07-03 Dachuan Shi , Chaofan Tao , Ying Jin , Zhendong Yang , Chun Yuan , Jiaqi Wang

We present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Qihang Fang , Chengcheng Tang , Bugra Tekin , Shugao Ma , Yanchao Yang

Inspired by recent work on compression with and for young humans, the success of transform-based approaches to information processing, and the rise of powerful language-based AI, we propose \emph{textual transform coding}. It shares some of…

信息论 · 计算机科学 2023-05-04 Tsachy Weissman

Robust reinforcement learning agents using high-dimensional observations must be able to identify relevant state features amidst many exogeneous distractors. A representation that captures controllability identifies these state elements by…

机器学习 · 计算机科学 2024-06-25 Max Rudolph , Caleb Chuck , Kevin Black , Misha Lvovsky , Scott Niekum , Amy Zhang

Imitation learning from human motion capture (MoCap) data provides a promising way to train humanoid robots. However, due to differences in morphology, such as varying degrees of joint freedom and force limits, exact replication of human…

机器人学 · 计算机科学 2024-10-04 Wenshuai Zhao , Yi Zhao , Joni Pajarinen , Michael Muehlebach

Rigid-bodied robots often lack compliance needed to adapt to unstructured environments, while fully soft robots, though highly adaptable, struggle with scalability and load capacity. In nature, musculoskeletal systems balance strength and…

计算工程、金融与科学 · 计算机科学 2026-05-29 Hiroki Kobayashi , Yuki Takaha , Changyoung Yuhn , Yuki Sato , Sunao Tomita , Atsushi Kawamoto , Tsuyoshi Nomura

Mode-based model-reduction is used to reduce the degrees of freedom of high dimensional systems, often by describing the system state by a linear combination of spatial modes. Transport dominated phenomena, ubiquitous in technical and…

数值分析 · 数学 2020-02-28 Julius Reiss

We propose novel motion representations for animating articulated objects consisting of distinct parts. In a completely unsupervised manner, our method identifies object parts, tracks them in a driving video, and infers their motions by…

计算机视觉与模式识别 · 计算机科学 2021-04-26 Aliaksandr Siarohin , Oliver J. Woodford , Jian Ren , Menglei Chai , Sergey Tulyakov

Recently, learning based video compression methods attract increasing attention. However, the previous works suffer from error propagation due to the accumulation of reconstructed error in inter predictive coding. Meanwhile, the previous…

图像与视频处理 · 电气工程与系统科学 2020-03-26 Guo Lu , Chunlei Cai , Xiaoyun Zhang , Li Chen , Wanli Ouyang , Dong Xu , Zhiyong Gao

In recent years, large visual language models (LVLMs) have shown impressive performance and promising generalization capability in multi-modal tasks, thus replacing humans as receivers of visual information in various application scenarios.…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Binzhe Li , Shurun Wang , Shiqi Wang , Yan Ye