English
Related papers

Related papers: ExpertEdit: Learning Skill-Aware Motion Editing fr…

200 papers

Embodied control requires agents to leverage multi-modal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details needed for low-level perception and control, and language…

Machine Learning · Computer Science 2023-04-20 Yao Mu , Shunyu Yao , Mingyu Ding , Ping Luo , Chuang Gan

Diverse and extensive work has recently been conducted on text-conditioned human motion generation. However, progress in the reverse direction, motion captioning, has seen less comparable advancement. In this paper, we introduce a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Karim Radouane , Julien Lagarde , Sylvie Ranwez , Andon Tchechmedjiev

Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained spatial manipulations, motivating a dedicated assessment…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Yicheng Xiao , Wenhu Zhang , Lin Song , Yukang Chen , Wenbo Li , Nan Jiang , Tianhe Ren , Haokun Lin , Wei Huang , Haoyang Huang , Xiu Li , Nan Duan , Xiaojuan Qi

Language-driven action localization in videos is a challenging task that involves not only visual-linguistic matching but also action boundary prediction. Recent progress has been achieved through aligning language query to video segments,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-13 Shuo Yang , Xinxiao Wu

Data-driven approaches for edge detection have proven effective and achieve top results on modern benchmarks. However, all current data-driven edge detectors require manual supervision for training in the form of hand-labeled region…

Computer Vision and Pattern Recognition · Computer Science 2016-04-12 Yin Li , Manohar Paluri , James M. Rehg , Piotr Dollár

Prompt tuning and adapter tuning have shown great potential in transferring pre-trained vision-language models (VLMs) to various downstream tasks. In this work, we design a new type of tuning method, termed as regularized mask tuning, which…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Kecheng Zheng , Wei Wu , Ruili Feng , Kai Zhu , Jiawei Liu , Deli Zhao , Zheng-Jun Zha , Wei Chen , Yujun Shen

Masked autoencoders (MAEs) have established themselves as a powerful method for unsupervised pre-training for computer vision tasks. While vanilla MAEs put equal emphasis on reconstructing the individual parts of the image, we propose to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-26 Leon Sick , Dominik Engel , Pedro Hermosilla , Timo Ropinski

In the context of fitness coaching or for rehabilitation purposes, the motor actions of a human participant must be observed and analyzed for errors in order to provide effective feedback. This task is normally carried out by human coaches,…

Artificial Intelligence · Computer Science 2017-09-27 Felix Hülsmann , Stefan Kopp , Mario Botsch

Knowledge editing, which aims to update the knowledge encoded in language models, can be deceptive. Despite the fact that many existing knowledge editing algorithms achieve near-perfect performance on conventional metrics, the models edited…

Computation and Language · Computer Science 2025-05-20 Jiakuan Xie , Pengfei Cao , Yubo Chen , Kang Liu , Jun Zhao

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Inversion-based image editing is rapidly gaining momentum while suffering from significant computation overhead, hindering its application in real-time interactive scenarios. In this paper, we rethink that the redundancy in inversion-based…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Zexuan Yan , Yue Ma , Chang Zou , Wenteng Chen , Qifeng Chen , Linfeng Zhang

We present ExAct, a new video-language benchmark for expert-level understanding of skilled physical human activities. Our new benchmark contains 3521 expert-curated video question-answer pairs spanning 11 physical activities in 6 domains:…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Han Yi , Yulu Pan , Feihong He , Xinyu Liu , Benjamin Zhang , Oluwatumininu Oguntola , Gedas Bertasius

Existing video self-supervised learning methods mainly rely on trimmed videos for model training. However, trimmed datasets are manually annotated from untrimmed videos. In this sense, these methods are not really self-supervised. In this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-07 Dezhao Luo , Bo Fang , Yu Zhou , Yucan Zhou , Dayan Wu , Weiping Wang

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Katsuyuki Nakamura , Hiroki Ohashi , Mitsuhiro Okada

Diffusion models have recently advanced video editing, yet controllable editing remains challenging due to the need for precise manipulation of diverse object properties. Current methods require different control signal for diverse editing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yuqing Chen , Junjie Wang , Lin Liu , Ruihang Chu , Xiaopeng Zhang , Qi Tian , Yujiu Yang

Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural language processing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Hongxia Xie , Chu-Jun Peng , Yu-Wen Tseng , Hung-Jen Chen , Chan-Feng Hsu , Hong-Han Shuai , Wen-Huang Cheng

Estimating 3D human motion from an egocentric video sequence plays a critical role in human behavior understanding and has various applications in VR/AR. However, naively learning a mapping between egocentric videos and human motions is…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Jiaman Li , C. Karen Liu , Jiajun Wu

Single view depth estimation models can be trained from video footage using a self-supervised end-to-end approach with view synthesis as the supervisory signal. This is achieved with a framework that predicts depth and camera motion, with a…

Computer Vision and Pattern Recognition · Computer Science 2019-08-30 Maarten Schellevis

Self-supervised video correspondence learning depends on the ability to accurately associate pixels between video frames that correspond to the same visual object. However, achieving reliable pixel matching without supervision remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Zihan Zhou , Changrui Dai , Aibo Song , Xiaolin Fang

Previous studies have established that language models manifest stereotyped biases. Existing debiasing strategies, such as retraining a model with counterfactual data, representation projection, and prompting often fail to efficiently…

Computation and Language · Computer Science 2025-03-12 Xin Xu , Wei Xu , Ningyu Zhang , Julian McAuley
‹ Prev 1 8 9 10 Next ›