English
Related papers

Related papers: Stitch-a-Demo: Video Demonstrations from Multistep…

200 papers

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Bizhu Wu , Jinheng Xie , Meidan Ding , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models are invariably…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Andrew Shin , Yusuke Mori , Kunitake Kaneko

Multi-Object Tracking (MOT) aims to associate multiple objects across video frames and is a challenging vision task due to inherent complexities in the tracking environment. Most existing approaches train and track within a single domain,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Run Luo , Zikai Song , Longze Chen , Yunshui Li , Min Yang , Wei Yang

We introduce the task of early mistake detection in video, where the goal is to determine whether a keystep in a procedural activity is performed correctly while observing as little of the streaming video as possible. To tackle this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Sagnik Majumder , Anish Nethi , Ziad Al-Halah , Kristen Grauman

We address the problem of extracting key steps from unlabeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and performance. We decompose the problem into two steps:…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Anshul Shah , Benjamin Lundell , Harpreet Sawhney , Rama Chellappa

Our goal in this paper is the adaptation of image-text models for long video retrieval. Recent works have demonstrated state-of-the-art performance in video retrieval by adopting CLIP, effectively hitchhiking on the image-text…

Computer Vision and Pattern Recognition · Computer Science 2022-05-18 Max Bain , Arsha Nagrani , Gül Varol , Andrew Zisserman

A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain…

Multimedia · Computer Science 2024-06-21 Yuchen Yang , Yingxuan Duan

Video is a powerful medium for communication and storytelling, yet reauthoring existing footage remains challenging. Even simple edits often demand expertise, time, and careful planning, constraining how creators envision and shape their…

Human-Computer Interaction · Computer Science 2026-04-07 Sitong Wang , Anh Truong , Lydia B. Chilton , Dingzeyu Li

We present SynthCLIP, a CLIP model trained on entirely synthetic text-image pairs. Leveraging recent text-to-image (TTI) networks and large language models (LLM), we generate synthetic datasets of images and corresponding captions at scale,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Hasan Abed Al Kader Hammoud , Hani Itani , Fabio Pizzati , Philip Torr , Adel Bibi , Bernard Ghanem

This paper explores the task of interactive image retrieval using natural language queries, where a user progressively provides input queries to refine a set of retrieval results. Moreover, our work explores this problem in the context of…

Computer Vision and Pattern Recognition · Computer Science 2019-11-12 Fuwen Tan , Paola Cascante-Bonilla , Xiaoxiao Guo , Hui Wu , Song Feng , Vicente Ordonez

Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique visual characteristics of particular subjects and ensure…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Shaobin Zhuang , Zhipeng Huang , Binxin Yang , Ying Zhang , Fangyikang Wang , Canmiao Fu , Chong Sun , Zheng-Jun Zha , Chen Li , Yali Wang

Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Hyelin Nam , Hyojun Go , Byeongjun Park , Byung-Hoon Kim , Hyungjin Chung

Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi-step video generation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Haodong Yan , Zhide Zhong , Jiaguan Zhu , Junjie He , Weilin Yuan , Wenxuan Song , Xin Gong , Yingjie Cai , Guanyi Zhao , Xu Yan , Bingbing Liu , Ying-Cong Chen , Haoang Li

Recently, one-stage trackers that use a joint model to predict both detections and appearance embeddings in one forward pass received much attention and achieved state-of-the-art results on the Multi-Object Tracking (MOT) benchmarks.…

Computer Vision and Pattern Recognition · Computer Science 2022-05-12 Shuzhi Yu , Guanhang Wu , Chunhui Gu , Mohammed E. Fathy

Video summarization helps turn long videos into clear, concise representations that are easier to review, document, and analyze, especially in high-stakes domains like surgical training. Prior work has progressed from using basic visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Shreya Rajpal , Michal Golovanevsky , Carsten Eickhoff

Previous work on visual storytelling mainly focused on exploring image sequence as evidence for storytelling and neglected textual evidence for guiding story generation. Motivated by human storytelling process which recalls stories for…

Computation and Language · Computer Science 2019-11-26 Tianyi Li , Sujian Li

In this work, following the intuition that adverbs describing scene-sequences are best identified by reasoning over high-level concepts of object-behavior, we propose the design of a new framework that reasons over object-behaviours…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Amrit Diggavi Seshadri , Alessandra Russo

Depth estimation and scene segmentation are two important tasks in intelligent transportation systems. A joint modeling of these two tasks will reduce the requirement for both the storage and training efforts. This work explores how the…

Machine Learning · Computer Science 2025-05-16 Tiancong Cheng , Ying Zhang , Yuxuan Liang , Roger Zimmermann , Zhiwen Yu , Bin Guo

Standard video and movie description tasks abstract away from person identities, thus failing to link identities across sentences. We propose a multi-sentence Identity-Aware Video Description task, which overcomes this limitation and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Jae Sung Park , Trevor Darrell , Anna Rohrbach

Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Honghui Yang , Di Huang , Wei Yin , Chunhua Shen , Haifeng Liu , Xiaofei He , Binbin Lin , Wanli Ouyang , Tong He