English
Related papers

Related papers: Dense Motion Captioning

200 papers

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Shuolin Xu , Siming Zheng , Ziyi Wang , HC Yu , Jinwei Chen , Huaqi Zhang , Daquan Zhou , Tong-Yee Lee , Bo Li , Peng-Tao Jiang

Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Gong Jingyu , Tong Kunkun , Chen Zhuoran , Yuan Chuanhan , Chen Mingang , Zhang Zhizhong , Tan Xin , Xie Yuan

Learning specific hands-on skills such as cooking, car maintenance, and home repairs increasingly happens via instructional videos. The user experience with such videos is known to be improved by meta-information such as time-stamped…

Computer Vision and Pattern Recognition · Computer Science 2020-11-25 Gabriel Huang , Bo Pang , Zhenhai Zhu , Clara Rivera , Radu Soricut

We introduce Cap3D, an automatic approach for generating descriptive text for 3D objects. This approach utilizes pretrained models from image captioning, image-text alignment, and LLM to consolidate captions from multiple views of a 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Tiange Luo , Chris Rockwell , Honglak Lee , Justin Johnson

Dense video captioning is a challenging video understanding task which aims to simultaneously segment the video into a sequence of meaningful consecutive events and to generate detailed captions to accurately describe each event. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 AJ Piergiovanni , Ganesh Satish Mallya , Dahun Kim , Anelia Angelova

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yiheng Li , Zhuo Li , Ruibing Hou , Yingjie Chen , Hong Chang , Hao Liu , Shiguang Shan

Semantic segmentation of motion capture sequences plays a key part in many data-driven motion synthesis frameworks. It is a preprocessing step in which long recordings of motion capture sequences are partitioned into smaller segments.…

Computer Vision and Pattern Recognition · Computer Science 2018-07-17 Noshaba Cheema , Somayeh Hosseini , Janis Sprenger , Erik Herrmann , Han Du , Klaus Fischer , Philipp Slusallek

Although significant progress has been achieved on monocular maker-less human motion capture in recent years, it is still hard for state-of-the-art methods to obtain satisfactory results in occlusion scenarios. There are two main reasons:…

Computer Vision and Pattern Recognition · Computer Science 2022-07-13 Buzhen Huang , Yuan Shu , Jingyi Ju , Yangang Wang

3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Shanlin Sun , Gabriel De Araujo , Jiaqi Xu , Shenghan Zhou , Hanwen Zhang , Ziheng Huang , Chenyu You , Xiaohui Xie

Existing video captioning approaches typically require to first sample video frames from a decoded video and then conduct a subsequent process (e.g., feature extraction and/or captioning model learning). In this pipeline, manual frame…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yaojie Shen , Xin Gu , Kai Xu , Heng Fan , Longyin Wen , Libo Zhang

Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub-actions or incoherent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Miaowei Wang , Qingxuan Yan , Zhi Cao , Yayuan Li , Oisin Mac Aodha , Jason J. Corso , Amir Vaxman

Current approaches for 3D human motion synthesis generate high quality animations of digital humans performing a wide variety of actions and gestures. However, a notable technological gap exists in addressing the complex dynamics of multi…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Anindita Ghosh , Rishabh Dabral , Vladislav Golyanik , Christian Theobalt , Philipp Slusallek

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations and relations for visual entities.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Xiangtai Li , Tao Zhang , Yanwei Li , Haobo Yuan , Shihao Chen , Yikang Zhou , Jiahao Meng , Yueyi Sun , Shilin Xu , Lu Qi , Tianheng Cheng , Yi Lin , Zilong Huang , Wenhao Huang , Jiashi Feng , Guang Shi

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal…

We introduce the task of dense captioning in 3D scans from commodity RGB-D sensors. As input, we assume a point cloud of a 3D scene; the expected output is the bounding boxes along with the descriptions for the underlying objects. To…

Computer Vision and Pattern Recognition · Computer Science 2020-12-07 Dave Zhenyu Chen , Ali Gholami , Matthias Nießner , Angel X. Chang

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

3D dense captioning, as an emerging vision-language task, aims to identify and locate each object from a set of point clouds and generate a distinctive natural language sentence for describing each located object. However, the existing…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Yufeng Zhong , Long Xu , Jiebo Luo , Lin Ma

Existing video captioning methods merely provide shallow or simplistic representations of object behaviors, resulting in superficial and ambiguous descriptions. However, object behavior is dynamic and complex. To comprehensively capture the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-20 Caihua Liu , Xu Li , Wenjing Xue , Wei Tang , Xia Feng

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent…

Graphics · Computer Science 2026-05-20 Yi-Yang Zhang , Tengjiao Sun , Pengcheng Fang , Deng-Bao Wang , Xiaohao Cai , Min-Ling Zhang , Hansung Kim

While image captioning provides isolated descriptions for individual images, and video captioning offers one single narrative for an entire video clip, our work explores an important middle ground: progress-aware video captioning at the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Zihui Xue , Joungbin An , Xitong Yang , Kristen Grauman
‹ Prev 1 3 4 5 6 7 10 Next ›