English
Related papers

Related papers: How to Correctly Make Mistakes: A Framework for Co…

200 papers

The key to video inpainting is to use correlation information from as many reference frames as possible. Existing flow-based propagation methods split the video synthesis process into multiple steps: flow completion -> pixel propagation ->…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Jaeyeon Kang , Seoung Wug Oh , Seon Joo Kim

With the surge in attention to Egocentric Hand-Object Interaction (Ego-HOI), large-scale datasets such as Ego4D and EPIC-KITCHENS have been proposed. However, most current research is built on resources derived from third-person video…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Yue Xu , Yong-Lu Li , Zhemin Huang , Michael Xu Liu , Cewu Lu , Yu-Wing Tai , Chi-Keung Tang

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior work is limited to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Zihui Xue , Kristen Grauman

While exocentric video synthesis has achieved great progress, egocentric video generation remains largely underexplored, which requires modeling first-person view content along with camera motion patterns induced by the wearer's body…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Jingqiao Xiu , Fangzhou Hong , Yicong Li , Mengze Li , Wentao Wang , Sirui Han , Liang Pan , Ziwei Liu

Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is the lack of…

Computer Vision and Pattern Recognition · Computer Science 2020-01-17 Antoine Miech , Ivan Laptev , Josef Sivic

Estimating 3D human motion from an egocentric video sequence plays a critical role in human behavior understanding and has various applications in VR/AR. However, naively learning a mapping between egocentric videos and human motions is…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Jiaman Li , C. Karen Liu , Jiajun Wu

Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Munan Ning , Bin Zhu , Yujia Xie , Bin Lin , Jiaxi Cui , Lu Yuan , Dongdong Chen , Li Yuan

Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation of continuous frames. While these models show promise for…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Yifan Li , Yukai Gu , Yingqian Min , Zikang Liu , Yifan Du , Kun Zhou , Min Yang , Wayne Xin Zhao , Minghui Qiu

Emotional Video Captioning (EVC) is an emerging task, which aims to describe factual content with the intrinsic emotions expressed in videos. Existing works perceive global emotional cues and then combine with video content to generate…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Weidong Chen , Cheng Ye , Zhendong Mao , Peipei Song , Xinyan Liu , Lei Zhang , Xiaojun Chang , Yongdong Zhang

ENIGMA-51 is a new egocentric dataset acquired in an industrial scenario by 19 subjects who followed instructions to complete the repair of electrical boards using industrial tools (e.g., electric screwdriver) and equipments (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Francesco Ragusa , Rosario Leonardi , Michele Mazzamuto , Claudia Bonanno , Rosario Scavo , Antonino Furnari , Giovanni Maria Farinella

Egocentric human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and scarce labeled data. We present EgoPoseFormer v2, a method…

Image scoring is a crucial task in numerous real-world applications. To trust a model's judgment, understanding its rationale is essential. This paper proposes a novel training method for Vision Language Models (VLMs) to generate not only…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Naoto Tanji , Toshihiko Yamasaki

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural language processing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Hongxia Xie , Chu-Jun Peng , Yu-Wen Tseng , Hung-Jen Chen , Chan-Feng Hsu , Hong-Han Shuai , Wen-Huang Cheng

Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent. To evaluate these video generative models, benchmarks such as VBench have…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Dian Zheng , Ziqi Huang , Hongbo Liu , Kai Zou , Yinan He , Fan Zhang , Lulu Gu , Yuanhan Zhang , Jingwen He , Wei-Shi Zheng , Yu Qiao , Ziwei Liu

Environment understanding in egocentric videos is an important step for applications like robotics, augmented reality and assistive technologies. These videos are characterized by dynamic interactions and a strong dependence on the wearer…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Lorenzo Mur-Labadia , Josechu Guerrero , Ruben Martinez-Cantin

Recent studies have shown that providing personalized explanations alongside recommendations increases trust and perceived quality. Furthermore, it gives users an opportunity to refine the recommendations by critiquing parts of the…

Information Retrieval · Computer Science 2021-07-09 Diego Antognini , Boi Faltings

We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus on factual accuracy. Unlike existing metrics that suffer from limited context handling, weak…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Shubhashis Roy Dipta , Tz-Ying Wu , Subarna Tripathi

Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing to capture the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Wei Chow , Jiachun Pan , Yongyuan Liang , Mingze Zhou , Xue Song , Liyu Jia , Saining Zhang , Siliang Tang , Juncheng Li , Fengda Zhang , Weijia Wu , Hanwang Zhang , Tat-Seng Chua

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories:…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Hui Han , Siyuan Li , Jiaqi Chen , Yiwen Yuan , Yuling Wu , Chak Tou Leong , Hanwen Du , Junchen Fu , Youhua Li , Jie Zhang , Chi Zhang , Li-jia Li , Yongxin Ni