中文
相关论文

相关论文: Personalizing Fast-Forward Videos Based on Visual …

200 篇论文

Every hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increased attention over…

计算机视觉与模式识别 · 计算机科学 2022-08-04 Alex Falcon , Giuseppe Serra , Oswald Lanz

Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generation suffer from tackling multiple subjects, which is a more…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Zhao Wang , Aoxue Li , Lingting Zhu , Yong Guo , Qi Dou , Zhenguo Li

In this paper, the problem of head movement prediction for virtual reality videos is studied. In the considered model, a deep learning network is introduced to leverage position data as well as video frame content to predict future head…

计算机视觉与模式识别 · 计算机科学 2020-03-03 Xinwei Chen , Ali Taleb Zadeh Kasgari , Walid Saad

Fr\'echet Video Distance (FVD), a prominent metric for evaluating video generation models, is known to conflict with human perception occasionally. In this paper, we aim to explore the extent of FVD's bias toward per-frame quality over…

计算机视觉与模式识别 · 计算机科学 2024-04-19 Songwei Ge , Aniruddha Mahapatra , Gaurav Parmar , Jun-Yan Zhu , Jia-Bin Huang

To date, the privacy-protection intended pixelation tasks are still labor-intensive and yet to be studied. With the prevailing of video live streaming, establishing an online face pixelation mechanism during streaming is an urgency. In this…

计算机视觉与模式识别 · 计算机科学 2021-01-06 Jizhe Zhou , Chi-Man Pun

Mass utilization of body-worn cameras has led to a huge corpus of available egocentric video. Existing video summarization algorithms can accelerate browsing such videos by selecting (visually) interesting shots from them. Nonetheless,…

计算机视觉与模式识别 · 计算机科学 2020-09-22 Aidean Sharghi , Niels da Vitoria Lobo , Mubarak Shah

The understanding of human-object interactions is fundamental in First Person Vision (FPV). Visual tracking algorithms which follow the objects manipulated by the camera wearer can provide useful information to effectively model such…

计算机视觉与模式识别 · 计算机科学 2022-10-20 Matteo Dunnhofer , Antonino Furnari , Giovanni Maria Farinella , Christian Micheloni

We wish to automatically predict the "speediness" of moving objects in videos---whether they move faster, at, or slower than their "natural" speed. The core component in our approach is SpeedNet---a novel deep network trained to detect if a…

计算机视觉与模式识别 · 计算机科学 2020-07-28 Sagie Benaim , Ariel Ephrat , Oran Lang , Inbar Mosseri , William T. Freeman , Michael Rubinstein , Michal Irani , Tali Dekel

Multimedia systems underpin modern digital interactions, facilitating seamless integration and optimization of resources across diverse multimedia applications. To meet growing personalization demands, multimedia systems must efficiently…

多媒体 · 计算机科学 2025-08-26 Yili Jin , Ling Pan , Rui-Xiao Zhang , Jiangchuan Liu , Xue Liu

Understanding human-object interactions is fundamental in First Person Vision (FPV). Tracking algorithms which follow the objects manipulated by the camera wearer can provide useful cues to effectively model such interactions. Visual…

计算机视觉与模式识别 · 计算机科学 2022-05-10 Matteo Dunnhofer , Antonino Furnari , Giovanni Maria Farinella , Christian Micheloni

In cinema, large camera lenses create beautiful shallow depth of field (DOF), but make focusing difficult and expensive. Accurate cinema focus usually relies on a script and a person to control focus in realtime. Casual videographers often…

计算机视觉与模式识别 · 计算机科学 2019-05-22 Xuaner Zhang , Kevin Matzen , Vivien Nguyen , Dillon Yao , You Zhang , Ren Ng

Videos captured by consumer cameras often exhibit temporal variations in color and tone that are caused by camera auto-adjustments like white-balance and exposure. When such videos are sub-sampled to play fast-forward, as in the…

图形学 · 计算机科学 2017-10-02 Xuaner Cecilia Zhang , Joon-Young Lee , Kalyan Sunkavalli , Zhaowen Wang

This paper addresses automatic summarization and search in visual data comprising of videos, live streams and image collections in a unified manner. In particular, we propose a framework for multi-faceted summarization which extracts…

计算机视觉与模式识别 · 计算机科学 2017-04-06 Anurag Sahoo , Vishal Kaushal , Khoshrav Doctor , Suyash Shetty , Rishabh Iyer , Ganesh Ramakrishnan

Recently, video captioning has been attracting an increasing amount of interest, due to its potential for improving accessibility and information retrieval. While existing methods rely on different kinds of visual features and model…

计算机视觉与模式识别 · 计算机科学 2016-12-02 Xiang Long , Chuang Gan , Gerard de Melo

We present an efficient framework that can generate a coherent paragraph to describe a given video. Previous works on video captioning usually focus on video clips. They typically treat an entire video as a whole and generate the caption…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Yilei Xiong , Bo Dai , Dahua Lin

Automatically describing video content with natural language has been attracting much attention in CV and NLP communities. Most existing methods predict one word at a time, and by feeding the last generated word back as input at the next…

计算机视觉与模式识别 · 计算机科学 2019-11-06 Huanhou Xiao , Jinglun Shi

We present an approach to robot learning from egocentric human videos by modeling human preferences in a reward function and optimizing robot behavior to maximize this reward. Prior work on reward learning from human videos attempts to…

机器人学 · 计算机科学 2026-02-13 Mrinal Verghese , Christopher G. Atkeson

The overwhelming amount and rate of information update in online social media is making it increasingly difficult for users to allocate their attention to their topics of interest, thus there is a strong need for prioritizing news feeds.…

社会与信息网络 · 计算机科学 2015-11-16 Mehrdad Farajtabar , Safoora Yousefi , Long Q. Tran , Le Song , Hongyuan Zha

Image captioning bridges the gap between vision and language by automatically generating natural language descriptions for images. Traditional image captioning methods often overlook the preferences and characteristics of users.…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Xuan Wang , Guanhong Wang , Wenhao Chai , Jiayu Zhou , Gaoang Wang

We present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multiple observable human-centric modalities in the videos, such as…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Uttaran Bhattacharya , Gang Wu , Stefano Petrangeli , Viswanathan Swaminathan , Dinesh Manocha