English
Related papers

Related papers: Fine-grained Video Attractiveness Prediction Using…

200 papers

The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Ming Hu , Lin Wang , Siyuan Yan , Don Ma , Qingli Ren , Peng Xia , Wei Feng , Peibo Duan , Lie Ju , Zongyuan Ge

The share of videos in the internet traffic has been growing, therefore understanding how videos capture attention on a global scale is also of growing importance. Most current research focus on modeling the number of views, but we argue…

Social and Information Networks · Computer Science 2018-04-12 Siqi Wu , Marian-Andrei Rizoiu , Lexing Xie

Video matting has traditionally been limited by the lack of high-quality ground-truth data. Most existing video matting datasets provide only human-annotated imperfect alpha and foreground annotations, which must be composited to background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Yongtao Ge , Kangyang Xie , Guangkai Xu , Mingyu Liu , Li Ke , Longtao Huang , Hui Xue , Hao Chen , Chunhua Shen

Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement…

Computer-Aided Design (CAD) is a time-consuming and complex process, requiring precise, long-horizon user interactions with intricate 3D interfaces. While recent advances in AI-driven user interface (UI) agents show promise, most existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Brandon Man , Ghadi Nehme , Md Ferdous Alam , Faez Ahmed

With the motivation of practical gait recognition applications, we propose to automatically create a large-scale synthetic gait dataset (called VersatileGait) by a game engine, which consists of around one million silhouette sequences of…

Computer Vision and Pattern Recognition · Computer Science 2021-01-06 Huanzhang Dou , Wenhu Zhang , Pengyi Zhang , Yuhan Zhao , Songyuan Li , Zequn Qin , Fei Wu , Lin Dong , Xi Li

Short-video platforms show an increasing impact on people's daily lives nowadays, with billions of active users spending plenty of time each day. The interactions between users and online platforms give rise to many scientific problems…

Multimedia · Computer Science 2025-02-11 Yu Shang , Chen Gao , Nian Li , Yong Li

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

Computer Vision and Pattern Recognition · Computer Science 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

The ability to predict, anticipate and reason about future outcomes is a key component of intelligent decision-making systems. In light of the success of deep learning in computer vision, deep-learning-based video prediction emerged as a…

Multimodal counterfactual reasoning is a vital yet challenging ability for AI systems. It involves predicting the outcomes of hypothetical circumstances based on vision and language inputs, which enables AI models to learn from failures and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Te-Lin Wu , Zi-Yi Dou , Qingyuan Hu , Yu Hou , Nischal Reddy Chandra , Marjorie Freedman , Ralph M. Weischedel , Nanyun Peng

Automatic understanding of human affect using visual signals is a problem that has attracted significant interest over the past 20 years. However, human emotional states are quite complex. To appraise such states displayed in real-world…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Dimitrios Kollias , Stefanos Zafeiriou

Visual parsing of images and videos is critical for a wide range of real-world applications. However, progress in this field is constrained by limitations of existing datasets: (1) insufficient annotation granularity, which impedes…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Minghao Zou , Qingtian Zeng , Yongping Miao , Shangkun Liu , Zilong Wang , Hantao Liu , Wei Zhou

Anomaly detection has attracted considerable search attention. However, existing anomaly detection databases encounter two major problems. Firstly, they are limited in scale. Secondly, training sets contain only video-level labels…

Computer Vision and Pattern Recognition · Computer Science 2021-06-17 Boyang Wan , Wenhui Jiang , Yuming Fang , Zhiyuan Luo , Guanqun Ding

Understanding long-form videos, such as movies and TV episodes ranging from tens of minutes to two hours, remains a significant challenge for multi-modal models. Existing benchmarks often fail to test the full range of cognitive skills…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Kirolos Ataallah , Eslam Abdelrahman , Mahmoud Ahmed , Chenhui Gou , Khushbu Pahwa , Jian Ding , Mohamed Elhoseiny

Deep video action recognition models have been highly successful in recent years but require large quantities of manually annotated data, which are expensive and laborious to obtain. In this work, we investigate the generation of synthetic…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Naila Murray , Antonio Manuel López

Data is the foundation for the development of computer vision, and the establishment of datasets plays an important role in advancing the techniques of fine-grained visual categorization~(FGVC). In the existing FGVC datasets used in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Shuo Ye , Shiming Chen , Ruxin Wang , Tianxu Wu , Jiamiao Xu , Salman Khan , Fahad Shahbaz Khan , Ling Shao

YouTube is the leading social media platform for sharing videos. As a result, it is plagued with misleading content that includes staged videos presented as real footages from an incident, videos with misrepresented context and videos where…

Computation and Language · Computer Science 2019-01-28 Priyank Palod , Ayush Patwari , Sudhanshu Bahety , Saurabh Bagchi , Pawan Goyal

Remote work and online courses have become important methods of knowledge dissemination, leading to a large number of document-based instructional videos. Unlike traditional video datasets, these videos mainly feature rich-text images and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Haochen Wang , Kai Hu , Liangcai Gao

Fr\'echet Video Distance (FVD), a prominent metric for evaluating video generation models, is known to conflict with human perception occasionally. In this paper, we aim to explore the extent of FVD's bias toward per-frame quality over…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Songwei Ge , Aniruddha Mahapatra , Gaurav Parmar , Jun-Yan Zhu , Jia-Bin Huang

Emotion AI is the ability of computers to understand human emotional states. Existing works have achieved promising progress, but two limitations remain to be solved: 1) Previous studies have been more focused on short sequential video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Deng Li , Xin Liu , Bohao Xing , Baiqiang Xia , Yuan Zong , Bihan Wen , Heikki Kälviäinen