中文
相关论文

相关论文: Extraction and Summarization of Explicit Video Con…

200 篇论文

In response to the ongoing COVID-19 pandemic, we present a robust deep learning pipeline that is capable of identifying correct and incorrect mask-wearing from real-time video streams. To accomplish this goal, we devised two separate…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Yuchen Ding , Zichen Li , David Yastremsky

This paper presents a novel approach to synthesize automatically age-progressed facial images in video sequences using Deep Reinforcement Learning. The proposed method models facial structures and the longitudinal face-aging process of…

计算机视觉与模式识别 · 计算机科学 2019-04-25 Chi Nhan Duong , Khoa Luu , Kha Gia Quach , Nghia Nguyen , Eric Patterson , Tien D. Bui , Ngan Le

In this paper, we address several inadequacies of current video object segmentation pipelines. Firstly, a cyclic mechanism is incorporated to the standard semi-supervised process to produce more robust representations. By relying on the…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Yuxi Li , Ning Xu , Jinlong Peng , John See , Weiyao Lin

Recent development of 3D sensors allows the acquisition of extremely dense 3D point clouds of large-scale scenes. The main challenge of processing such large point clouds remains in the size of the data, which induce expensive computational…

计算机视觉与模式识别 · 计算机科学 2021-09-24 Thomas Richard , Florent Dupont , Guillaume Lavoue

Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging when dealing with long-form videos, lasting from minutes to…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Soumya Shamarao Jahagirdar , Jayasree Saha , C V Jawahar

To date, machine learning for human action recognition in video has been widely implemented in sports activities. Although some studies have been successful in the past, precision is still the most significant concern. In this study, we…

计算机视觉与模式识别 · 计算机科学 2020-12-02 Cheng Yan , Xin Li , Guoqiang Li

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

Classification of document images is a critical step for archival of old manuscripts, online subscription and administrative procedures. Computer vision and deep learning have been suggested as a first solution to classify documents based…

计算机视觉与模式识别 · 计算机科学 2019-07-16 Nicolas Audebert , Catherine Herold , Kuider Slimani , Cédric Vidal

In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are…

计算机视觉与模式识别 · 计算机科学 2015-05-05 Vignesh Ramanathan , Kevin Tang , Greg Mori , Li Fei-Fei

The number of videos being produced and consequently stored in databases for video streaming platforms has been increasing exponentially over time. This vast database should be easily index-able to find the requisite clip or video to match…

机器学习 · 计算机科学 2021-01-01 Sriram Krishna , Siddarth Vinay , Srinivas K S

In the current information era, customer analytics play a key role in the success of any business. Since customer demographics primarily dictate their preferences, identification and utilization of age & gender information of customers in…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Earnest Paul Ijjina , Goutham Kanahasabai , Aniruddha Srinivas Joshi

In recent years, there have been numerous developments towards solving multimodal tasks, aiming to learn a stronger representation than through a single modality. Certain aspects of the data can be particularly useful in this case - for…

机器学习 · 统计学 2023-09-06 Cătălina Cangea , Petar Veličković , Pietro Liò

Videos represent the primary source of information for surveillance applications and are available in large amounts but in most cases contain little or no annotation for supervised learning. This article reviews the state-of-the-art deep…

计算机视觉与模式识别 · 计算机科学 2018-01-31 B Ravi Kiran , Dilip Mathew Thomas , Ranjith Parakkal

The real-world is inherently multi-modal at its core. Our tools observe and take snapshots of it, in digital form, such as videos or sounds, however much of it is lost. Similarly for actions and information passing between humans, languages…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Mihai-Cristian Pîrvu , Marius Leordeanu

Film media is a rich form of artistic expression. Unlike photography, and short videos, movies contain a storyline that is deliberately complex and intricate in order to engage its audience. In this paper we present a large scale study…

计算机视觉与模式识别 · 计算机科学 2019-08-09 Paola Cascante-Bonilla , Kalpathy Sitaraman , Mengjia Luo , Vicente Ordonez

Video summarization is among challenging tasks in computer vision, which aims at identifying highlight frames or shots over a lengthy video input. In this paper, we propose an novel attention-based framework for video summarization with…

计算机视觉与模式识别 · 计算机科学 2020-06-04 Yen-Ting Liu , Yu-Jhe Li , Yu-Chiang Frank Wang

The exponential increase in video content poses significant challenges in terms of efficient navigation, search, and retrieval, thus requiring advanced video summarization techniques. Existing video summarization methods, which heavily rely…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Min Jung Lee , Dayoung Gong , Minsu Cho

Automated Machine Learning (AutoML) is a promising direction for democratizing AI by automatically deploying Machine Learning systems with minimal human expertise. The core technical challenge behind AutoML is optimizing the pipelines of…

机器学习 · 计算机科学 2023-05-26 Sebastian Pineda Arango , Josif Grabocka

Extreme Multimodal Summarization with Multimodal Output (XMSMO) becomes an attractive summarization approach by integrating various types of information to create extremely concise yet informative summaries for individual modalities.…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Sicheng Liu , Lintao Wang , Xiaogang Zhu , Xuequan Lu , Zhiyong Wang , Kun Hu

Contemporary deep learning based video captioning follows encoder-decoder framework. In encoder, visual features are extracted with 2D/3D Convolutional Neural Networks (CNNs) and a transformed version of those features is passed to the…

计算机视觉与模式识别 · 计算机科学 2019-11-22 Nayyer Aafaq , Naveed Akhtar , Wei Liu , Ajmal Mian