English
Related papers

Related papers: Visual Summary of Egocentric Photostreams by Repre…

200 papers

Event cameras are biologically-inspired sensors that gather the temporal evolution of the scene. They capture pixel-wise brightness variations and output a corresponding stream of asynchronous events. Despite having multiple advantages with…

Computer Vision and Pattern Recognition · Computer Science 2019-12-11 Stefano Pini , Guido Borghi , Roberto Vezzani

Recent Transformer-based summarization models have provided a promising approach to abstractive summarization. They go beyond sentence selection and extractive strategies to deal with more complicated tasks such as novel word generation and…

Computation and Language · Computer Science 2023-02-09 Sajad Sotudeh , Hanieh Deilamsalehy , Franck Dernoncourt , Nazli Goharian

In this paper we introduce LifelongMemory, a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval. LifelongMemory generates concise video activity descriptions of the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Ying Wang , Yanlai Yang , Mengye Ren

Monocular egocentric 3D human motion capture is a challenging and actively researched problem. Existing methods use synchronously operating visual sensors (e.g. RGB cameras) and often fail under low lighting and fast motions, which can be…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Christen Millerdurai , Hiroyasu Akada , Jian Wang , Diogo Luvizon , Christian Theobalt , Vladislav Golyanik

In this paper, we address the problem of unsupervised video summarization that automatically extracts key-shots from an input video. Specifically, we tackle two critical issues based on our empirical observations: (i) Ineffective feature…

Computer Vision and Pattern Recognition · Computer Science 2018-11-27 Yunjae Jung , Donghyeon Cho , Dahun Kim , Sanghyun Woo , In So Kweon

Video summarization aims to select keyframes that are visually diverse and can represent the whole story of a given video. Previous approaches have focused on global interlinkability between frames in a video by temporal modeling. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Jungin Park , Jiyoung Lee , Kwanghoon Sohn

We present a domain- and user-preference-agnostic approach to detect highlightable excerpts from human-centric videos. Our method works on the graph-based representation of multiple observable human-centric modalities in the videos, such as…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Uttaran Bhattacharya , Gang Wu , Stefano Petrangeli , Viswanathan Swaminathan , Dinesh Manocha

Research in child development has shown that embodied experience handling physical objects contributes to many cognitive abilities, including visual learning. One characteristic of such experience is that the learner sees the same object…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Deepayan Sanyal , Joel Michelson , Yuan Yang , James Ainooson , Maithilee Kunda

Event cameras are bio-inspired sensors that capture the per-pixel intensity changes asynchronously and produce event streams encoding the time, pixel position, and polarity (sign) of the intensity changes. Event cameras possess a myriad of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Xu Zheng , Yexin Liu , Yunfan Lu , Tongyan Hua , Tianbo Pan , Weiming Zhang , Dacheng Tao , Lin Wang

Sentence extraction based summarization methods has some limitations as it doesn't go into the semantics of the document. Also, it lacks the capability of sentence generation which is intuitive to humans. Here we present a novel method to…

Computation and Language · Computer Science 2014-06-06 Divyanshu Bhartiya , Ashudeep Singh

Event cameras produce asynchronous event streams that are spatially sparse yet temporally dense. Mainstream event representation learning algorithms typically use event frames, voxels, or tensors as input. Although these approaches have…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Futian Wang , Fan Zhang , Xiao Wang , Mengqi Wang , Dexing Huang , Jin Tang

Keyframes are LiDAR scans saved for future reference in Simultaneous Localization And Mapping (SLAM), but despite their central importance most algorithms leave choices of which scans to save and how to use them to wasteful heuristics. This…

Robotics · Computer Science 2025-04-18 David Thorne , Nathan Chan , Yanlong Ma , Christa S. Robison , Philip R. Osteen , Brett T. Lopez

Large collections of videos are grouped into clusters by a topic keyword, such as Eiffel Tower or Surfing, with many important visual concepts repeating across them. Such a topically close set of videos have mutual influence on each other,…

Computer Vision and Pattern Recognition · Computer Science 2017-06-13 Rameswar Panda , Amit K. Roy-Chowdhury

This paper introduces neck-mounted view gaze estimation, a new task that estimates user gaze from the neck-mounted camera perspective. Prior work on egocentric gaze estimation, which predicts device wearer's gaze location within the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Haoyu Huang , Yoichi Sato

Key frame extraction algorithms consider the problem of selecting a subset of the most informative frames from a video to summarize its content.

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Chinh Dang , Abdolreza Moghadam , Hayder Radha

We propose a novel agglomerative clustering method based on unmasking, a technique that was previously used for authorship verification of text documents and for abnormal event detection in videos. In order to join two clusters, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-03 Mariana-Iuliana Georgescu , Radu Tudor Ionescu

Video summarization aims at generating concise video summaries from the lengthy videos, to achieve better user watching experience. Due to the subjectivity, purely supervised methods for video summarization may bring the inherent errors…

Computer Vision and Pattern Recognition · Computer Science 2020-07-30 Tianyu Liu

We consider the problem of localizing visitors in a cultural site from egocentric (first person) images. Localization information can be useful both to assist the user during his visit (e.g., by suggesting where to go and what to see next)…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Francesco Ragusa , Antonino Furnari , Sebastiano Battiato , Giovanni Signorello , Giovanni Maria Farinella

Supervised object detection has been proven to be successful in many benchmark datasets achieving human-level performances. However, acquiring a large amount of labeled image samples for supervised detection training is tedious,…

Computer Vision and Pattern Recognition · Computer Science 2021-05-12 Bishwo Adhikari , Esa Rahtu , Heikki Huttunen

This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker's view of the…

Sound · Computer Science 2023-05-04 Roshan Sharma , Weipeng He , Ju Lin , Egor Lakomkin , Yang Liu , Kaustubh Kalgaonkar