English
Related papers

Related papers: CoSight: Exploring Viewer Contributions to Online …

200 papers

Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Existing methods typically apply decoupled attention mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Hannan Lu , Xiaohe Wu , Shudong Wang , Xiameng Qin , Xinyu Zhang , Junyu Han , Wangmeng Zuo , Ji Tao

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang

Humans can watch a continuous video stream and effortlessly perform continual acquisition and transfer of new knowledge with minimal supervision yet retaining previously learnt experiences. In contrast, existing continual learning (CL)…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Jay Zhangjie Wu , David Junhao Zhang , Wynne Hsu , Mengmi Zhang , Mike Zheng Shou

Research in the Vision and Language area encompasses challenging topics that seek to connect visual and textual information. When the visual information is related to videos, this takes us into Video-Text Research, which includes several…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Jesus Perez-Martin , Benjamin Bustos , Silvio Jamil F. Guimarães , Ivan Sipiran , Jorge Pérez , Grethel Coello Said

The ever-increasing amount of user-generated content online has led, in recent years, to an expansion in research and investment in automated content analysis tools. Scrutiny of automated content analysis has accelerated during the COVID-19…

Multimedia · Computer Science 2022-01-27 Carey Shenkman , Dhanaraj Thakur , Emma Llansó

Multi-modal reasoning requires the seamless integration of visual and linguistic cues, yet existing Chain-of-Thought methods suffer from two critical limitations in cross-modal scenarios: (1) over-reliance on single coarse-grained image…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Wenting Lu , Didi Zhu , Tao Shen , Donglin Zhu , Ayong Ye , Chao Wu

Video summarization aims at choosing parts of a video that narrate a story as close as possible to the original one. Most of the existing video summarization approaches focus on hand-crafted labels. As the number of videos grows…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Ivan Sosnovik , Artem Moskalev , Cees Kaandorp , Arnold Smeulders

Social VR has increased in popularity due to its affordances for rich, embodied, and nonverbal communication. However, nonverbal communication remains inaccessible for blind and low vision people in social VR. We designed accessible cues…

Human-Computer Interaction · Computer Science 2024-10-30 Crescentia Jung , Jazmin Collins , Ricardo E. Gonzalez Penuela , Jonathan Isaac Segal , Andrea Stevenson Won , Shiri Azenkot

YouTube is a valuable source of user-generated content on a wide range of topics, and it encourages user participation through the use of a comment system. Video content is increasingly addressing scientific topics, and there is evidence…

Computers and Society · Computer Science 2024-05-22 Sören Striewski , Olga Zagovora , Isabella Peters

Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging when dealing with long-form videos, lasting from minutes to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Soumya Shamarao Jahagirdar , Jayasree Saha , C V Jawahar

Low-code applications are gaining popularity across various fields, enabling non-developers to participate in the software development process. However, due to the strong reliance on graphical user interfaces, they may unintentionally…

Software Engineering · Computer Science 2025-08-11 Mohammadali Mohammadkhani , Sara Zahedi Movahed , Hourieh Khalajzadeh , Mojtaba Shahin , Khuong Tran Hoang

Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich context present in the video such as speaking scene and…

Online services often require users to agree to lengthy and obscure Terms of Service (ToS), leading to information asymmetry and legal risks. This paper proposes TOSense-a Chrome extension that allows users to ask questions about ToS in…

Cryptography and Security · Computer Science 2025-08-04 Xinzhang Chen , Hassan Ali , Arash Shaghaghi , Salil S. Kanhere , Sanjay Jha

Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect, while,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Chenglizhao Chen , Mengke Song , Wenfeng Song , Li Guo , Muwei Jian

Watching movies and TV shows with subtitles enabled is not simply down to audibility or speech intelligibility. A variety of evolving factors related to technological advances, cinema production and social behaviour challenge our perception…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-28 Helard Becerra Martinez , Alessandro Ragano , Diptasree Debnath , Asad Ullah , Crisron Rudolf Lucas , Martin Walsh , Andrew Hines

In this paper we present an approach for supporting users in the difficult task of searching for video. We use collaborative feedback mined from the interactions of earlier users of a video search system to help users in their current…

Information Retrieval · Computer Science 2009-08-07 Frank Hopfgartner , David Vallet , Martin Halvey , Joemon Jose

Multimedia content, such as advertisements and story videos, exhibit a rich blend of creativity and multiple modalities. They incorporate elements like text, visuals, audio, and storytelling techniques, employing devices like emotions,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Aanisha Bhattacharya , Yaman K Singla , Balaji Krishnamurthy , Rajiv Ratn Shah , Changyou Chen

Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has propelled advancements in VLN, the significance of foreground…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Yunbo Xu , Xuesong Zhang , Jia Li , Zhenzhen Hu , Richang Hong

Existing video captioning approaches typically require to first sample video frames from a decoded video and then conduct a subsequent process (e.g., feature extraction and/or captioning model learning). In this pipeline, manual frame…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yaojie Shen , Xin Gu , Kai Xu , Heng Fan , Longyin Wen , Libo Zhang

This paper investigates the accessibility of cookie notices on websites for users with visual impairments (VI) via a set of system studies on top UK websites (n=46) and a user study (n=100). We use a set of methods and tools--including…

Human-Computer Interaction · Computer Science 2024-01-18 James M. Clarke , Maryam Mehrnezhad , Ehsan Toreini