English
Related papers

Related papers: AVscript: Accessible Video Editing with Audio-Visu…

200 papers

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news studios, which are…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Eric Zhongcong Xu , Zeyang Song , Satoshi Tsutsui , Chao Feng , Mang Ye , Mike Zheng Shou

Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs). However, the focus of prior research has been predominantly on devising…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Yizhou Wang , Ruiyi Zhang , Haoliang Wang , Uttaran Bhattacharya , Yun Fu , Gang Wu

Machine learning is transforming the video editing industry. Recent advances in computer vision have leveled-up video editing tasks such as intelligent reframing, rotoscoping, color grading, or applying digital makeups. However, most of the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Dawit Mureja Argaw , Fabian Caba Heilbron , Joon-Young Lee , Markus Woodson , In So Kweon

Accessibility of tables on websites for Visually Impaired Persons (VIP) is not optimal with screen readers which are not always effective for the recovery of visual information (2D). Actual Web/Multimedia technologies are not taking in…

Human-Computer Interaction · Computer Science 2019-11-11 Katerine Romeo , E Pissaloux , F Serin

Authors make their videos visually accessible by adding audio descriptions (AD), and auditorily accessible by adding closed captions (CC). However, creating AD and CC is challenging and tedious, especially for non-professional describers…

Human-Computer Interaction · Computer Science 2025-02-19 Xingyu "Bruce" Liu , Ruolin Wang , Dingzeyu Li , Xiang 'Anthony' Chen , Amy Pavel

Effective time management during presentations is challenging, particularly for Blind and Low-Vision (BLV) individuals, as existing tools often lack accessibility and multimodal feedback. To address this gap, we developed vashTimer: a free,…

Human-Computer Interaction · Computer Science 2025-09-25 Aziz N Zeidieh , Sanchita S. Kamath , JooYoung Seo

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their adaptability in open-ended analytical scenarios. The recent…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yuxuan Yan , Shiqi Jiang , Ting Cao , Yifan Yang , Qianqian Yang , Yuanchao Shu , Yuqing Yang , Lili Qiu

Visual Question Answering (VQA) holds great potential for assisting Blind and Low Vision (BLV) users, yet real-world usage remains challenging. Due to visual impairments, BLV users often take blurry or poorly framed photos and face…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Wanyin Cheng , Zanxi Ruan

The development of AI-Generated Video (AIGV) technology has been remarkable in recent years, significantly transforming the paradigm of video content production. However, AIGVs still suffer from noticeable visual quality defects, such as…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Zelu Qi , Ping Shi , Chaoyang Zhang , Shuqi Wang , Fei Zhao , Da Pan , Zefeng Ying

Despite the recent advancement in video stylization, most existing methods struggle to render any video with complex transitions, based on an open style description of user query. To fill this gap, we introduce a generic multi-agent system…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Zhengrong Yue , Shaobin Zhuang , Kunchang Li , Yanbo Ding , Yali Wang

Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop framework that jointly understands intrinsic scene properties…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Ye Fang , Tong Wu , Valentin Deschaintre , Duygu Ceylan , Iliyan Georgiev , Chun-Hao Paul Huang , Yiwei Hu , Xuelin Chen , Tuanfeng Yang Wang

Detecting video deepfakes has become increasingly urgent in recent years. Given the audio-visual information in videos, existing methods typically expose deepfakes by modeling cross-modal correspondence using specifically designed…

Multimedia · Computer Science 2026-04-13 Zihe Wei , Yuezun Li

In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detects the temporal…

People who are blind or have low vision (BLV) encounter numerous challenges in their daily lives and work. To support them, various haptic assistive tools have been developed. Despite these advancements, the effective utilization of these…

Human-Computer Interaction · Computer Science 2024-12-30 Chutian Jiang , Emily Kuang , Mingming Fan

Audiovisual (AV) archives are invaluable for holistically preserving the past. Unlike other forms, AV archives can be difficult to explore. This is not only because of its complex modality and sheer volume but also the lack of appropriate…

Multimedia · Computer Science 2023-10-10 Yuchen Yang , Linyida Zhang

As short videos have risen in popularity, the role of video content in advertising has become increasingly significant. Typically, advertisers record a large amount of raw footage about the product and then create numerous different…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Dongjun Qian , Kai Su , Yiming Tan , Qishuai Diao , Xian Wu , Chang Liu , Bingyue Peng , Zehuan Yuan

Even though large-scale text-to-image generative models show promising performance in synthesizing high-quality images, applying these models directly to image editing remains a significant challenge. This challenge is further amplified in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Shutong Jin , Ruiyu Wang , Florian T. Pokorny

Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, existing diffusion-based video editing approaches lack the ability to offer precise control over generated content that…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Paul Couairon , Clément Rambour , Jean-Emmanuel Haugeard , Nicolas Thome

As information becomes more accessible, user-generated videos are increasing in length, placing a burden on viewers to sift through vast content for valuable insights. This trend underscores the need for an algorithm to extract key video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Lingfeng Yang , Zhenyuan Chen , Xiang Li , Peiyang Jia , Liangqu Long , Jian Yang

Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short video understanding. However, the understanding of long videos…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Haoji Zhang , Yiqin Wang , Yansong Tang , Yong Liu , Jiashi Feng , Xiaojie Jin