English
Related papers

Related papers: SPICA: Interactive Video Content Exploration throu…

200 papers

This paper provides recommendations on how to integrate accessibility solutions, like subtitling, audio description and sign language, with immersive media services, with a focus on 360-degree video and spatial audio. It provides an…

Multimedia · Computer Science 2020-05-11 Peter tho Pesch , Romain Bouqueau , Mario Montagud

While audio guides can offer rich information about an exhibit, it is challenging for visitors to focus on specific exhibit details based only on the verbal description. We present \textit{CLIO}, a tour guide robot with co-speech actions to…

Human-Computer Interaction · Computer Science 2025-12-08 Yuxuan Chen , Ian Leong Ting Lo , Bao Guo , Netitorn Kawmali , Chun Kit Chan , Ruoyu Wang , Jia Pan , Lei Yang

Modern information querying systems are progressively incorporating multimodal inputs like vision and audio. However, the integration of gaze -- a modality deeply linked to user intent and increasingly accessible via gaze-tracking wearables…

Human-Computer Interaction · Computer Science 2024-05-14 Zeyu Wang , Yuanchun Shi , Yuntao Wang , Yuchen Yao , Kun Yan , Yuhan Wang , Lei Ji , Xuhai Xu , Chun Yu

In recent years, online lecture videos have become an increasingly popular resource for acquiring new knowledge. Systems capable of effectively understanding/indexing lecture videos are thus highly desirable, enabling downstream tasks like…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Kangda Wei , Zhengyu Zhou , Bingqing Wang , Jun Araki , Lukas Lange , Ruihong Huang , Zhe Feng

Large shared displays, such as digital whiteboards, are useful for supporting co-located team collaborations by helping members perform cognitive tasks such as brainstorming, organizing ideas, and making comparisons. While recent…

Human-Computer Interaction · Computer Science 2025-02-10 Zheng Zhang , Weirui Peng , Xinyue Chen , Luke Cao , Toby Jia-Jun Li

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Xiaoyan Cong , Haotian Yang , Angtian Wang , Yizhi Wang , Yiding Yang , Canyu Zhang , Chongyang Ma

Blind Image Quality Assessment (BIQA) is essential for automatically evaluating the perceptual quality of visual signals without access to the references. In this survey, we provide a comprehensive analysis and discussion of recent…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Miaohui Wang

Recent advancements in video understanding within visual large language models (VLLMs) have led to notable progress. However, the complexity of video data and contextual processing limitations still hinder long-video comprehension. A common…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Yanan Guo , Wenhui Dong , Jun Song , Shiding Zhu , Xuan Zhang , Hanqing Yang , Yingbo Wang , Yang Du , Xianing Chen , Bo Zheng

Long video understanding (LVU) is challenging because answering real-world queries often depends on sparse, temporally dispersed cues buried in hours of mostly redundant and irrelevant content. While agentic pipelines improve video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Ziyang Wang , Honglu Zhou , Shijie Wang , Junnan Li , Caiming Xiong , Silvio Savarese , Mohit Bansal , Michael S. Ryoo , Juan Carlos Niebles

Universal video understanding requires modeling fine-grained visual and audio information over time in diverse real-world scenarios. However, the performance of existing models is primarily constrained by video-instruction data that…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yunheng Li , Hengrui Zhang , Meng-Hao Guo , Wenzhao Gao , Shaoyong Jia , Shaohui Jiao , Qibin Hou , Ming-Ming Cheng

Blind image quality assessment (BIQA) aims to predict perceptual image quality scores without access to reference images. State-of-the-art BIQA methods typically require subjects to score a large number of images to train a robust model.…

Computer Vision and Pattern Recognition · Computer Science 2019-04-25 Fei Gao , Dacheng Tao , Xinbo Gao , Xuelong Li

Audio-visual video parsing focuses on classifying videos through weak labels while identifying events as either visible, audible, or both, alongside their respective temporal boundaries. Many methods ignore that different modalities often…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Langyu Wang , Bingke Zhu , Yingying Chen , Jinqiao Wang

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, multi-action…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Alejandro Aparcedo , Akash Kumar , Aaryan Garg , Dalton Pham , Wen-Kai Chen , Anirudh Bharadwaj , Aman Chadha , Yogesh Rawat

Augmented Reality (AR) is transforming the way we interact with virtual information in the physical world. By overlaying digital content in real-world environments, AR enables new forms of immersive and engaging experiences. However,…

Human-Computer Interaction · Computer Science 2025-04-24 Julian Rasch , Florian Müller , Francesco Chiossi

Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs). However, the focus of prior research has been predominantly on devising…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Yizhou Wang , Ruiyi Zhang , Haoliang Wang , Uttaran Bhattacharya , Yun Fu , Gang Wu

As virtual 3D environments become more prevalent, equitable access is essential for blind and low-vision (BLV) users, who face challenges with spatial awareness, navigation, and interaction. Prior work has explored supplementing visual…

Human-Computer Interaction · Computer Science 2026-02-10 Xinyun Cao , Kexin Phyllis Ju , Chenglin Li , Venkatesh Potluri , Dhruv Jain

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Xinlong Chen , Yue Ding , Weihong Lin , Jingyun Hua , Linli Yao , Yang Shi , Bozhou Li , Yuanxing Zhang , Qiang Liu , Pengfei Wan , Liang Wang , Tieniu Tan

Class-Incremental Learning (CIL) enables models to continuously integrate new knowledge while mitigating catastrophic forgetting. Driven by the remarkable generalization of CLIP, leveraging pre-trained vision-language models has become a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Hao Sun , Zi-Jun Ding , Da-Wei Zhou

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs)…

Human-Computer Interaction · Computer Science 2025-07-28 Mina Huh , Zihui Xue , Ujjaini Das , Kumar Ashutosh , Kristen Grauman , Amy Pavel

The panoramic video is widely used to build virtual reality (VR) and is expected to be one of the next generation Killer-Apps. Transmitting panoramic VR videos is a challenging task because of two problems: 1) panoramic VR videos are…

Multimedia · Computer Science 2017-04-24 Lun Wang , Damai Dai , Jie Jiang , Tong Yang , Xiaoke Jiang , Zekun Cai , Yang Li , Xiaoming Li
‹ Prev 1 8 9 10 Next ›