中文
相关论文

相关论文: SPICA: Interactive Video Content Exploration throu…

200 篇论文

Advances in text-based image generation and editing have revolutionized content creation, enabling users to create impressive content from imaginative text prompts. However, existing methods are not designed to work well with the…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Tiancheng Shen , Jun Hao Liew , Long Mai , Lu Qi , Jiashi Feng , Jiaya Jia

There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece is temporal…

计算机视觉与模式识别 · 计算机科学 2024-03-29 De-An Huang , Shijia Liao , Subhashree Radhakrishnan , Hongxu Yin , Pavlo Molchanov , Zhiding Yu , Jan Kautz

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer…

人机交互 · 计算机科学 2026-02-20 Ricardo E. Gonzalez Penuela , Crescentia Jung , Sharon Y Lin , Ruiying Hu , Shiri Azenkot

The Audio Visual Question Answering (AVQA) task aims to answer questions related to various visual objects, sounds, and their interactions in videos. Such naturally multimodal videos contain rich and complex dynamic audio-visual components,…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Guangyao Li , Henghui Du , Di Hu

It has always been a rather tough task to communicate with someone possessing a hearing impairment. One of the most tested ways to establish such a communication is through the use of sign based languages. However, not many people are aware…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Sharanya Mukherjee , Md Hishaam Akhtar , Kannadasan R

Immersive audio-visual perception relies on the spatial integration of both auditory and visual information which are heterogeneous sensing modalities with different fields of reception and spatial resolution. This study investigates the…

音频与语音处理 · 电气工程与系统科学 2020-03-17 Davide Berghi , Hanne Stenzel , Marco Volino , Adrian Hilton , Philip J. B. Jackson

The large-scale digitization of historical archives has created a paradox: "dark data"-digital objects lacking metadata for retrieval. Manual archival description is slow and expensive, limiting discovery and reuse. We propose Vidya, a…

数字图书馆 · 计算机科学 2026-05-19 Cloter Migliorini Filho , Julia Graciela Machado , Edson Armando Silva , Marcella Scoczynski

We introduce VIBA, a novel approach for explainable video classification by adapting Information Bottlenecks for Attribution (IBA) to video sequences. While most traditional explainability methods are designed for image models, our IBA…

计算机视觉与模式识别 · 计算机科学 2025-01-29 Veronika Solopova , Lucas Schmidt , Dorothea Kolossa

Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video question answering systems employ rigid pipelines that…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Zixuan Dong , Baoyun Peng , Yufei Wang , Lin Liu , Xinxin Dong , Yunlong Cao , Xiaodong Wang

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

多媒体 · 计算机科学 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

360$^\circ$ videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond pre-determined normal field of views and displays distinctive spatial relations on a sphere. However, previous benchmark tasks…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Heeseung Yun , Youngjae Yu , Wonsuk Yang , Kangil Lee , Gunhee Kim

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite…

计算与语言 · 计算机科学 2025-10-02 Kimihiro Hasegawa , Wiradee Imrattanatrai , Masaki Asada , Ken Fukuda , Teruko Mitamura

Vision and language understanding has emerged as a subject undergoing intense study in Artificial Intelligence. Among many tasks in this line of research, visual question answering (VQA) has been one of the most successful ones, where the…

计算机视觉与模式识别 · 计算机科学 2017-12-05 Yunseok Jang , Yale Song , Youngjae Yu , Youngjin Kim , Gunhee Kim

Movie Audio Description (AD) aims to narrate visual content during dialogue-free segments, particularly benefiting blind and visually impaired (BVI) audiences. Compared with general video captioning, AD demands plot-relevant narration with…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Xiaojun Ye , Chun Wang , Yiren Song , Sheng Zhou , Liangcheng Li , Jiajun Bu

Non-invasive steady-state visual evoked potential (SSVEP) based brain-computer interface (BCI) systems offer high bandwidth compared to other BCI types and require only minimal calibration and training. Virtual reality (VR) has been already…

Question-asking is one of the key indicators of cognitive engagement. However, understanding how the distinct psychological affordances of presentation media shape learners' spoken inquiries with embodied Intelligent Virtual Agents (IVAs)…

人机交互 · 计算机科学 2026-03-17 Hyerim Park , Jinseok Hong , Heejeong Ko , Woontack Woo

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Guangyao Li , Yake Wei , Yapeng Tian , Chenliang Xu , Ji-Rong Wen , Di Hu

As social virtual reality (VR) grows more popular, addressing accessibility for blind and low vision (BLV) users is increasingly critical. Researchers have proposed an AI "sighted guide" to help users navigate VR and answer their questions,…

人机交互 · 计算机科学 2026-03-31 Jazmin Collins , Sharon Y Lin , Tianqi Liu , Andrea Stevenson Won , Shiri Azenkot

LLaVA-Plus is a general-purpose multimodal assistant that expands the capabilities of large multimodal models. It maintains a skill repository of pre-trained vision and vision-language models and can activate relevant tools based on users'…

计算机视觉与模式识别 · 计算机科学 2023-11-10 Shilong Liu , Hao Cheng , Haotian Liu , Hao Zhang , Feng Li , Tianhe Ren , Xueyan Zou , Jianwei Yang , Hang Su , Jun Zhu , Lei Zhang , Jianfeng Gao , Chunyuan Li

Biometric authentication techniques are more consistent and efficient than conventional authentication techniques and can be used in monitoring, transaction authentication, information retrieval, access control, forensics, etc. In this…

声音 · 计算机科学 2010-04-27 Anuj Mehra , Anupam Shukla , Mahender Kumawat , Rajiv Ranjan , Ritu Tiwari