中文
相关论文

相关论文: Learning Audio-guided Video Representation with Ga…

200 篇论文

Understanding human instructions is essential for enabling smooth human-robot interaction. In this work, we focus on object grounding, i.e., localizing an object of interest in a visual scene (e.g., an image) based on verbal human…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Joel Alberto Santos , Zongwei Wu , Xavier Alameda-Pineda , Radu Timofte

Diverse and extensive work has recently been conducted on text-conditioned human motion generation. However, progress in the reverse direction, motion captioning, has seen less comparable advancement. In this paper, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Karim Radouane , Julien Lagarde , Sylvie Ranwez , Andon Tchechmedjiev

Video-text retrieval (VTR) is an attractive yet challenging task for multi-modal understanding, which aims to search for relevant video (text) given a query (video). Existing methods typically employ completely heterogeneous visual-textual…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Haoran Wang , Di Xu , Dongliang He , Fu Li , Zhong Ji , Jungong Han , Errui Ding

Voice, as input, has progressively become popular on mobiles and seems to transcend almost entirely text input. Through voice, the voice search (VS) system can provide a more natural way to meet user's information needs. However, errors…

信息检索 · 计算机科学 2023-09-06 Yi-Cheng Wang , Tzu-Ting Yang , Hsin-Wei Wang , Bi-Cheng Yan , Berlin Chen

With the development of internet of things technologies, tremendous sensor audio data has been produced, which poses great challenges to audio-based event detection in smart cities. In this paper, we target a challenging audio-based event…

声音 · 计算机科学 2023-12-27 Haoyu Tang , Yunxiao Wang , Jihua Zhu , Shuaike Zhang , Mingzhu Xu , Qinghai Zheng , Yupeng Hu

Video-text retrieval has many real-world applications such as media analytics, surveillance, and robotics. This paper presents the 1st place solution to the video retrieval track of the ICCV VALUE Challenge 2021. We present a simple yet…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Aiden Seungjoon Lee , Hanseok Oh , Minjoon Seo

In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed video that uses audio as a preview mechanism to eliminate both…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Ruohan Gao , Tae-Hyun Oh , Kristen Grauman , Lorenzo Torresani

We introduce a novel deep learning-based audio-visual quality (AVQ) prediction model that leverages internal features from state-of-the-art unimodal predictors. Unlike prior approaches that rely on simple fusion strategies, our model…

音频与语音处理 · 电气工程与系统科学 2026-01-23 Ina Salaj , Arijit Biswas

Generating videos for visual storytelling can be a tedious and complex process that typically requires either live-action filming or graphics animation rendering. To bypass these challenges, our key idea is to utilize the abundance of…

计算机视觉与模式识别 · 计算机科学 2023-07-14 Yingqing He , Menghan Xia , Haoxin Chen , Xiaodong Cun , Yuan Gong , Jinbo Xing , Yong Zhang , Xintao Wang , Chao Weng , Ying Shan , Qifeng Chen

Visual texts embedded in videos carry rich semantic information, which is crucial for both holistic video understanding and fine-grained reasoning about local human actions. However, existing video understanding benchmarks largely overlook…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Zhoufaran Yang , Yan Shu , Jing Wang , Zhifei Yang , Yan Zhang , Yu Li , Keyang Lu , Gangyan Zeng , Shaohui Liu , Yu Zhou , Nicu Sebe

The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cost of fine-grained manual annotations. Recent works attempt…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Kyungbok Lee , You Zhang , Zhiyao Duan

Weakly supervised multimodal video anomaly detection has gained significant attention, yet the potential of the text modality remains under-explored. Text provides explicit semantic information that can enhance anomaly characterization and…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Shengyang Sun , Jiashen Hua , Junyi Feng , Xiaojin Gong

Video generation is experiencing rapid growth, driven by advances in diffusion models and the development of better and larger datasets. However, producing high-quality videos remains challenging due to the high-dimensional data and the…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Elia Peruzzo , Dejia Xu , Xingqian Xu , Humphrey Shi , Nicu Sebe

The joint understanding of vision and language has been recently gaining a lot of attention in both the Computer Vision and Natural Language Processing communities, with the emergence of tasks such as image captioning, image-text matching,…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Matteo Stefanini , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

Partially Relevant Video Retrieval~(PRVR) aims to retrieve a video where a specific segment is relevant to a given text query. Typical training processes of PRVR assume a one-to-one relationship where each text query is relevant to only one…

计算机视觉与模式识别 · 计算机科学 2025-06-10 CH Cho , WJ Moon , W Jun , MS Jung , JP Heo

Video restoration aims at restoring multiple high-quality frames from multiple low-quality frames. Existing video restoration methods generally fall into two extreme cases, i.e., they either restore all frames in parallel or restore the…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Jingyun Liang , Yuchen Fan , Xiaoyu Xiang , Rakesh Ranjan , Eddy Ilg , Simon Green , Jiezhang Cao , Kai Zhang , Radu Timofte , Luc Van Gool

Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Mohamed Eltahir , Osamah Sarraj , Mohammed Bremoo , Mohammed Khurd , Abdulrahman Alfrihidi , Taha Alshatiri , Mohammad Almatrafi , Tanveer Hussain

Despite recent advancements in computer vision research, object detection in aerial images still suffers from several challenges. One primary challenge to be mitigated is the presence of multiple types of variation in aerial images, for…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Sungjune Park , Hyunjun Kim , Beomchan Park , Yong Man Ro

Multimodal Large Language Models have advanced AI in applications like text-to-video generation and visual question answering. These models rely on visual encoders to convert non-text data into vectors, but current encoders either lack…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Junjie Li , Jianghong Ma , Xiaofeng Zhang , Yuhang Li , Jianyang Shi

The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and…

多媒体 · 计算机科学 2025-08-06 Baoquan Zhao , Xiaofan Ma , Qianshi Pang , Ruomei Wang , Fan Zhou , Shujin Lin