中文
相关论文

相关论文: MEGA: Multimodal Alignment Aggregation and Distill…

200 篇论文

The proliferation of online short video platforms has driven a surge in user demand for short video editing. However, manually selecting, cropping, and assembling raw footage into a coherent, high-quality video remains laborious and…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Zhihui Yin , Ye Ma , Xipeng Cao , Bo Wang , Quan Chen , Peng Jiang

Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematically analyse the performance of existing multimodal fusion…

多媒体 · 计算机科学 2025-10-10 Han Hu , Dongheng Lin , Qiming Huang , Yuqi Hou , Hyung Jin Chang , Jianbo Jiao

Most existing real-time deep models trained with each frame independently may produce inconsistent results across the temporal axis when tested on a video sequence. A few methods take the correlations in the video sequence into…

计算机视觉与模式识别 · 计算机科学 2022-02-28 Yifan Liu , Chunhua Shen , Changqian Yu , Jingdong Wang

In this work, we propose an efficient Video-Language Alignment (ViLA) network. Our ViLA model addresses both efficient frame sampling and effective cross-modal alignment in a unified way. In our ViLA network, we design a new learnable…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Xijun Wang , Junbang Liang , Chun-Kai Wang , Kenan Deng , Yu Lou , Ming Lin , Shan Yang

Recently, learning open-vocabulary semantic segmentation from text supervision has achieved promising downstream performance. Nevertheless, current approaches encounter an alignment granularity gap owing to the absence of dense annotations,…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Yajie Liu , Pu Ge , Qingjie Liu , Di Huang

Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, applying these techniques directly to video diffusion often results in unsatisfactory frame quality due to the limited visual quality in…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Yuanhao Zhai , Kevin Lin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Chung-Ching Lin , David Doermann , Junsong Yuan , Lijuan Wang

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of…

计算机视觉与模式识别 · 计算机科学 2020-04-29 Anyi Rao , Linning Xu , Yu Xiong , Guodong Xu , Qingqiu Huang , Bolei Zhou , Dahua Lin

Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Han Fang , Xianghao Zang , Chao Ban , Zerun Feng , Lanxiang Zhou , Zhongjiang He , Yongxiang Li , Hao Sun

Video synthesis has recently made remarkable strides benefiting from the rapid development of diffusion models. However, it still encounters challenges in terms of semantic accuracy, clarity and spatio-temporal continuity. They primarily…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Shiwei Zhang , Jiayu Wang , Yingya Zhang , Kang Zhao , Hangjie Yuan , Zhiwu Qin , Xiang Wang , Deli Zhao , Jingren Zhou

Video data is increasingly used alongside conventional data for interactive data exploration, necessitating interfaces for exploring and presenting mixed-modality data. However, integrating video into visualizations remains difficult due to…

人机交互 · 计算机科学 2026-04-29 Dominik Winecki , Arnab Nandi

Dataset distillation has demonstrated remarkable effectiveness in high-compression scenarios for image datasets. While video datasets inherently contain greater redundancy, existing video dataset distillation methods primarily focus on…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Ning Li , Antai Andy Liu , Jingran Zhang , Justin Cui

We propose Compressed Video Aggregator (CVA), a lightweight micro-video recommendation module that decouples video information from preference learning. It aggregates frozen VFM embeddings, and uses latent reasoning without cross-attention…

机器学习 · 计算机科学 2026-05-12 Yang Xiao , Huiyuan Chen , Kaiyuan Deng , Chao Jiang , Zinan Ling , Ruimeng Ye , Xiaolong Ma , Bo Hui

In the era of large-scale pre-trained models, effectively adapting general knowledge to specific affective computing tasks remains a challenge, particularly regarding computational efficiency and multimodal heterogeneity. While…

人工智能 · 计算机科学 2026-03-20 Yan Li , Yifei Xing , Xiangyuan Lan , Xin Li , Haifeng Chen , Dongmei Jiang

Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Thanh-Dat Truong , Christophe Bobda , Nitin Agarwal , Khoa Luu

Multimodal alignment is commonly learned from isolated image-text pairs via CLIP-style dual encoders, leaving the relational context among entities largely unused. Multimodal attributed graphs (MAGs), where nodes carry multimodal attributes…

机器学习 · 计算机科学 2026-05-18 Xu Wang , Xunkai Li , Yinlin Zhu , Rong-Hua Li , Guoren Wang

Event cameras have the potential to revolutionize the field of robot vision, particularly in areas like stereo disparity estimation, owing to their high temporal resolution and high dynamic range. Many studies use deep learning for event…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Junjie Jiang , Hao Zhuang , Xinjie Huang , Delei Kong , Zheng Fang

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity,…

声音 · 计算机科学 2025-09-09 Xiaoran Yang , Jianxuan Yang , Xinyue Guo , Haoyu Wang , Ningning Pan , Gongping Huang

Real-time computational speed and a high degree of precision are requirements for computer-assisted interventions. Applying a segmentation network to a medical video processing task can introduce significant inter-frame prediction noise.…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Robert Mendel , Tobias Rueckert , Dirk Wilhelm , Daniel Rueckert , Christoph Palm

Weakly-supervised action segmentation is a task of learning to partition a long video into several action segments, where training videos are only accompanied by transcripts (ordered list of actions). Most of existing methods need to infer…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Angchi Xu , Wei-Shi Zheng

Temporal modeling is key for action recognition in videos. It normally considers both short-range motions and long-range aggregations. In this paper, we propose a Temporal Excitation and Aggregation (TEA) block, including a motion…

计算机视觉与模式识别 · 计算机科学 2020-04-06 Yan Li , Bin Ji , Xintian Shi , Jianguo Zhang , Bin Kang , Limin Wang