中文
相关论文

相关论文: VideoMaMa: Mask-Guided Video Matting via Generativ…

200 篇论文

Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Ruihao Xia , Yu Liang , Peng-Tao Jiang , Hao Zhang , Qianru Sun , Yang Tang , Bo Li , Pan Zhou

Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections. Existing…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Yao-Chih Lee , Erika Lu , Sarah Rumbley , Michal Geyer , Jia-Bin Huang , Tali Dekel , Forrester Cole

The recent segmentation foundation model, Segment Anything Model (SAM), exhibits strong zero-shot segmentation capabilities, but it falls short in generating fine-grained precise masks. To address this limitation, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Beomyoung Kim , Chanyong Shin , Joonhyun Jeong , Hyungsik Jung , Se-Yun Lee , Sewhan Chun , Dong-Hyun Hwang , Joonsang Yu

Masked Autoencoders (MAEs) learn generalizable representations for image, text, audio, video, etc., by reconstructing masked input data from tokens of the visible data. Current MAE approaches for videos rely on random patch, tube, or…

计算机视觉与模式识别 · 计算机科学 2022-11-17 Wele Gedara Chaminda Bandara , Naman Patel , Ali Gholami , Mehdi Nikkhah , Motilal Agrawal , Vishal M. Patel

Real-world image matting is essential for applications in content creation and augmented reality. However, it remains challenging due to the complex nature of scenes and the scarcity of high-quality datasets. To address these limitations,…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Rui Liu

Self-supervised video transformer pre-training has recently benefited from the mask-and-predict pipeline. They have demonstrated outstanding effectiveness on downstream video tasks and superior data efficiency on small datasets. However,…

计算机视觉与模式识别 · 计算机科学 2022-10-12 Yuxin Song , Min Yang , Wenhao Wu , Dongliang He , Fu Li , Jingdong Wang

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of…

计算机视觉与模式识别 · 计算机科学 2018-12-04 Ting-Chun Wang , Ming-Yu Liu , Jun-Yan Zhu , Guilin Liu , Andrew Tao , Jan Kautz , Bryan Catanzaro

Image-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accurately modeling…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Chenhui Zhu , Yilu Wu , Shuai Wang , Gangshan Wu , Limin Wang

Masked autoencoding has shown excellent performance on self-supervised video representation learning. Temporal redundancy has led to a high masking ratio and customized masking strategy in VideoMAE. In this paper, we aim to further improve…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Bingkun Huang , Zhiyu Zhao , Guozhen Zhang , Yu Qiao , Limin Wang

Pre-training video transformers generally requires a large amount of data, presenting significant challenges in terms of data collection costs and concerns related to privacy, licensing, and inherent biases. Synthesizing data is one of the…

计算机视觉与模式识别 · 计算机科学 2024-09-11 Yuchi Ishikawa , Masayoshi Kondo , Yoshimitsu Aoki

Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous demonstrations, and end-to-end pixel-to-action pipelines may degenerate under background and viewpoint…

机器学习 · 计算机科学 2025-12-23 Yao Feng , Hengkai Tan , Xinyi Mao , Chendong Xiang , Guodong Liu , Shuhe Huang , Hang Su , Jun Zhu

Transformer-based architectures have become competitive across a variety of visual domains, most notably images and videos. While prior work studies these modalities in isolation, having a common architecture suggests that one can train a…

计算机视觉与模式识别 · 计算机科学 2023-06-01 Rohit Girdhar , Alaaeldin El-Nouby , Mannat Singh , Kalyan Vasudev Alwala , Armand Joulin , Ishan Misra

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Yuanhan Zhang , Jinming Wu , Wei Li , Bo Li , Zejun Ma , Ziwei Liu , Chunyuan Li

Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture temporal dynamics and…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Chen Shi , Jinrui Xu , Shaoshuai Shi , Kehua Sheng , Bo Zhang , Li Jiang

Despite the significant progress made by deep learning in natural image matting, there has been so far no representative work on deep learning for video matting due to the inherent technical challenges in reasoning temporal domain and lack…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Yanan Sun , Guanzhi Wang , Qiao Gu , Chi-Keung Tang , Yu-Wing Tai

Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos remains limited. Foundation segmentation models such as the SAM…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Hao Wang , Limeng Qiao , Chi Zhang , Lin Ma , Guanglu Wan , Xiangyuan Lan , Xiaodan Liang

Video matting has broad applications, from adding interesting effects to casually captured movies to assisting video production professionals. Matting with associated effects such as shadows and reflections has also attracted increasing…

计算机视觉与模式识别 · 计算机科学 2023-09-15 Geng Lin , Chen Gao , Jia-Bin Huang , Changil Kim , Yipeng Wang , Matthias Zwicker , Ayush Saraf

This paper proposes a novel deep learning-based video object matting method that can achieve temporally coherent matting results. Its key component is an attention-based temporal aggregation module that maximizes image matting networks'…

计算机视觉与模式识别 · 计算机科学 2021-07-30 Yunke Zhang , Chi Wang , Miaomiao Cui , Peiran Ren , Xuansong Xie , Xian-sheng Hua , Hujun Bao , Qixing Huang , Weiwei Xu

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos…

多媒体 · 计算机科学 2024-09-12 Yan-Bo Lin , Yu Tian , Linjie Yang , Gedas Bertasius , Heng Wang

Natural image matting algorithms aim to predict the transparency map (alpha-matte) with the trimap guidance. However, the production of trimap often requires significant labor, which limits the widespread application of matting algorithms…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Jingfeng Yao , Xinggang Wang , Lang Ye , Wenyu Liu