中文
相关论文

相关论文: UniMMVSR: A Unified Multi-Modal Framework for Casc…

200 篇论文

Diffusion models have recently advanced video restoration, but applying them to real-world video super-resolution (VSR) remains challenging due to high latency, prohibitive computation, and poor generalization to ultra-high resolutions. Our…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Junhao Zhuang , Shi Guo , Xin Cai , Xiaohui Li , Yihao Liu , Chun Yuan , Tianfan Xue

This paper tackles the challenge of robust reconstruction, i.e., the task of reconstructing a 3D scene from a set of inconsistent multi-view images. Some recent works have attempted to simultaneously remove image inconsistencies and perform…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Jin Cao , Hongrui Wu , Ziyong Feng , Hujun Bao , Xiaowei Zhou , Sida Peng

Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existing codebook-based…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Yanzhe Chen , Huasong Zhong , Yan Li , Zhenheng Yang

With the rapid advancement in video generation, people can conveniently use video generation models to create videos tailored to their specific desires. As a result, there are also growing concerns about the potential misuse of video…

密码学与安全 · 计算机科学 2025-07-08 Yan Pang , Baicheng Chen , Yang Zhang , Tianhao Wang

Composed image retrieval, multi-turn composed image retrieval, and composed video retrieval all share a common paradigm: composing the reference visual with modification text to retrieve the desired target. Despite this shared structure,…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Haokun Wen , Xuemeng Song , Haoyu Zhang , Xiangyu Zhao , Weili Guan , Liqiang Nie

Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Zizhong Li , Haopeng Zhang , Jiawei Zhang

Video super-resolution (VSR) is a critical task for enhancing low-bitrate and low-resolution videos, particularly in streaming applications. While numerous solutions have been developed, they often suffer from high computational demands,…

图像与视频处理 · 电气工程与系统科学 2024-09-27 Marcos V Conde , Zhijun Lei , Wen Li , Christos Bampis , Ioannis Katsavounidis , Radu Timofte

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

High-resolution (HR) medical videos are vital for accurate diagnosis, yet are hard to acquire due to hardware limitations and physiological constraints. Clinically, the collected low-resolution (LR) medical videos present unique challenges…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Xinyu Liu , Guolei Sun , Cheng Wang , Yixuan Yuan , Ender Konukoglu

Multimodal retrieval-augmented Generation (MM-RAG) is a key approach for applying large language models (LLMs) and agents to real-world knowledge bases, yet current evaluations are fragmented -- focusing on either text or images in…

计算与语言 · 计算机科学 2026-01-06 Xiangyu Peng , Can Qin , Zeyuan Chen , Ran Xu , Caiming Xiong , Chien-Sheng Wu

Unified Multimodal Models (UMMs) built on shared autoregressive (AR) transformers are attractive for their architectural simplicity. However, we identify a critical limitation: when trained on multimodal inputs, modality-shared transformers…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Jitai Hao , Hao Liu , Xinyan Xiao , Qiang Huang , Jun Yu

Videos can be created by first outlining a global view of the scene and then adding local details. Inspired by this idea we propose a cascaded model for video generation which follows a coarse to fine approach. First our model generates a…

计算机视觉与模式识别 · 计算机科学 2022-06-03 Lluis Castrejon , Nicolas Ballas , Aaron Courville

We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models. Given a text prompt, Imagen Video generates high definition videos using a base video generation model and a sequence of…

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, limiting their…

In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventional two-stage…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Woongjib Choi , Sangmin Lee , Hyungseob Lim , Hong-Goo Kang

The video generation field has witnessed rapid improvements with the introduction of recent diffusion models. While these models have successfully enhanced appearance quality, they still face challenges in generating coherent and natural…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Yaosi Hu , Zhenzhong Chen , Chong Luo

This paper presents OmniDataComposer, an innovative approach for multimodal data fusion and unlimited data generation with an intent to refine and uncomplicate interplay among diverse data modalities. Coming to the core breakthrough, it…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Dongyang Yu , Shihao Wang , Yuan Fang , Wangpeng An

We present UniRef-Image-Edit, a high-performance multi-modal generation system that unifies single-image editing and multi-image composition within a single framework. Existing diffusion-based editing methods often struggle to maintain…

Perceptual studies demonstrate that conditional diffusion models excel at reconstructing video content aligned with human visual perception. Building on this insight, we propose a video compression framework that leverages conditional…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Fangqiu Yi , Jingyu Xu , Jiawei Shao , Chi Zhang , Xuelong Li

We introduce a novel diffusion-based video generation method, generating a video showing multiple events given multiple individual sentences from the user. Our method does not require a large-scale video dataset since our method uses a…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Gyeongrok Oh , Jaehwan Jeong , Sieun Kim , Wonmin Byeon , Jinkyu Kim , Sungwoong Kim , Sangpil Kim