中文
相关论文

相关论文: MoFu: Scale-Aware Modulation and Fourier Fusion fo…

200 篇论文

Diffusion models are proficient at generating high-quality images. They are however effective only when operating at the resolution used during training. Inference at a scaled resolution leads to repetitive patterns and structural…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Haosen Yang , Adrian Bulat , Isma Hadji , Hai X. Pham , Xiatian Zhu , Georgios Tzimiropoulos , Brais Martinez

Large Multimodal Models (LMMs) are powerful tools that are capable of reasoning and understanding multimodal information beyond text and language. Despite their entrenched impact, the development of LMMs is hindered by the higher…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Vittorio Pippi , Matthieu Guillaumin , Silvia Cascianelli , Rita Cucchiara , Maximilian Jaritz , Loris Bazzani

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Lijie Liu , Tianxiang Ma , Bingchuan Li , Zhuowei Chen , Jiawei Liu , Gen Li , Siyu Zhou , Qian He , Xinglong Wu

Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. While recent advances allow these models to handle varying input sizes during training,…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Bocheng Zou , Mu Cai , Mark Stanley , Dingfu Lu , Yong Jae Lee

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusion induces modality…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Bonan Ding , Umair Nawaz , Ufaq Khan , Abdelrahman M. Shaker , Muhammad Haris Khan , Jiale Cao , Jin Xie , Fahad Shahbaz Khan

Multimodal molecular models often suffer from 3D conformer unreliability and modality collapse, limiting their robustness and generalization. We propose MuMo, a structured multimodal fusion framework that addresses these challenges in…

机器学习 · 计算机科学 2025-10-29 Zihao Jing , Yan Sun , Yan Yi Li , Sugitha Janarthanan , Alana Deng , Pingzhao Hu

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Zhaoyang Li , Dongjun Qian , Kai Su , Qishuai Diao , Xiangyang Xia , Chang Liu , Wenfei Yang , Tianzhu Zhang , Zehuan Yuan

Consecutive frames in a video contain redundancy, but they may also contain relevant complementary information for the detection task. The objective of our work is to leverage this complementary information to improve detection. Therefore,…

计算机视觉与模式识别 · 计算机科学 2024-02-19 Noreen Anwar , Guillaume-Alexandre Bilodeau , Wassim Bouachir

Most existing learning-based multi-modality image fusion (MMIF) methods suffer from significant structure inconsistency due to their inappropriate usage of structural features at the semantic level. To alleviate these issues, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Qiao Yang , Yu Zhang , Yutong Chen , Jian Zhang , Shunli Zhang

Image-based geometric modeling and novel view synthesis based on sparse, large-baseline samplings are challenging but important tasks for emerging multimedia applications such as virtual reality and immersive telepresence. Existing methods…

计算机视觉与模式识别 · 计算机科学 2022-10-05 Wenpeng Xing , Jie Chen , Zaifeng Yang , Qiang Wang

Recent advancements in personalizing text-to-image (T2I) diffusion models have shown the capability to generate images based on personalized visual concepts using a limited number of user-provided examples. However, these models often…

计算机视觉与模式识别 · 计算机科学 2024-02-20 Yan Hong , Jianfu Zhang

Multimodal learning faces a fundamental tension between deep, fine-grained fusion and computational scalability. While cross-attention models achieve strong performance through exhaustive pairwise fusion, their quadratic complexity is…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yusuf Shihata

Multi-modal MRI offers valuable complementary information for diagnosis and treatment; however, its utility is limited by prolonged scanning times. To accelerate the acquisition process, a practical approach is to reconstruct images of the…

图像与视频处理 · 电气工程与系统科学 2024-07-09 Jing Zou , Lanqing Liu , Qi Chen , Shujun Wang , Zhanli Hu , Xiaohan Xing , Jing Qin

Large-scale pre-training has brought unimodal fields such as computer vision and natural language processing to a new era. Following this trend, the size of multi-modal learning models constantly increases, leading to an urgent need to…

计算机视觉与模式识别 · 计算机科学 2023-05-16 Yaowei Li , Ruijie Quan , Linchao Zhu , Yi Yang

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Yufeng Cheng , Wenxu Wu , Shaojin Wu , Mengqi Huang , Fei Ding , Qian He

Visual SLAM is particularly challenging in environments affected by noise, varying lighting conditions, and darkness. Learning-based optical flow algorithms can leverage multiple modalities to address these challenges, but traditional…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Youjie Zhou , Guofeng Mei , Yiming Wang , Yi Wan , Fabio Poiesi

Video semantic segmentation aims to generate accurate semantic maps for each video frame. To this end, many works dedicate to integrate diverse information from consecutive frames to enhance the features for prediction, where a feature…

计算机视觉与模式识别 · 计算机科学 2023-01-11 Jiafan Zhuang , Zilei Wang , Junjie Li

Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Delong Liu , Zhaohui Hou , Mingjie Zhan , Shihao Han , Zhicheng Zhao , Fei Su

Current multi-modal image fusion methods typically rely on task-specific models, leading to high training costs and limited scalability. While generative methods provide a unified modeling perspective, they often suffer from slow inference…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Huayi Zhu , Xiu Shu , Youqiang Xiong , Qiao Liu , Rui Chen , Di Yuan , Xiaojun Chang , Zhenyu He

Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine-grained control remains a key bottleneck that limits practical customizability for content creation. For AI video creators, three forms of…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhenghong Zhou , Xiaohang Zhan , Zhiqin Chen , Soo Ye Kim , Nanxuan Zhao , Haitian Zheng , Qing Liu , He Zhang , Zhe Lin , Yuqian Zhou , Jiebo Luo
‹ 上一页 1 2 3 10 下一页 ›