English
Related papers

Related papers: UniForm: A Unified Multi-Task Diffusion Transforme…

200 papers

When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing, segmentation and question answering are mostly explored…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Guangyao Li , Xin Wang , Wenwu Zhu

Removing various degradations from damaged documents greatly benefits digitization, downstream document analysis, and readability. Previous methods often treat each restoration task independently with dedicated models, leading to a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Fangmin Zhao , Weichao Zeng , Zhenhang Li , Dongbao Yang , Binbin Li , Xiaojun Bi , Yu Zhou

We introduce UniCon, a novel architecture designed to enhance control and efficiency in training adapters for large-scale diffusion models. Unlike existing methods that rely on bidirectional interaction between the diffusion model and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Fanghua Yu , Jinjin Gu , Jinfan Hu , Zheyuan Li , Chao Dong

Diffusion models, though originally designed for generative tasks, have demonstrated impressive self-supervised representation learning capabilities. A particularly intriguing phenomenon in these models is the emergence of unimodal…

Machine Learning · Computer Science 2026-02-04 Xiao Li , Zekai Zhang , Xiang Li , Siyi Chen , Zhihui Zhu , Peng Wang , Qing Qu

Diffusion-based generative models have exhibited powerful generative performance in recent years. However, as many attributes exist in the data distribution and owing to several limitations of sharing the model parameters across all levels…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-26 Ha-Yeong Choi , Sang-Hoon Lee , Seong-Whan Lee

Unified multimodal models have recently attracted considerable attention for their remarkable abilities in jointly understanding and generating diverse content. However, as contexts integrate increasingly numerous interleaved multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Yanzuo Lu , Xin Xia , Manlin Zhang , Huafeng Kuang , Jianbin Zheng , Yuxi Ren , Xuefeng Xiao

Video Diffusion Models have been developed for video generation, usually integrating text and image conditioning to enhance control over the generated content. Despite the progress, ensuring consistency across frames remains a challenge,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Tian Xia , Xuweiyi Chen , Sihan Xu

With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text related downstream tasks. However, most of the existing…

Computer Vision and Pattern Recognition · Computer Science 2020-09-16 Huaishao Luo , Lei Ji , Botian Shi , Haoyang Huang , Nan Duan , Tianrui Li , Jason Li , Taroon Bharti , Ming Zhou

Generating realistic motions for digital humans is a core but challenging part of computer animations and games, as human motions are both diverse in content and rich in styles. While the latest deep learning approaches have made…

Computer Vision and Pattern Recognition · Computer Science 2022-12-19 Ziyi Chang , Edmund J. C. Findlay , Haozheng Zhang , Hubert P. H. Shum

We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. Ming-Omni employs dedicated encoders to extract tokens from…

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical…

Sound · Computer Science 2026-04-07 Weiguo Pian , Saksham Singh Kushwaha , Zhimin Chen , Shijian Deng , Kai Wang , Yunhui Guo , Yapeng Tian

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress in the domain of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zeyinzi Jiang , Zhen Han , Chaojie Mao , Jingfeng Zhang , Yulin Pan , Yu Liu

Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Junwen Xiong , Peng Zhang , Tao You , Chuanyue Li , Wei Huang , Yufei Zha

There is a recent trend in the LiDAR perception field towards unifying multiple tasks in a single strong network with improved performance, as opposed to using separate networks for each task. In this paper, we introduce a new LiDAR…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Zixiang Zhou , Dongqiangzi Ye , Weijia Chen , Yufei Xie , Yu Wang , Panqu Wang , Hassan Foroosh

Although recent advances in visual generation have been remarkable, most existing architectures still depend on distinct encoders for images and text. This separation constrains diffusion models' ability to perform cross-modal reasoning and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Kevin Li , Manuel Brack , Sudeep Katakol , Hareesh Ravi , Ajinkya Kale

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a…

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Wei Xu , Chunsheng Shi , Sifan Tu , Xin Zhou , Dingkang Liang , Xiang Bai

Text-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. However, despite their success, text descriptions often…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Shihao Zhao , Dongdong Chen , Yen-Chun Chen , Jianmin Bao , Shaozhe Hao , Lu Yuan , Kwan-Yee K. Wong

Diffusion models have achieved remarkable success across a range of generative tasks. Recent efforts to enhance diffusion model architectures have reimagined them as a form of multi-task learning, where each task corresponds to a denoising…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Byeongjun Park , Hyojun Go , Jin-Young Kim , Sangmin Woo , Seokil Ham , Changick Kim

Biological intelligence systems of animals perceive the world by integrating information in different modalities and processing simultaneously for various tasks. In contrast, current machine learning research follows a task-specific…

Computer Vision and Pattern Recognition · Computer Science 2021-12-03 Xizhou Zhu , Jinguo Zhu , Hao Li , Xiaoshi Wu , Xiaogang Wang , Hongsheng Li , Xiaohua Wang , Jifeng Dai