English
Related papers

Related papers: MAGE: Modality-Agnostic Music Generation and Editi…

200 papers

We showcase an unsupervised method that repurposes deep models trained for music generation and music tagging for audio source separation, without any retraining. An audio generation model is conditioned on an input mixture, producing a…

Sound · Computer Science 2021-10-26 Ethan Manilow , Patrick O'Reilly , Prem Seetharaman , Bryan Pardo

We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transformer encoder with…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Hassan Akbari , Dan Kondratyuk , Yin Cui , Rachel Hornung , Huisheng Wang , Hartwig Adam

Generating human portraits is a hot topic in the image generation area, e.g. mask-to-face generation and text-to-face generation. However, these unimodal generation methods lack controllability in image generation. Controllability can be…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Debin Meng , Christos Tzelepis , Ioannis Patras , Georgios Tzimiropoulos

Text-to-image (T2I) generation has been actively studied using Diffusion Models and Autoregressive Models. Recently, Masked Generative Transformers have gained attention as an alternative to Autoregressive Models to overcome the inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Wonjun Kang , Byeongkeun Ahn , Minjae Lee , Kevin Galim , Seunghyuk Oh , Hyung Il Koo , Nam Ik Cho

Generative models in Autonomous Driving (AD) enable diverse scene creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yanhao Wu , Haoyang Zhang , Tianwei Lin , Lichao Huang , Shujie Luo , Rui Wu , Congpei Qiu , Wei Ke , Tong Zhang

Automatically generating symbolic music-music scores tailored to specific human needs-can be highly beneficial for musicians and enthusiasts. Recent studies have shown promising results using extensive datasets and advanced transformer…

Sound · Computer Science 2024-07-08 Yangyang Shu , Haiming Xu , Ziqin Zhou , Anton van den Hengel , Lingqiao Liu

When dealing with the task of fine-grained scene image classification, most previous works lay much emphasis on global visual features when doing multi-modal feature fusion. In other words, models are deliberately designed based on prior…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Yiqun Wang , Zhao Zhou , Xiangcheng Du , Xingjiao Wu , Yingbin Zheng , Cheng Jin

The goal of this work is to enhance balanced multimodal understanding in audio-visual large language models (AV-LLMs) by addressing modality bias without additional training. In current AV-LLMs, audio and video features are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Chaeyoung Jung , Youngjoon Jang , Jongmin Choi , Joon Son Chung

Multimodal sentiment analysis has been studied under the assumption that all modalities are available. However, such a strong assumption does not always hold in practice, and most of multimodal fusion models may fail when partial modalities…

Machine Learning · Computer Science 2022-05-02 Jiandian Zeng , Tianyi Liu , Jiantao Zhou

Automated incident management is critical for microservice reliability. While recent unified frameworks leverage multimodal data for joint optimization, they unrealistically assume perfect data completeness. In practice, network…

Machine Learning · Computer Science 2026-03-30 Wenzhuo Qian , Hailiang Zhao , Ziqi Wang , Zhipeng Gao , Jiayi Chen , Zhiwei Ling , Shuiguang Deng

Current generative models are able to generate high-quality artefacts but have been shown to struggle with compositional reasoning, which can be defined as the ability to generate complex structures from simpler elements. In this paper, we…

Machine Learning · Computer Science 2024-08-20 Giovanni Bindi , Philippe Esling

Multimodal-attributed graphs (MAGs) are a fundamental data structure for multimodal graph learning (MGL), enabling both graph-centric and modality-centric tasks. However, our empirical analysis reveals inherent topology quality limitations…

Machine Learning · Computer Science 2026-03-31 Yinlin Zhu , Xunkai Li , Di Wu , Wang Luo , Miao Hu , Di Wu

We present TALE, a novel training-free framework harnessing the generative capabilities of text-to-image diffusion models to address the cross-domain image composition task that focuses on flawlessly incorporating user-specified objects…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Kien T. Pham , Jingye Chen , Qifeng Chen

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu

Instruction tuning of large vision-language models (LVLMs) increasingly depends on massive multimodal corpora, yet these datasets contain samples with substantial redundancy, low visual dependency, and highly imbalanced coverage of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Shristi Das Biswas , Kaushik Roy

Recent advances in unified multimodal models indicate a clear trend towards comprehensive content generation. However, the auditory domain remains a significant challenge, with music and speech often developed in isolation, hindering…

In extreme scenarios such as nighttime or low-visibility environments, achieving reliable perception is critical for applications like autonomous driving, robotics, and surveillance. Multi-modality image fusion, particularly integrating…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yuchen Guo , Ruoxiang Xu , Rongcheng Li , Weifeng Su

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

This study aims to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video. To achieve this, we propose a novel method that guides single-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Akio Hayakawa , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Emotion alignment between music and palettes is crucial for effective multimedia content, yet misalignment creates confusion that weakens the intended message. However, existing methods often generate only a single dominant color, missing…

Multimedia · Computer Science 2025-09-18 Jiayun Hu , Yueyi He , Tianyi Liang , Changbo Wang , Chenhui Li
‹ Prev 1 8 9 10 Next ›