中文
相关论文

相关论文: Emu: Generative Pretraining in Multimodality

200 篇论文

With the rapid penetration of artificial intelligence across industries and scenarios, a key challenge in building the next-generation intelligent core lies in effectively integrating the language understanding capabilities of foundation…

人工智能 · 计算机科学 2025-06-03 Liang Geng

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus on natural language tasks, GEM is a large-scale…

计算与语言 · 计算机科学 2021-06-21 Lin Su , Nan Duan , Edward Cui , Lei Ji , Chenfei Wu , Huaishao Luo , Yongfei Liu , Ming Zhong , Taroon Bharti , Arun Sacheti

Conventional methods for the image-text generation tasks mainly tackle the naturally bidirectional generation tasks separately, focusing on designing task-specific frameworks to improve the quality and fidelity of the generated samples.…

计算机视觉与模式识别 · 计算机科学 2022-01-03 Han Zhang , Weichong Yin , Yewei Fang , Lanxin Li , Boqiang Duan , Zhihua Wu , Yu Sun , Hao Tian , Hua Wu , Haifeng Wang

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent…

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, limiting their…

Real-world perception and interaction are inherently multimodal, encompassing not only language but also vision and speech, which motivates the development of "Omni" MLLMs that support both multimodal inputs and multimodal outputs. While a…

机器学习 · 计算机科学 2026-01-27 Dongjie Cheng , Ruifeng Yuan , Yongqi Li , Runyang You , Wenjie Wang , Liqiang Nie , Lei Zhang , Wenjie Li

Recent advancements in image generation have enabled the creation of high-quality images from text conditions. However, when facing multi-modal conditions, such as text combined with reference appearances, existing methods struggle to…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Yucheng Han , Rui Wang , Chi Zhang , Juntao Hu , Pei Cheng , Bin Fu , Hanwang Zhang

Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to characterize and study…

In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, image editing, and in-context generation. Unlike OmniGen v1,…

Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the two tasks and preventing fully end-to-end optimization from…

This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the…

计算与语言 · 计算机科学 2024-10-10 Soyeon Caren Han , Feiqi Cao , Josiah Poon , Roberto Navigli

We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image, video, audio, text, depth, point cloud, time series,…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Siddharth Srivastava , Gaurav Sharma

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Jiepeng Wang , Zhaoqing Wang , Hao Pan , Yuan Liu , Dongdong Yu , Changhu Wang , Wenping Wang

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired…

声音 · 计算机科学 2024-10-08 Han Yang , Kun Su , Yutong Zhang , Jiaben Chen , Kaizhi Qian , Gaowen Liu , Chuang Gan

Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This imbalance motivates…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Luozheng Qin , Jia Gong , Qian Qiao , Tianjiao Li , Li Xu , Haoyu Pan , Chao Qu , Zhiyu Tan , Hao Li

Using a natural language sentence to describe the content of an image is a challenging but very important task. It is challenging because a description must not only capture objects contained in the image and the relationships among them,…

计算机视觉与模式识别 · 计算机科学 2019-04-23 Xuelong Li , Aihong Yuan , Xiaoqiang Lu

Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Wei Li , Xue Xu , Jiachen Liu , Xinyan Xiao

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging to design a unified network for processing various…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Yiyuan Zhang , Kaixiong Gong , Kaipeng Zhang , Hongsheng Li , Yu Qiao , Wanli Ouyang , Xiangyu Yue

Unified multimodal models hold the promise of generating extensive, interleaved narratives, weaving text and imagery into coherent long-form stories. However, current systems suffer from a critical reliability gap: as sequences grow,…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Haoyu Chen , Qing Liu , Yuqian Zhou , He Zhang , Zhaowen Wang , Mengwei Ren , Jingjing Ren , Xiang Wang , Zhe Lin , Lei Zhu