中文
相关论文

相关论文: Dynin-Omni: Omnimodal Unified Large Diffusion Lang…

200 篇论文

Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet incorporating speech with 3D facial animation remains largely unexplored despite its importance for natural interaction. A key challenge…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Haoyu Zhang , Zhipeng Li , Yiwen Guo , Tianshu Yu

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts. Qwen3-Omni matches the…

Recent advances in generative medical models are constrained by modality-specific scenarios that hinder the integration of complementary evidence from imaging, pathology, and clinical notes. This fragmentation limits their evolution into…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Jiawei Mao , Yuhan Wang , Lifeng Chen , Can Zhao , Yucheng Tang , Dong Yang , Liangqiong Qu , Daguang Xu , Yuyin Zhou

Magnetic Resonance Imaging (MRI) is indispensable in clinical practice but remains constrained by fragmented, multi-stage workflows encompassing acquisition, reconstruction, segmentation, detection, diagnosis, and reporting. While deep…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Xingxin He , Aurora Rofena , Ruimin Feng , Haozhe Liao , Zhaoye Zhou , Albert Jang , Fang Liu

We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by utilizing a fully discrete diffusion modeling to handle inputs…

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Xu Guo , Fulong Ye , Qichao Sun , Liyang Chen , Bingchuan Li , Pengze Zhang , Jiawei Liu , Songtao Zhao , Qian He , Xiangwang Hou

The recently developed discrete diffusion models perform extraordinarily well in the text-to-image task, showing significant promise for handling the multi-modality signals. In this work, we harness these traits and present a unified…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Minghui Hu , Chuanxia Zheng , Heliang Zheng , Tat-Jen Cham , Chaoyue Wang , Zuopeng Yang , Dacheng Tao , Ponnuthurai N. Suganthan

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study…

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations, or unrealistic…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Tianxiao Li , Zhenglin Huang , Haiquan Wen , Yiwei He , Xinze Li , Bingyu Zhu , Wuhui Duan , Congang Chen , Zeyu Fu , Yi Dong , Baoyuan Wu , Jason Li , Guangliang Cheng

We introduce MarDini, a new family of video diffusion models that integrate the advantages of masked auto-regression (MAR) into a unified diffusion model (DM) framework. Here, MAR handles temporal planning, while DM focuses on spatial…

We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its architecture combines a fully semantic discrete tokenizer, a…

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming…

In this study, we present a multimodal framework for predicting neuro-facial disorders by capturing both vocal and facial cues. We hypothesize that explicitly disentangling shared and modality-specific representations within multimodal…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Mohd Mujtaba Akhtar , Girish , Muskaan Singh

Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the lack of multilingual speech data, unified multimodal architectures, and comprehensive evaluation benchmarks. To address these…

计算与语言 · 计算机科学 2025-12-12 Bo Yang , Lanfei Feng , Yunkui Chen , Yu Zhang , Jianyu Zhang , Xiao Xu , Nueraili Aierken , Shijian Li

We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks and low-resolution…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Shufan Li , Jiuxiang Gu , Kangning Liu , Zhe Lin , Zijun Wei , Aditya Grover , Jason Kuen

With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited exploration of unified generative architectures. In addition, many…

多媒体 · 计算机科学 2025-07-08 Lei Zhao , Linfeng Feng , Dongxu Ge , Rujin Chen , Fangqiu Yi , Chi Zhang , Xiao-Lei Zhang , Xuelong Li

Multimodal Large Languages models have been progressing from uni-modal understanding toward unifying visual, audio and language modalities, collectively termed omni models. However, the correlation between uni-modal and omni-modal remains…

计算与语言 · 计算机科学 2025-10-31 Chen Chen , ZeYang Hu , Fengjiao Chen , Liya Ma , Jiaxing Liu , Xiaoyu Li , Ziwen Wang , Xuezhi Cao , Xunliang Cai

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective…

计算与语言 · 计算机科学 2026-01-21 Qian Chen , Jinlan Fu , Changsong Li , See-Kiong Ng , Xipeng Qiu

Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities. However, the capacity of these models to concurrently process and reason about multiple modalities remains…