English
Related papers

Related papers: U-Mind: A Unified Framework for Real-Time Multimod…

200 papers

While Unified Multimodal Models (UMMs) have achieved remarkable success in cross-modal comprehension, a significant gap persists in their ability to leverage such internal knowledge for high-quality generation. We formalize this discrepancy…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Ruiyan Han , Zhen Fang , XinYu Sun , Yuchen Ma , Ziheng Wang , Yu Zeng , Zehui Chen , Lin Chen , Wenxuan Huang , Wei-Jie Xu , Yi Cao , Feng Zhao

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Tao Feng , Yifan Xie , Xun Guan , Jiyuan Song , Zhou Liu , Fei Ma , Fei Yu

Generating realistic, context-aware two-person motion conditioned on diverse modalities remains a fundamental challenge for graphics, animation and embodied AI systems. Real-world applications such as VR/AR companions, social robotics and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Prerit Gupta , Shourya Verma , Ananth Grama , Aniket Bera

We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene-within a single, task-agnostic architecture. In contrast to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Sheng Liu , Yuanzhi Liang , Jiepeng Wang , Sidan Du , Chi Zhang , Xuelong Li

Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works largely rely on implicit token-level alignment via supervised fine-tuning, which fails to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zijun Shen , Sihan Yang , Ruichuan An , Ziyu Guo , Hao Liang , Ming Lu , Renrui Zhang , Wentao Zhang

Large language model (LLM)-based agents have been successfully deployed in many tool-augmented settings, but their scalability is fundamentally constrained by context length. Existing context-folding methods mitigate this issue by…

Computation and Language · Computer Science 2026-01-27 Jin Su , Runnan Fang , Yeqiu Li , Xiaobin Wang , Shihao Cai , Pengjun Xie , Ningyu Zhang , Fajie Yuan

The goal-oriented document-grounded dialogue aims at responding to the user query based on the dialogue context and supporting document. Existing studies tackle this problem by decomposing it into two sub-tasks: knowledge identification and…

Computation and Language · Computer Science 2022-04-19 Chang Gao , Wenxuan Zhang , Wai Lam

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised…

Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nuanced emotional expressions due to the lack of fine-grained emotion control. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Jiayi Lyu , Leigang Qu , Wenjing Zhang , Hanyu Jiang , Kai Liu , Zhenglin Zhou , Xiaobo Xia , Jian Xue , Tat-Seng Chua

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yu Zhang , Kaiyuan Shen , Yang Li

Unified Multimodal Large Language Models (U-MLLMs) integrate understanding and generation within a single architecture. However, existing evaluations typically assess these capabilities separately, overlooking semantic equivalence, i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Hongbo Jiang , Jie Li , Yunhang Shen , Pingyang Dai , Xing Sun , Haoyu Cao , Liujuan Cao

The recently developed discrete diffusion models perform extraordinarily well in the text-to-image task, showing significant promise for handling the multi-modality signals. In this work, we harness these traits and present a unified…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Minghui Hu , Chuanxia Zheng , Heliang Zheng , Tat-Jen Cham , Chaoyue Wang , Zuopeng Yang , Dacheng Tao , Ponnuthurai N. Suganthan

Multimodal large language models (MLLMs) have shown strong capabilities but remain limited to fixed modality pairs and require costly fine-tuning with large aligned datasets. Building fully omni-capable models that can integrate text,…

Artificial Intelligence · Computer Science 2025-11-06 Huawei Lin , Yunzhi Shi , Tong Geng , Weijie Zhao , Wei Wang , Ravender Pal Singh

Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descriptions,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Junchao Yi , Rui Zhao , Jiahao Tang , Weixian Lei , Linjie Li , Qisheng Su , Zhengyuan Yang , Lijuan Wang , Xiaofeng Zhu , Alex Jinpeng Wang

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jiadong Pan , Liang Li , Yuxin Peng , Yu-Ming Tang , Shuohuan Wang , Yu Sun , Hua Wu , Qingming Huang , Haifeng Wang

Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come…

Sound · Computer Science 2025-01-31 Joanna Hong , Sanjeel Parekh , Honglie Chen , Jacob Donley , Ke Tan , Buye Xu , Anurag Kumar

Decoding visual information from time-resolved brain recordings, such as EEG and MEG, plays a pivotal role in real-time brain-computer interfaces. However, existing approaches primarily focus on direct brain-image feature alignment and are…

Human-Computer Interaction · Computer Science 2025-11-12 Chengjian Xu , Yonghao Song , Zelin Liao , Haochuan Zhang , Qiong Wang , Qingqing Zheng

We propose a unified Implicit Dialog framework for goal-oriented, information seeking tasks of Conversational Search applications. It aims to enable dialog interactions with domain data without replying on explicitly encoded the rules but…

Computation and Language · Computer Science 2018-02-14 Song Feng , R. Chulaka Gunasekara , Sunil Shashidhara , Kshitij P. Fadnis , Lazaros C. Polymenakos

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging to design a unified network for processing various…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Yiyuan Zhang , Kaixiong Gong , Kaipeng Zhang , Hongsheng Li , Yu Qiao , Wanli Ouyang , Xiangyu Yue