English
Related papers

Related papers: MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-…

200 papers

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts. Qwen3-Omni matches the…

Multimodal Large Language Models (MLLMs) have achieved strong performance across vision-language tasks, but suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Yingqi Fan , Anhao Zhao , Jinlan Fu , Junlong Tong , Hui Su , Yijie Pan , Wei Zhang , Xiaoyu Shen

As AR/VR technologies become integral to daily life, there's a growing need for AI that understands human social dynamics from an egocentric perspective. However, current LLMs often lack the social awareness to discern when to intervene as…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Xijun Wang , Tanay Sharma , Achin Kulshrestha , Abhimitra Meka , Aveek Purohit , Dinesh Manocha

In this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task designed to produce synchronized verbal and non-verbal listener feedback online, based on the speaker's multimodal inputs. OMCRG captures…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Cheng Luo , Jianghui Wang , Bing Li , Siyang Song , Bernard Ghanem

Spoken dialogue modeling poses challenges beyond text-based language modeling, requiring real-time interaction, turn-taking, and backchanneling. While most Spoken Dialogue Models (SDMs) operate in half-duplex mode-processing one turn at a…

Computation and Language · Computer Science 2025-08-19 Guan-Ting Lin , Jiachen Lian , Tingle Li , Qirui Wang , Gopala Anumanchipalli , Alexander H. Liu , Hung-yi Lee

Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users. While emerging benchmarks address proactivity, they are largely confined to alert…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Dongchuan Ran , Linyu Ou , Xueheng Li , Wenwen Tong , Chenxu Guo , Hewei Guo , Kaibing Wang , Lewei Lu

We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining paradigm, named Multimodal Context (MiCo), which can scale up…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Yiyuan Zhang , Handong Li , Jing Liu , Xiangyu Yue

Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures…

Computation and Language · Computer Science 2025-08-08 Qianli Ma , Yaowei Zheng , Zhelun Shi , Zhongkai Zhao , Bin Jia , Ziyue Huang , Zhiqi Lin , Youjie Li , Jiacheng Yang , Yanghua Peng , Zhi Zhang , Xin Liu

Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among…

Computation and Language · Computer Science 2025-06-06 Haonan Zhang , Run Luo , Xiong Liu , Yuchuan Wu , Ting-En Lin , Pengpeng Zeng , Qiang Qu , Feiteng Fang , Min Yang , Lianli Gao , Jingkuan Song , Fei Huang , Yongbin Li

As multimodal large language models (MLLMs) advance, their large-scale architectures pose challenges for deployment in resource-constrained environments. In the age of large models, where energy efficiency, computational scalability and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Sike Xiang , Shuang Chen , Amir Atapour-Abarghouei

Multimodal Large Language Models (MLLMs) have recently made rapid progress toward unified Omni models that integrate vision, language, and audio. However, existing environments largely focus on 2D or 3D visual context and vision-language…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yurui Dong , Ziyue Wang , Shuyun Lu , Dairu Liu , Xuechen Liu , Fuwen Luo , Peng Li , Yang Liu

In this paper, we introduce LiveMind, a novel low-latency inference framework for large language model (LLM) inference which enables LLMs to perform inferences with incomplete user input. By reallocating computational processes to the input…

Artificial Intelligence · Computer Science 2024-11-07 Chuangtao Chen , Grace Li Zhang , Xunzhao Yin , Cheng Zhuo , Ulf Schlichtmann , Bing Li

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…

In this paper, we introduce a new problem, Online-MMSI, where the model must perform multimodal social interaction understanding (MMSI) using only historical information. Given a recorded video and a multi-party dialogue, the AI assistant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xinpeng Li , Shijian Deng , Bolin Lai , Weiguo Pian , James M. Rehg , Yapeng Tian

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipelines, high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Zehan Wang , Ziang Zhang , Hang Zhang , Luping Liu , Rongjie Huang , Xize Cheng , Hengshuang Zhao , Zhou Zhao

Large language models (LLMs) have demonstrated remarkable performance on single-turn text-to-SQL tasks, but real-world database applications predominantly require multi-turn interactions to handle ambiguous queries, execution errors, and…

Computational resource constraints on edge devices make it difficult to develop a fully embedded AI companion system with a satisfactory user experience. AI companion and memory systems detailed in existing literature cannot be directly…

Artificial Intelligence · Computer Science 2026-01-14 Rahul Gupta , Stephen D. H. Hsu

In an era defined by the explosive growth of data and rapid technological advancements, Multimodal Large Language Models (MLLMs) stand at the forefront of artificial intelligence (AI) systems. Designed to seamlessly integrate diverse data…

Large Language Models (LLMs) have demonstrated remarkable prowess in generating contextually coherent responses, yet their fixed context windows pose fundamental challenges for maintaining consistency over prolonged multi-session dialogues.…

Computation and Language · Computer Science 2025-04-29 Prateek Chhikara , Dev Khant , Saket Aryan , Taranjeet Singh , Deshraj Yadav

Pre-trained conversation models (PCMs) have demonstrated remarkable results in task-oriented dialogue (TOD) systems. Many PCMs focus predominantly on dialogue management tasks like dialogue state tracking, dialogue generation tasks like…

Computation and Language · Computer Science 2023-12-29 Mingtao Yang , See-Kiong Ng , Jinlan Fu