English
Related papers

Related papers: Stream-Omni: Simultaneous Multimodal Interactions …

200 papers

Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring synergistic omni-modal reasoning. Beyond understanding global…

Computation and Language · Computer Science 2026-04-08 Hongcheng Liu , Yuhao Wang , Zhe Chen , Pingjie Wang , Zhiyuan Zhu , Yixuan Hou , Yanfeng Wang , Yu Wang

This research introduces a transformative framework for integrating Vision-Enhanced Large Language Models (LLMs) with advanced transformer-based architectures to tackle challenges in high-resolution image synthesis and multimodal data…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Karthikeya KV

This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limited tri-modal datasets, high computational costs, and complex…

Large Language Models (LLMs) have showcased impressive capabilities in text comprehension and generation, prompting research efforts towards video LLMs to facilitate human-AI interaction at the video level. However, how to effectively…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Ruyang Liu , Chen Li , Haoran Tang , Yixiao Ge , Ying Shan , Ge Li

The advent of Large Language Models (LLMs) has significantly reshaped the trajectory of the AI revolution. Nevertheless, these LLMs exhibit a notable limitation, as they are primarily adept at processing textual information. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Akash Ghosh , Arkadeep Acharya , Sriparna Saha , Vinija Jain , Aman Chadha

Real-world perception and interaction are inherently multimodal, encompassing not only language but also vision and speech, which motivates the development of "Omni" MLLMs that support both multimodal inputs and multimodal outputs. While a…

Machine Learning · Computer Science 2026-01-27 Dongjie Cheng , Ruifeng Yuan , Yongqi Li , Runyang You , Wenjie Wang , Liqiang Nie , Lei Zhang , Wenjie Li

Large Language Models (LLMs) have demonstrated extraordinary performance across a broad array of applications, from traditional language processing tasks to interpreting structured sequences like time-series data. Yet, their effectiveness…

Databases · Computer Science 2023-07-18 Shuhao Zhang , Xianzhi Zeng , Yuhao Wu , Zhonghao Yang

While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Lijiang Li , Zuwei Long , Yunhang Shen , Heting Gao , Haoyu Cao , Xing Sun , Caifeng Shan , Ran He , Chaoyou Fu

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strategy that transitions…

Multimedia · Computer Science 2025-12-01 Meituan LongCat Team , Bairui Wang , Bayan , Bin Xiao , Bo Zhang , Bolin Rong , Borun Chen , Chang Wan , Chao Zhang , Chen Huang , Chen Chen , Chen Chen , Chengxu Yang , Chengzuo Yang , Cong Han , Dandan Peng , Delian Ruan , Detai Xin , Disong Wang , Dongchao Yang , Fanfan Liu , Fengjiao Chen , Fengyu Yang , Gan Dong , Gang Huang , Gang Xu , Guanglu Wan , Guoqiang Tan , Guoqiao Yu , Haibo Qiu , Hao Lu , Hongbo Liu , Hongyu Xiang , Jiaheng Wu , Jian Yang , Jiaxing Liu , Jing Huang , Jingang Wang , Jinrui Ding , Juchao Jiang , Jun Kuang , Jun Wang , Junhui Mei , Ke Ding , Kefeng Zhang , Lei Chen , Liang Shi , Limeng Qiao , Liming Zheng , Lin Ma , Liuyang Guo , Liya Ma , Luying Sun , Man Gao , Mengshen Zhu , Miao Cao , Minliang Lin , Nuo Xu , Peng Shi , Qi Zhang , Qian Fang , Qian Wang , Qian Yang , Quanxiu Wang , Rongxiang Weng , Rongxin Guo , Ruoxuan Liang , Senbin Yang , Shanbo Xu , Shanglin Lei , Shengze Ye , Shimin Chen , Shuaiqi Chen , Shujie Hu , Shuo Li , Siqi Yang , Siyu Xu , Siyu Ren , Song Li , Songxiang Liu , Tianhao Bai , Tianye Dai , Wei Hong , Wei Wang , Weixiao Zhao , Wengang Cao , Wenlong Zhu , Wenlong He , Xi Su , Xi Nan , Xiaohan Zhao , Xiaohao Wang , Xiaoyu Zhao , Xiaoyu Wang , Xiaoyu Li , Xin Pan , Xin Chen , Xiusong Sun , Xu Xiang , Xudong Xing , Xuezhi Cao , Xunliang Cai , Yang Yang , Yanli Tan , Yao Yao , Yerui Sun , Yi Chen , Yifan Lu , Yin Gong , Yining Zhang , Yitian Chen , Yiyang Gan , Yuchen Tang , Yuchen Xie , Yueqian Wang , Yuewen Zheng , Yufei Zhang , Yufeng Zhong , Yulei Qian , Yuqi Peng , Yuqian Li , Yuwei Jiang , Zeyang Hu , Zheng Zhang , Zhengkun Tian , Zhiqing Hong , Zhixiong Zeng , Zhuqi Mi , Ziran Li , Ziwen Wang , Ziyi Zhao , Ziyuan Zhuang , Zizhe Zhao

Next-generation multimodal foundation models capable of any-to-any cross-modal generation and multi-turn interaction will serve as core components of artificial general intelligence systems, playing a pivotal role in human-machine…

Computation and Language · Computer Science 2025-10-17 Run Luo , Xiaobo Xia , Lu Wang , Longze Chen , Renke Shan , Jing Luo , Min Yang , Tat-Seng Chua

Recent Large Language Models have been enhanced with vision capabilities, enabling them to comprehend images, videos, and interleaved vision-language content. However, the learning methods of these large multimodal models typically treat…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Joya Chen , Zhaoyang Lv , Shiwei Wu , Kevin Qinghong Lin , Chenan Song , Difei Gao , Jia-Wei Liu , Ziteng Gao , Dongxing Mao , Mike Zheng Shou

Any-to-any multimodal models that jointly handle text, images, video, and audio represent a significant advance in multimodal AI. However, their complex architectures (typically combining multiple autoregressive LLMs, diffusion…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-02-03 Peiqi Yin , Jiangyun Zhu , Han Gao , Chenguang Zheng , Yongxiang Huang , Taichang Zhou , Ruirui Yang , Weizhi Liu , Weiqing Chen , Canlin Guo , Didan Deng , Zifeng Mo , Cong Wang , James Cheng , Roger Wang , Hongsheng Liu

Multi-Object Tracking (MOT) is evolving from geometric localization to Semantic MOT (SMOT) to answer complex relational queries, yet progress is hindered by semantic data scarcity and a structural disconnect between tracking architectures…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Pan Liao , Feng Yang , Di Wu , Jinwen Yu , Yuhua Zhu , Wenhui Zhao , Dingwen Zhang

Currently, large language models (LLMs) predominantly focus on the text modality. To enable more natural human-AI interaction, speech LLMs are emerging, but building effective end-to-end speech LLMs remains challenging due to limited data…

Computation and Language · Computer Science 2026-04-14 Yan Zhou , Qingkai Fang , Yun Hong , Yang Feng

This paper presents StreamChat, a novel approach that enhances the interaction capabilities of Large Multimodal Models (LMMs) with streaming video content. In streaming interaction scenarios, existing methods rely solely on visual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Jihao Liu , Zhiding Yu , Shiyi Lan , Shihao Wang , Rongyao Fang , Jan Kautz , Hongsheng Li , Jose M. Alvare

Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation models remain fragmented, specializing narrowly in image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yibin Yan , Jilan Xu , Shangzhe Di , Haoning Wu , Weidi Xie

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Haomiao Xiong , Zongxin Yang , Jiazuo Yu , Yunzhi Zhuge , Lu Zhang , Jiawen Zhu , Huchuan Lu

Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions. This enrichment…

Computation and Language · Computer Science 2024-07-08 Chang-Sheng Kao , Yun-Nung Chen

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, limiting their…

GPT-4o, an all-encompassing model, represents a milestone in the development of large multi-modal language models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-06 Zhifei Xie , Changqiao Wu