English
Related papers

Related papers: Baichuan-Omni-1.5 Technical Report

200 papers

The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world…

Sound · Computer Science 2026-03-10 Wenjie Tian , Zhixian Zhao , Jingbin Hu , Huakang Chen , Haohe Liu , Binshen Mu , Lei Xie

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Wendong Bu , Kaihang Pan , Yuze Lin , Jiacheng Li , Kai Shen , Wenqiao Zhang , Juncheng Li , Jun Xiao , Siliang Tang

Accurate and timely prediction of tool conditions is critical for intelligent manufacturing systems, where unplanned tool failures can lead to quality degradation and production downtime. In modern industrial environments, predictive…

Artificial Intelligence · Computer Science 2025-11-04 Ziqi Wang , Hailiang Zhao , Yuhao Yang , Daojiang Hu , Cheng Bao , Mingyi Liu , Kai Di , Schahram Dustdar , Zhongjie Wang , Shuiguang Deng

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yiheng Li , Zhuo Li , Ruibing Hou , Yingjie Chen , Hong Chang , Hao Liu , Shiguang Shan

The use of omni-LLMs (large language models that accept any modality as input), particularly for multimodal cognitive state tasks involving speech, is understudied. We present OmniVox, the first systematic evaluation of four omni-LLMs on…

Computation and Language · Computer Science 2025-03-31 John Murzaku , Owen Rambow

We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Inclusion AI , Biao Gong , Cheng Zou , Dandan Zheng , Hu Yu , Jingdong Chen , Jianxin Sun , Junbo Zhao , Jun Zhou , Kaixiang Ji , Lixiang Ru , Libin Wang , Qingpei Guo , Rui Liu , Weilong Chai , Xinyu Xiao , Ziyuan Huang

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Keda Tao , Wenjie Du , Bohan Yu , Weiqiang Wang , Jian Liu , Huan Wang

Omnimodal Large Language Models (Omni-LLMs) incur substantial computational overhead due to the large number of multimodal input tokens they process, making token reduction essential for real-world deployment. Existing Omni-LLM pruning…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Chaeyoung Jung , Kyeongha Rho , Joon Son Chung

Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that…

The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimodal customization for the simultaneous preservation of visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yuheng Chen , Qingdong He , Teng Hu , Yuji Wang , Yabiao Wang , Lizhuang Ma , Jiangning Zhang

We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image, video, audio, text, depth, point cloud, time series,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Siddharth Srivastava , Gaurav Sharma

We introduce Baichuan Alignment, a detailed analysis of the alignment techniques employed in the Baichuan series of models. This represents the industry's first comprehensive account of alignment methodologies, offering valuable insights…

This tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video. Participants will gain an understanding of the…

Computation and Language · Computer Science 2024-10-10 Soyeon Caren Han , Feiqi Cao , Josiah Poon , Roberto Navigli

Research on multi-modal learning dominantly aligns the modalities in a unified space at training, and only a single one is taken for prediction at inference. However, for a real machine, e.g., a robot, sensors could be added or removed at…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Yuanhuiyi Lyu , Xu Zheng , Dahun Kim , Lin Wang

Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evaluation frameworks for such models suffer from critical limitations, including modality…

Computation and Language · Computer Science 2026-04-29 Seunghee Kim , Ingyu Bang , Seokgyu Jang , Changhyeon Kim , Sanghwan Bae , Jihun Choi , Richeng Xuan , Taeuk Kim

Federated learning (FL) has become a promising paradigm for collaborative medical image analysis, yet existing frameworks remain tightly coupled to task-specific backbones and are fragile under heterogeneous imaging modalities. Such…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Meilin Liu , Jiaying Wang , Jing Shan

We introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jack Hong , Shilin Yan , Jiayin Cai , Xiaolong Jiang , Yao Hu , Weidi Xie

Multimodal spatiotemporal learning on real-world experimental data is constrained by two challenges: within-modality measurements are sparse, irregular, and noisy (QA/QC artifacts) but cross-modally correlated; the set of available…

Machine Learning · Computer Science 2025-11-05 Kevin Valencia , Thilina Balasooriya , Xihaier Luo , Shinjae Yoo , David Keetae Park