English
Related papers

Related papers: Intern-S1: A Scientific Multimodal Foundation Mode…

200 papers

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond…

Machine Learning · Computer Science 2026-04-03 Yicheng Zou , Dongsheng Zhu , Lin Zhu , Tong Zhu , Yunhua Zhou , Peiheng Zhou , Xinyu Zhou , Dongzhan Zhou , Zhiwang Zhou , Yuhao Zhou , Bowen Zhou , Zhanping Zhong , Zhijie Zhong , Haiteng Zhao , Penghao Zhao , Xiaomeng Zhao , Zhiyuan Zhao , Yechen Zhang , Jin Zhang , Wenwei Zhang , Hongjie Zhang , Zhuo Zhang , Wenlong Zhang , Bo Zhang , Chao Zhang , Chen Zhang , Yuhang Zang , Fei Yuan , Jiakang Yuan , Jiashuo Yu , Jinhui Yin , Haochen Ye , Qian Yao , Bowen Yang , Danni Yang , Kaichen Yang , Ziang Yan , Jun Xu , Yicheng Xu , Wanghan Xu , Xuenan Xu , Chao Xu , Ruiliang Xu , Shuhao Xing , Long Xing , Xinchen Xie , Ling-I Wu , Zijian Wu , Zhenyu Wu , Lijun Wu , Yue Wu , Jianyu Wu , Wen Wu , Fan Wu , Xilin Wei , Qi Wei , Bingli Wang , Rui Wang , Ziyi Wang , Zun Wang , Yi Wang , Haomin Wang , Yizhou Wang , Lintao Wang , Yiheng Wang , Longjiang Wang , Bin Wang , Jian Tong , Zhongbo Tian , Huanze Tang , Chen Tang , Shixiang Tang , Yu Sun , Qiushi Sun , Xuerui Su , Qisheng Su , Chenlin Su , Demin Song , Jin Shi , Fukai Shang , Yuchen Ren , Pengli Ren , Xiaoye Qu , Yuan Qu , Jiantao Qiu , Yu Qiao , Biqing Qi , Runyu Peng , Tianshuo Peng , Jiahui Peng , Qizhi Pei , Zhuoshi Pan , Linke Ouyang , Wenchang Ning , Yichuan Ma , Zerun Ma , Ningsheng Ma , Runyuan Ma , Chengqi Lyu , Haijun Lv , Han Lv , Lindong Lu , Kuikun Liu , Jiangning Liu , Yuhong Liu , Kai Liu , Hongwei Liu , Zhoumianze Liu , Mengjie Liu , Ziyu Liu , Wenran Liu , Yang Liu , Liwei Liu , Kaiwen Liu , Junyao Lin , Junming Lin , Tianyang Lin , Dahua Lin , Jianze Liang , Linyang Li , Peiji Li , Zonglin Li , Zehao Li , Pengze Li , Guoyan Li , Lingkai Kong , Linglin Jing , Zhenjiang Jin , Feifei Jiang , Qian Jiang , Junhao Huang , Zixian Huang , Haian Huang , Zhouqi Hua , Ermo Hua , Han Hu , Linfeng Hou , Yinan He , Conghui He , Tianyao He , Xu Guo , Qipeng Guo , Aijia Guo , Yuzhe Gu , Lixin Gu , Jingyang Gong , Qiming Ge , Jiaye Ge , Songyang Gao , Jianfei Gao , Xinyu Fang , Caihua fan , Yue Fan , Yanhui Duan , Zichen Ding , Shengyuan Ding , Ning Ding , Xuanlang Dai , Erfei Cui , Ganqu Cui , Pei Chu , Tao Chu , Guangran Cheng , Yu Cheng , Kai Chen , Yongkang Chen , Chiyu Chen , Guanzhou Chen , Qiaosheng Chen , Sitao Chen , Xin Chen , Haojiong Chen , Yicheng Chen , Weihan Cao , Yuhang Cao , Qinglong Cao , Lei Bai

The rapid advancement of foundation models has revolutionized visual representation learning in a self-supervised manner. However, their application in remote sensing (RS) remains constrained by a fundamental gap: existing models…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Hanbo Bi , Yingchao Feng , Boyuan Tong , Mengyu Wang , Haichen Yu , Yongqiang Mao , Hao Chang , Wenhui Diao , Peijin Wang , Yue Yu , Hanyang Peng , Yehong Zhang , Kun Fu , Xian Sun

We introduce Timer-S1, a strong Mixture-of-Experts (MoE) time series foundation model with 8.3B total parameters, 0.75B activated parameters for each token, and a context length of 11.5K. To overcome the scalability bottleneck in existing…

Artificial Intelligence · Computer Science 2026-04-10 Yong Liu , Xingjian Su , Shiyu Wang , Haoran Zhang , Haixuan Liu , Yuxuan Wang , Zhou Ye , Yang Xiang , Jianmin Wang , Mingsheng Long

Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this report, we present dots.llm1, a large-scale MoE model that…

Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits,…

Machine Learning · Computer Science 2024-05-01 Paul Pu Liang

Modern applications increasingly involve many heterogeneous input streams, such as clinical sensors, wearable device data, imaging, and text, each with distinct measurement models, sampling rates, and noise characteristics. We define this…

Machine Learning · Computer Science 2026-03-03 Xing Han , Hsing-Huan Chung , Joydeep Ghosh , Paul Pu Liang , Suchi Saria

Recent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE)…

Artificial Intelligence · Computer Science 2024-05-21 Yunxin Li , Shenyuan Jiang , Baotian Hu , Longyue Wang , Wanqi Zhong , Wenhan Luo , Lin Ma , Min Zhang

Artificial intelligence (AI) has achieved astonishing successes in many domains, especially with the recent breakthroughs in the development of foundational large models. These large models, leveraging their extensive training data, provide…

Machine Learning · Computer Science 2026-01-27 Siyuan Mu , Sen Lin

Mixture-of-Experts (MoE) architectures have emerged as a promising direction, offering efficiency and scalability by activating only a subset of parameters during inference. However, current research remains largely performance-centric,…

Machine Learning · Computer Science 2025-09-30 Jiahao Ying , Mingbao Lin , Qianru Sun , Yixin Cao

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across multi-modal tasks by scaling model size and training data. However, these dense LVLMs incur significant computational costs and motivate the exploration of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Dianyi Wang , Siyuan Wang , Zejun Li , Yikun Wang , Yitong Li , Duyu Tang , Xiaoyu Shen , Xuanjing Huang , Zhongyu Wei

The Mixture of Experts (MoE) models are an emerging class of sparsely activated deep learning models that have sublinear compute costs with respect to their parameters. In contrast with dense models, the sparse architecture of MoE offers…

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 billion per token. Training such models at a…

The rapid advancement of autonomous systems, including self-driving vehicles and drones, has intensified the need to forge true Spatial Intelligence from multi-modal onboard sensor data. While foundation models excel in single-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Song Wang , Lingdong Kong , Xiaolu Liu , Hao Shi , Wentong Li , Jianke Zhu , Steven C. H. Hoi

Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. We present S1-MMAlign, a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 He Wang , Longteng Guo , Pengkang Huo , Xuanxu Lin , Yichen Yuan , Jie Jiang , Jing Liu

Modality fusion is a cornerstone of multimodal learning, enabling information integration from diverse data sources. However, vanilla fusion methods are limited by (1) inability to account for heterogeneous interactions between modalities…

Machine Learning · Computer Science 2025-05-27 Jiayi Xin , Sukwon Yun , Jie Peng , Inyoung Choi , Jenna L. Ballard , Tianlong Chen , Qi Long

To help the open-source community have a better understanding of Mixture-of-Experts (MoE) based large language models (LLMs), we train and release OpenMoE, a series of fully open-sourced and reproducible decoder-only MoE LLMs, ranging from…

Computation and Language · Computer Science 2024-03-28 Fuzhao Xue , Zian Zheng , Yao Fu , Jinjie Ni , Zangwei Zheng , Wangchunshu Zhou , Yang You

Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of capabilities, e.g., code, math, or domain-specific knowledge. Mixture-of-Experts (MoEs)…

Computation and Language · Computer Science 2026-05-12 Ryan Wang , Akshita Bhagia , Sewon Min

Mixture-of-Experts (MoE) models enable scalable performance by activating large parameter sets sparsely, minimizing computational overhead. To mitigate the prohibitive cost of training MoEs from scratch, recent work employs upcycling,…

Machine Learning · Computer Science 2025-11-13 Qi Wang , Hanyang Peng , Yue Yu

Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borrow representational…

Machine Learning · Computer Science 2026-05-12 Xing Han , Shravan Chaudhari , Tanvi Ranade , Rama Chellappa , Suchi Saria
‹ Prev 1 2 3 10 Next ›