English
Related papers

Related papers: Xuanwu: Evolving General Multimodal Models into an…

200 papers

We present ZAYA1-VL-8B, a compact mixture-of-experts vision-language model built upon our in-house language model, ZAYA1-8B. Despite its compact size, ZAYA1-VL achieves performance competitive with leading base models such as Molmo2-4B and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Hassan Shapourian , Kasra Hejazi , Olabode M. Sule , Beren Millidge

We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead the field of lightweight models by offering comprehensive…

Vision-Language Models (VLMs) are increasingly tasked with ultra-long multimodal understanding. While linear architectures offer constant computation and memory footprints, they often struggle with high-frequency visual perception compared…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Hongyuan Tao , Bencheng Liao , Shaoyu Chen , Haoran Yin , Qian Zhang , Wenyu Liu , Xinggang Wang

Long-document topic segmentation plays an important role in information retrieval and document understanding, yet existing methods still show clear shortcomings in ultra-long text settings. Traditional discriminative models are constrained…

Computation and Language · Computer Science 2026-03-02 Kaifeng Wu , Junyan Wu , Qiang Liu , Jiarui Zhang , Wen Xu

We introduce NVLM 1.0, a family of frontier-class multimodal large language models (LLMs) that achieve state-of-the-art results on vision-language tasks, rivaling the leading proprietary models (e.g., GPT-4o) and open-access models (e.g.,…

Computation and Language · Computer Science 2024-10-24 Wenliang Dai , Nayeon Lee , Boxin Wang , Zhuolin Yang , Zihan Liu , Jon Barker , Tuomas Rintamaki , Mohammad Shoeybi , Bryan Catanzaro , Wei Ping

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Le Thien Phuc Nguyen , Zhuoran Yu , Samuel Low Yu Hang , Subin An , Jeongik Lee , Yohan Ban , SeungEun Chung , Thanh-Huy Nguyen , JuWan Maeng , Soochahn Lee , Yong Jae Lee

The impressive development of large language models (LLMs) is expanding into the realm of large multimodal models (LMMs), which incorporate multiple types of data beyond text. However, the nature of multimodal models leads to significant…

Computation and Language · Computer Science 2024-08-05 Dongjae Shin , Hyeonseok Lim , Inho Won , Changsu Choi , Minjun Kim , Seungwoo Song , Hangyeol Yoo , Sangmin Kim , Kyungtae Lim

Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy. Existing token reduction methods often ignore the textual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Kaitong Cai , Jusheng Zhang , Jing Yang , Yijia Fan , Pengtao Xie , Jian Wang , Keze Wang

We introduce Motif-2-12.7B-Reasoning, a 12.7B parameter language model designed to bridge the gap between open-weight systems and proprietary frontier models in complex reasoning and long-context understanding. Addressing the common…

Recent advances in vision-language models (VLMs) have sparked growing interest in using them to automate web tasks, yet their feasibility as independent agents that reason and act purely from visual input remains underexplored. We…

Human-Computer Interaction · Computer Science 2026-04-14 Alexandra Yakovleva , Henrik Pärssinen , Harri Valpola , Juho Kannala , Alexander Ilin

We introduce Eagle 2.5, a family of frontier vision-language models (VLMs) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist…

As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mamba's long-sequence processing efficiency with Transformer's…

Computation and Language · Computer Science 2025-07-08 Tencent Hunyuan Team , Ao Liu , Botong Zhou , Can Xu , Chayse Zhou , ChenChen Zhang , Chengcheng Xu , Chenhao Wang , Decheng Wu , Dengpeng Wu , Dian Jiao , Dong Du , Dong Wang , Feng Zhang , Fengzong Lian , Guanghui Xu , Guanwei Zhang , Hai Wang , Haipeng Luo , Han Hu , Huilin Xu , Jiajia Wu , Jianchen Zhu , Jianfeng Yan , Jiaqi Zhu , Jihong Zhang , Jinbao Xue , Jun Xia , Junqiang Zheng , Kai Liu , Kai Zhang , Kai Zheng , Kejiao Li , Keyao Wang , Lan Jiang , Lixin Liu , Lulu Wu , Mengyuan Huang , Peijie Yu , Peiqi Wang , Qian Wang , Qianbiao Xiang , Qibin Liu , Qingfeng Sun , Richard Guo , Ruobing Xie , Saiyong Yang , Shaohua Chen , Shihui Hu , Shuai Li , Shuaipeng Li , Shuang Chen , Suncong Zheng , Tao Yang , Tian Zhang , Tinghao Yu , Weidong Han , Weijie Liu , Weijin Zhou , Weikang Wang , Wesleye Chen , Xiao Feng , Xiaoqin Ren , Xingwu Sun , Xiong Kuang , Xuemeng Huang , Xun Cao , Yanfeng Chen , Yang Du , Zhen Yang , Yangyu Tao , Yaping Deng , Yi Shen , Yigeng Hong , Yiqi Chen , Yiqing Huang , Yuchi Deng , Yue Mao , Yulong Wang , Yuyuan Zeng , Zenan Xu , Zhanhui Kang , Zhe Zhao , ZhenXiang Yan , Zheng Fang , Zhichao Hu , Zhongzhi Chen , Zhuoyu Li , Zongwei Li , Alex Yan , Ande Liang , Baitong Liu , Beiping Pan , Bin Xing , Binghong Wu , Bingxin Qu , Bolin Ni , Boyu Wu , Chen Li , Cheng Jiang , Cheng Zhang , Chengjun Liu , Chengxu Yang , Chengzhong Xu , Chiyu Wang , Chong Zha , Daisy Yi , Di Wang , Fanyang Lu , Fei Chen , Feifei Liu , Feng Zheng , Guanghua Yu , Guiyang Li , Guohua Wang , Haisheng Lin , Han Liu , Han Wang , Hao Fei , Hao Lu , Haoqing Jiang , Haoran Sun , Haotian Zhu , Huangjin Dai , Huankui Chen , Huawen Feng , Huihui Cai , Huxin Peng , Jackson Lv , Jiacheng Shi , Jiahao Bu , Jianbo Li , Jianglu Hu , Jiangtao Guan , Jianing Xu , Jianwei Cai , Jiarong Zhang , Jiawei Song , Jie Jiang , Jie Liu , Jieneng Yang , Jihong Zhang , Jin lv , Jing Zhao , Jinjian Li , Jinxing Liu , Jun Zhao , Juntao Guo , Kai Wang , Kan Wu , Lei Fu , Lei He , Lei Wang , Li Liu , Liang Dong , Liya Zhan , Long Cheng , Long Xu , Mao Zheng , Meng Liu , Mengkang Hu , Nanli Chen , Peirui Chen , Peng He , Pengju Pan , Pengzhi Wei , Qi Yang , Qi Yi , Roberts Wang , Rongpeng Chen , Rui Sun , Rui Yang , Ruibin Chen , Ruixu Zhou , Shaofeng Zhang , Sheng Zhang , Shihao Xu , Shuaishuai Chang , Shulin Liu , SiQi Wang , Songjia Feng , Songling Yuan , Tao Zhang , Tianjiao Lang , Tongkai Li , Wei Deng , Wei Li , Weichao Wang , Weigang Zhang , Weixuan Sun , Wen Ouyang , Wenxiang Jiao , Wenzhi Sun , Wenzhuo Jia , Xiang Zhang , Xiangyu He , Xianshun Ren , XiaoYing Zhu , Xiaolong Guo , Xiaoxue Li , Xiaoyu Ma , Xican Lu , Xinhua Feng , Xinting Huang , Xinyu Guan , Xirui Li , Xu Zhang , Xudong Gao , Xun Luo , Xuxiang Qi , Yangkun Chen , Yangyu Tao , Yanling Xiao , Yantao Mai , Yanze Chen , Yao Ding , Yeting Yang , YiFan Song , Yifan Yang , Yijiao Zhu , Yinhe Wu , Yixian Liu , Yong Yang , Yuanjun Cai , Yuanlin Tu , Yue Zhang , Yufei Huang , Yuhang Zhou , Yuhao Jiang , Yuhong Liu , Yuhui Hu , Yujin Lin , Yun Yang , Yunhao Wang , Yusong Zhang , Zekun Wu , Zelong Zhang , Zhan Yu , Zhaoliang Yang , Zhe Zhao , Zheng Li , Zhenyu Huang , Zhiguang Liu , Zhijiang Xu , Zhiqing Kui , Zhiyin Zeng , Zhiyuan Xiong , Zhuo Han , Zifan Wu , Zigang Geng , Zilong Zhao , Ziyan Tang , Ziyuan Zhu , Zonglei Zhu , Zhijiang Xu

We present Chameleon, a family of early-fusion token-based mixed-modal models capable of understanding and generating images and text in any arbitrary sequence. We outline a stable training approach from inception, an alignment recipe, and…

Computation and Language · Computer Science 2025-03-24 Chameleon Team

Recent advancements in unified multimodal understanding and visual generation (or multimodal generation) models have been hindered by their quadratic computational complexity and dependence on large-scale training data. We present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

We study how to endow GUI agents with scalable memory that help generalize across unfamiliar interfaces and long-horizon tasks. Prior GUI agents compress past trajectories into text tokens, which balloons context length and misses decisive…

Artificial Intelligence · Computer Science 2025-10-13 Wenyi Wu , Kun Zhou , Ruoxin Yuan , Vivian Yu , Stephen Wang , Zhiting Hu , Biwei Huang

We present Ruyi2.5, a multimodal familial model built on the AI Flow framework. Extending Ruyi2's "Train Once, Deploy Many" paradigm to the multimodal domain, Ruyi2.5 constructs a shared-backbone architecture that co-trains models of…

Computation and Language · Computer Science 2026-03-19 Huan Song , Shuyu Tian , Qingfei Zhao , Wenhao Hong , Jiang Liu , Ting Long , Jiawei Shao , Xuelong Li

Multimodal Large Reasoning Models (MLRMs) demonstrate impressive cross-modal reasoning but often amplify safety risks under adversarial or unsafe prompts, a phenomenon we call the \textit{Reasoning Tax}. Existing defenses mainly act at the…

Machine Learning · Computer Science 2025-10-10 Huahui Yi , Kun Wang , Qiankun Li , Miao Yu , Liang Lin , Gongli Xi , Hao Wu , Xuming Hu , Kang Li , Yang Liu

Modern automotive infotainment systems necessitate intelligent and adaptive solutions to manage frequent User Interface (UI) updates and diverse design variations. This work introduces a vision-language framework to facilitate the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Benjamin Raphael Ernhofer , Daniil Prokhorov , Jannica Langner , Dominik Bollmann

With the rapid penetration of artificial intelligence across industries and scenarios, a key challenge in building the next-generation intelligent core lies in effectively integrating the language understanding capabilities of foundation…

Artificial Intelligence · Computer Science 2025-06-03 Liang Geng

Multimodal Large Language Models (MLLMs) based agents have demonstrated remarkable potential in autonomous web navigation. However, handling long-horizon tasks remains a critical bottleneck. Prevailing strategies often rely heavily on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Dawei Yan , Haokui Zhang , Guangda Huzhang , Yang Li , Yibo Wang , Qing-Guo Chen , Zhao Xu , Weihua Luo , Ying Li , Wei Dong , Chunhua Shen