English
Related papers

Related papers: GLM-5V-Turbo: Toward a Native Foundation Model for…

200 papers

We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to…

Machine Learning · Computer Science 2026-02-25 GLM-5-Team , : , Aohan Zeng , Xin Lv , Zhenyu Hou , Zhengxiao Du , Qinkai Zheng , Bin Chen , Da Yin , Chendi Ge , Chenghua Huang , Chengxing Xie , Chenzheng Zhu , Congfeng Yin , Cunxiang Wang , Gengzheng Pan , Hao Zeng , Haoke Zhang , Haoran Wang , Huilong Chen , Jiajie Zhang , Jian Jiao , Jiaqi Guo , Jingsen Wang , Jingzhao Du , Jinzhu Wu , Kedong Wang , Lei Li , Lin Fan , Lucen Zhong , Mingdao Liu , Mingming Zhao , Pengfan Du , Qian Dong , Rui Lu , Shuang-Li , Shulin Cao , Song Liu , Ting Jiang , Xiaodong Chen , Xiaohan Zhang , Xuancheng Huang , Xuezhen Dong , Yabo Xu , Yao Wei , Yifan An , Yilin Niu , Yitong Zhu , Yuanhao Wen , Yukuo Cen , Yushi Bai , Zhongpei Qiao , Zihan Wang , Zikang Wang , Zilin Zhu , Ziqiang Liu , Zixuan Li , Bojie Wang , Bosi Wen , Can Huang , Changpeng Cai , Chao Yu , Chen Li , Chengwei Hu , Chenhui Zhang , Dan Zhang , Daoyan Lin , Dayong Yang , Di Wang , Ding Ai , Erle Zhu , Fangzhou Yi , Feiyu Chen , Guohong Wen , Hailong Sun , Haisha Zhao , Haiyi Hu , Hanchen Zhang , Hanrui Liu , Hanyu Zhang , Hao Peng , Hao Tai , Haobo Zhang , He Liu , Hongwei Wang , Hongxi Yan , Hongyu Ge , Huan Liu , Huanpeng Chu , Jia'ni Zhao , Jiachen Wang , Jiajing Zhao , Jiamin Ren , Jiapeng Wang , Jiaxin Zhang , Jiayi Gui , Jiayue Zhao , Jijie Li , Jing An , Jing Li , Jingwei Yuan , Jinhua Du , Jinxin Liu , Junkai Zhi , Junwen Duan , Kaiyue Zhou , Kangjian Wei , Ke Wang , Keyun Luo , Laiqiang Zhang , Leigang Sha , Liang Xu , Lindong Wu , Lintao Ding , Lu Chen , Minghao Li , Nianyi Lin , Pan Ta , Qiang Zou , Rongjun Song , Ruiqi Yang , Shangqing Tu , Shangtong Yang , Shaoxiang Wu , Shengyan Zhang , Shijie Li , Shuang Li , Shuyi Fan , Wei Qin , Wei Tian , Weining Zhang , Wenbo Yu , Wenjie Liang , Xiang Kuang , Xiangmeng Cheng , Xiangyang Li , Xiaoquan Yan , Xiaowei Hu , Xiaoying Ling , Xing Fan , Xingye Xia , Xinyuan Zhang , Xinze Zhang , Xirui Pan , Xu Zou , Xunkai Zhang , Yadi Liu , Yandong Wu , Yanfu Li , Yidong Wang , Yifan Zhu , Yijun Tan , Yilin Zhou , Yiming Pan , Ying Zhang , Yinpei Su , Yipeng Geng , Yong Yan , Yonglin Tan , Yuean Bi , Yuhan Shen , Yuhao Yang , Yujiang Li , Yunan Liu , Yunqing Wang , Yuntao Li , Yurong Wu , Yutao Zhang , Yuxi Duan , Yuxuan Zhang , Zezhen Liu , Zhengtao Jiang , Zhenhe Yan , Zheyu Zhang , Zhixiang Wei , Zhuo Chen , Zhuoer Feng , Zijun Yao , Ziwei Chai , Ziyuan Wang , Zuzhou Zhang , Bin Xu , Minlie Huang , Hongning Wang , Juanzi Li , Yuxiao Dong , Jie Tang

Multimodal large language models (MLLMs) have shown strong capabilities but remain limited to fixed modality pairs and require costly fine-tuning with large aligned datasets. Building fully omni-capable models that can integrate text,…

Artificial Intelligence · Computer Science 2025-11-06 Huawei Lin , Yunzhi Shi , Tong Geng , Weijie Zhao , Wei Wang , Ravender Pal Singh

We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this report, we share our key findings in the development of the…

Mobile agents are essential for automating tasks in complex and dynamic mobile environments. As foundation models evolve, the demands for agents that can adapt in real-time and process multimodal data have grown. This survey provides a…

Artificial Intelligence · Computer Science 2025-09-16 Biao Wu , Yanda Li , Zhiwei Zhang , Yunchao Wei , Meng Fang , Ling Chen

The advancement of large language models (LLMs) prompts the development of multi-modal agents, which are used as a controller to call external tools, providing a feasible way to solve practical tasks. In this paper, we propose a multi-modal…

Artificial Intelligence · Computer Science 2025-02-04 Zhi Gao , Bofei Zhang , Pengxiang Li , Xiaojian Ma , Tao Yuan , Yue Fan , Yuwei Wu , Yunde Jia , Song-Chun Zhu , Qing Li

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and language comprehension…

A key challenge in training Vision-Language Model (VLM) agents, compared to Language Model (LLM) agents, lies in the shift from textual states to complex visual observations. This transition introduces partial observability and demands…

Unmanned Aerial Vehicles (UAVs) are increasingly used in defense, surveillance, and disaster response, yet most systems still operate at SAE Level 2 to 3 autonomy. Their dependence on rule-based control and narrow AI limits adaptability in…

Artificial Intelligence · Computer Science 2025-12-03 Anis Koubaa , Khaled Gabr

The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven and rule-based, have been hindered by their inability to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Licheng Wen , Xuemeng Yang , Daocheng Fu , Xiaofeng Wang , Pinlong Cai , Xin Li , Tao Ma , Yingxuan Li , Linran Xu , Dengke Shang , Zheng Zhu , Shaoyan Sun , Yeqi Bai , Xinyu Cai , Min Dou , Shuanglu Hu , Botian Shi , Yu Qiao

This paper introduces a multi-agent framework for comprehensive highway scene understanding, designed around a mixture-of-experts strategy. In this framework, a large generic vision-language model (VLM), such as GPT-4o, is contextualized…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Yunxiang Yang , Ningning Xu , Jidong J. Yang

Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Tajamul Ashraf , Umair Nawaz , Abdelrahman M. Shaker , Rao Anwer , Philip Torr , Fahad Shahbaz Khan , Salman Khan

Open-source pre-trained Large Language Models (LLMs) exhibit strong language understanding and generation capabilities, making them highly successful in a variety of tasks. However, when used as agents for dealing with complex problems in…

Computation and Language · Computer Science 2024-04-01 Qinhao Zhou , Zihan Zhang , Xiang Xiang , Ke Wang , Yuchuan Wu , Yongbin Li

Extending the capabilities of Large Language Models (LLMs) with functions or tools for environment interaction has led to the emergence of the agent paradigm. In industry, training an LLM is not always feasible because of the scarcity of…

Computation and Language · Computer Science 2025-01-14 Saptarshi Sengupta , Harsh Vashistha , Kristal Curtis , Akshay Mallipeddi , Abhinav Mathur , Joseph Ross , Liang Gou

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

Computation and Language · Computer Science 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

The move toward open Sixth-Generation (6G) networks necessitates a novel approach to full-stack simulation environments for evaluating complex technology developments before prototyping and real-world implementation. This paper introduces…

Networking and Internet Architecture · Computer Science 2025-03-18 Farhad Rezazadeh , Amir Ashtari Gargari , Sandra Lagen , Houbing Song , Dusit Niyato , Lingjia Liu

Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper, we analyze the latest model, GPT-4V(ision), to deepen the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Zhengyuan Yang , Linjie Li , Kevin Lin , Jianfeng Wang , Chung-Ching Lin , Zicheng Liu , Lijuan Wang

Foundation models, including large language models (LLMs) and vision-language models (VLMs), have recently enabled novel approaches to robot autonomy and human-robot interfaces. In parallel, vision-language-action models (VLAs) or large…

Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Rohan Wadhawan , Hritik Bansal , Kai-Wei Chang , Nanyun Peng

With their prominent scene understanding and reasoning capabilities, pre-trained visual-language models (VLMs) such as GPT-4V have attracted increasing attention in robotic task planning. Compared with traditional task planning strategies,…

Robotics · Computer Science 2024-05-24 Aoran Mei , Jianhua Wang , Guo-Niu Zhu , Zhongxue Gan

Recent advances in foundation models, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs), have facilitated the development of intelligent agents capable of performing complex tasks. By leveraging the…

‹ Prev 1 2 3 10 Next ›