English
Related papers

Related papers: GPT-4V(ision) is a Generalist Web Agent, if Ground…

200 papers

Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper, we analyze the latest model, GPT-4V(ision), to deepen the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Zhengyuan Yang , Linjie Li , Kevin Lin , Jianfeng Wang , Chung-Ching Lin , Zicheng Liu , Lijuan Wang

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Xinlu Zhang , Yujie Lu , Weizhi Wang , An Yan , Jun Yan , Lianke Qin , Heng Wang , Xifeng Yan , William Yang Wang , Linda Ruth Petzold

Large Multimodal Models (LMMs) have demonstrated impressive performance across various vision and language tasks, yet their potential applications in recommendation tasks with visual assistance remain unexplored. To bridge this gap, we…

Information Retrieval · Computer Science 2023-11-08 Peilin Zhou , Meng Cao , You-Liang Huang , Qichen Ye , Peiyan Zhang , Junling Liu , Yueqi Xie , Yining Hua , Jaeboum Kim

The rapid advancement of large language models (LLMs) has led to a new era marked by the development of autonomous applications in real-world scenarios, which drives innovation in creating advanced web agents. Existing web agents typically…

Computation and Language · Computer Science 2024-06-10 Hongliang He , Wenlin Yao , Kaixin Ma , Wenhao Yu , Yong Dai , Hongming Zhang , Zhenzhong Lan , Dong Yu

Recent research has offered insights into the extraordinary capabilities of Large Multimodal Models (LMMs) in various general vision and language tasks. There is growing interest in how LMMs perform in more specialized domains. Social media…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Hanjia Lyu , Jinfa Huang , Daoan Zhang , Yongsheng Yu , Xinyi Mou , Jinsheng Pan , Zhengyuan Yang , Zhongyu Wei , Jiebo Luo

We explore the use of GPT-4 on a humanoid robot in simulation and the real world as proof of concept of a novel large language model (LLM) driven behaviour method. LLMs have shown the ability to perform various tasks, including robotic…

Robotics · Computer Science 2025-04-01 Thomas O'Brien , Ysobel Sims

The emergence of multimodal large models (MLMs) has significantly advanced the field of visual understanding, offering remarkable capabilities in the realm of visual question answering (VQA). Yet, the true challenge lies in the domain of…

Computation and Language · Computer Science 2024-08-27 Yunxin Li , Longyue Wang , Baotian Hu , Xinyu Chen , Wanqi Zhong , Chenyang Lyu , Wei Wang , Min Zhang

The advent of large language models (LLMs) has heightened interest in their potential for multimodal applications that integrate language and vision. This paper explores the capabilities of GPT-4V in the realms of geography, environmental…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Chenjiao Tan , Qian Cao , Yiwei Li , Jielu Zhang , Xiao Yang , Huaqin Zhao , Zihao Wu , Zhengliang Liu , Hao Yang , Nemin Wu , Tao Tang , Xinyue Ye , Lilong Chai , Ninghao Liu , Changying Li , Lan Mu , Tianming Liu , Gengchen Mai

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in…

Autonomous web navigation requires agents to perceive complex visual environments and maintain long-term context, yet current Large Language Model (LLM) based agents often struggle with spatial disorientation and navigation loops. In this…

Artificial Intelligence · Computer Science 2026-03-04 Xinjun Wang , Shengyao Wang , Aimin Zhou , Hao Hao

The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven and rule-based, have been hindered by their inability to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Licheng Wen , Xuemeng Yang , Daocheng Fu , Xiaofeng Wang , Pinlong Cai , Xin Li , Tao Ma , Yingxuan Li , Linran Xu , Dengke Shang , Zheng Zhu , Shaoyan Sun , Yeqi Bai , Xinyu Cai , Min Dou , Shuanglu Hu , Botian Shi , Yu Qiao

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

Computation and Language · Computer Science 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

While there is much excitement about the potential of large multimodal models (LMM), a comprehensive evaluation is critical to establish their true capabilities and limitations. In support of this aim, we evaluate two state-of-the-art LMMs,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Mengchen Liu , Chongyan Chen , Danna Gurari

Large language models (LLMs) have demonstrated a powerful ability to answer various queries as a general-purpose assistant. The continuous multi-modal large language models (MLLM) empower LLMs with the ability to perceive visual signals.…

Computation and Language · Computer Science 2024-01-05 Ziqiang Zheng , Yiwei Chen , Jipeng Zhang , Tuan-Anh Vu , Huimin Zeng , Yue Him Wong Tim , Sai-Kit Yeung

In recent years, large language models (LLMs) have become increasingly capable and can now interact with tools (i.e., call functions), read documents, and recursively call themselves. As a result, these LLMs can now function autonomously as…

Cryptography and Security · Computer Science 2024-02-19 Richard Fang , Rohan Bindu , Akul Gupta , Qiusi Zhan , Daniel Kang

Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by their poor performance in tasks that require a combination of…

Computation and Language · Computer Science 2025-12-15 Zhoutong Ye , Mingze Sun , Huan-ang Gao , Xutong Wang , Xiangyang Wang , Yu Mei , Chang Liu , Qinwei Li , Chengwen Zhang , Qinghuan Lan , Chun Yu , Yuanchun Shi

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and language comprehension…

Multimodal large language models (MLLMs) are designed to process and integrate information from multiple sources, such as text, speech, images, and videos. Despite its success in language understanding, it is critical to evaluate the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Hao Lu , Xuesong Niu , Jiyao Wang , Yin Wang , Qingyong Hu , Jiaqi Tang , Yuting Zhang , Kaishen Yuan , Bin Huang , Zitong Yu , Dengbo He , Shuiguang Deng , Hao Chen , Yingcong Chen , Shiguang Shan

Multimodal large language models (MLLMs) are transforming the capabilities of graphical user interface (GUI) agents, facilitating their transition from controlled simulations to complex, real-world applications across various platforms.…

Artificial Intelligence · Computer Science 2025-06-18 Boyu Gou , Ruohan Wang , Boyuan Zheng , Yanan Xie , Cheng Chang , Yiheng Shu , Huan Sun , Yu Su

We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability…

‹ Prev 1 2 3 10 Next ›