English
Related papers

Related papers: MobileFlow: A Multimodal LLM For Mobile GUI Agent

200 papers

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

Computation and Language · Computer Science 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

Recent advances in large language models (LLMs) have enabled new applications in e-commerce customer service. However, their capabilities remain constrained in complex, multimodal scenarios. We present MindFlow, the first open-source…

Computation and Language · Computer Science 2025-07-09 Ming Gong , Xucheng Huang , Chenghan Yang , Xianhan Peng , Haoxin Wang , Yang Liu , Ling Jiang

The rise of Large Language Models (LLMs) and generative visual analytics systems has transformed data-driven insights, yet significant challenges persist in accurately interpreting users' analytical and interaction intents. While language…

Human-Computer Interaction · Computer Science 2025-04-17 Juntong Chen , Jiang Wu , Jiajing Guo , Vikram Mohanty , Xueming Li , Jorge Piazentin Ono , Wenbin He , Liu Ren , Dongyu Liu

In past years, we have been dedicated to automating user acceptance testing (UAT) process of WeChat Pay, one of the most influential mobile payment applications in China. A system titled XUAT has been developed for this purpose. However,…

Artificial Intelligence · Computer Science 2024-01-11 Zhitao Wang , Wei Wang , Zirao Li , Long Wang , Can Yi , Xinjie Xu , Luyang Cao , Hanjing Su , Shouzhi Chen , Jun Zhou

Text-rich visual understanding-the ability to process environments where dense textual content is integrated with visuals-is crucial for multimodal large language models (MLLMs) to interact effectively with structured environments. To…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Junpeng Liu , Tianyue Ou , Yifan Song , Yuxiao Qu , Wai Lam , Chenyan Xiong , Wenhu Chen , Graham Neubig , Xiang Yue

In recent years, rapid advances in Multimodal Large Language Models (MLLMs) have increasingly stimulated research on ancient Chinese scripts. As the evolution of written characters constitutes a fundamental pathway for understanding…

Computation and Language · Computer Science 2026-04-14 Rui Song , Lida Shi , Ruihua Qi , Yingji Li , Hao Xu

Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking interactivity and adaptability for diverse analytical…

Artificial Intelligence · Computer Science 2025-02-28 Lei Li , Sen Jia , Jianhao Wang , Zhaochong An , Jiaang Li , Jenq-Neng Hwang , Serge Belongie

Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled powerful autonomous agents capable of complex reasoning and multi-modal tool use. Despite their growing capabilities, today's agent frameworks…

Artificial Intelligence · Computer Science 2025-06-12 Peiran Li , Xinkai Zou , Zhuohang Wu , Ruifeng Li , Shuo Xing , Hanwen Zheng , Zhikai Hu , Yuping Wang , Haoxi Li , Qin Yuan , Yingmo Zhang , Zhengzhong Tu

The Large Vision-Language Models (LVLMs) have demonstrated great abilities in image perception and language understanding. However, existing multimodal benchmarks focus on primary perception abilities and commonsense knowledge which are…

Computation and Language · Computer Science 2024-08-07 Yi Zong , Xipeng Qiu

Mobile Graphical User Interface (GUI) World Models (WMs) offer a promising path for improving mobile GUI agent performance at train- and inference-time. However, current approaches face a critical trade-off: text-based WMs sacrifice visual…

Machine Learning · Computer Science 2026-05-26 Woosung Koh , Sungjun Han , Segyu Lee , Se-Young Yun , Jamin Shin

Multimodal Large Language Models (MLLMs) show strong potential for interpreting and interacting with complex, pixel-rich Graphical User Interface (GUI) environments. However, building agents that are both efficient for high-level tasks and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Sarthak Mehrotra , Sairam V C Rebbapragada , Mani Hemanth Reddy Bonthu , Vineeth N Balasubramanian

The recent success of large vision language models shows great potential in driving the agent system operating on user interfaces. However, we argue that the power multimodal models like GPT-4V as a general agent on multiple operating…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Yadong Lu , Jianwei Yang , Yelong Shen , Ahmed Awadallah

Recent advancements in Large Language Models (LLMs) have shown significant progress in understanding complex natural language. One important application of LLM is LLM-based AI Agent, which leverages the ability of LLM as well as external…

Computation and Language · Computer Science 2024-07-19 Zelong Li , Shuyuan Xu , Kai Mei , Wenyue Hua , Balaji Rama , Om Raheja , Hao Wang , He Zhu , Yongfeng Zhang

Graphical User Interface (GUI) agents extend large language models from text generation to action execution in real-world digital environments. Unlike conversational systems, GUI agents perform irreversible operations such as submitting…

Machine Learning · Computer Science 2026-02-25 Yucheng Shi , Wenhao Yu , Jingyuan Huang , Wenlin Yao , Wenhu Chen , Ninghao Liu

GUI agents have emerged as a powerful paradigm for automating interactions in digital environments, yet achieving both broad generality and consistently strong task performance remains challenging. In this report, we present UI-Venus-1.5, a…

The rapid progress of large language models (LLMs) has sparked growing interest in building Artificial General Intelligence (AGI) within Graphical User Interface (GUI) environments. However, existing GUI agents based on LLMs or…

Artificial Intelligence · Computer Science 2025-05-27 Runliang Niu , Jinglong Ji , Yi Chang , Qi Wang

Recent advancements in multimodal large language models (MLLMs) have shown exceptional potential in enabling mobile-using agents to autonomously execute human instructions. However, fully automated agents often try to execute tasks even…

Computation and Language · Computer Science 2026-05-28 Zheng Wu , Pengzhou Cheng , Zongru Wu , Yuan Guo , Tianjie Ju , Aston Zhang , Gongshen Liu , Zhuosheng Zhang

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models in multimodal understanding. We introduce three simple…

Vision-language model-based mobile agents have gained the ability to understand complex instructions and mobile screenshots, benefiting from reinforcement learning paradigms like Group Relative Policy Optimization (GRPO). However, existing…

Artificial Intelligence · Computer Science 2026-04-28 Jihao Gu , Qihang Ai , Yingyao Wang , Pi Bu , Jingxuan Xing , Zekun Zhu , Wei Jiang , Ziming Wang , Yingxiu Zhao , Ming-Liang Zhang , Jun Song , Yuning Jiang , Bo Zheng

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer…

Human-Computer Interaction · Computer Science 2026-02-20 Ricardo E. Gonzalez Penuela , Crescentia Jung , Sharon Y Lin , Ruiying Hu , Shiri Azenkot
‹ Prev 1 4 5 6 7 8 10 Next ›