中文
相关论文

相关论文: Understanding Mobile GUI: from Pixel-Words to Scre…

200 篇论文

Mobile UI understanding is important for enabling various interaction tasks such as UI automation and accessibility. Previous mobile UI modeling often depends on the view hierarchy information of a screen, which directly provides the…

计算机视觉与模式识别 · 计算机科学 2023-02-27 Gang Li , Yang Li

With the metaverse slowly becoming a reality and given the rapid pace of developments toward the creation of digital humans, the need for a principled style editing pipeline for human faces is bound to increase manifold. We cater to this…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Snehal Singh Tomar , A. N. Rajagopalan

The histopathological analysis of whole-slide images (WSIs) is fundamental to cancer diagnosis but is a time-consuming and expert-driven process. While deep learning methods show promising results, dominant patch-based methods artificially…

图像与视频处理 · 电气工程与系统科学 2025-10-08 Alexander Weers , Alexander H. Berger , Laurin Lux , Peter Schüffler , Daniel Rueckert , Johannes C. Paetzold

The goal of this work is to bring semantics into the tasks of text recognition and retrieval in natural images. Although text recognition and retrieval have received a lot of attention in recent years, previous works have focused on…

计算机视觉与模式识别 · 计算机科学 2015-09-22 Albert Gordo , Jon Almazan , Naila Murray , Florent Perronnin

This study aims to explore efficient tuning methods for the screenshot captioning task. Recently, image captioning has seen significant advancements, but research in captioning tasks for mobile screens remains relatively scarce. Current…

机器学习 · 计算机科学 2023-09-27 Ching-Yu Chiang , I-Hua Chang , Shih-Wei Liao

Visual Question-Answering, a technology that generates textual responses from an image and natural language question, has progressed significantly. Notably, it can aid in tracking and inquiring about daily activities, crucial in healthcare…

机器学习 · 计算机科学 2024-10-29 Wenqiang Chen , Jiaxuan Cheng , Leyao Wang , Wei Zhao , Wojciech Matusik

The proliferation of mobile sensing technologies has enabled the study of various physiological and behavioural phenomena through unobtrusive data collection from smartphone sensors. This approach offers real-time insights into individuals'…

人机交互 · 计算机科学 2024-08-26 Songyan Teng , Tianyi Zhang , Simon D'Alfonso , Vassilis Kostakos

Spatial understanding is a fundamental aspect of computer vision and integral for human-level reasoning about images, making it an important component for grounded language understanding. While recent text-to-image synthesis (T2I) models…

计算机视觉与模式识别 · 计算机科学 2023-10-30 Tejas Gokhale , Hamid Palangi , Besmira Nushi , Vibhav Vineet , Eric Horvitz , Ece Kamar , Chitta Baral , Yezhou Yang

Despite remarkable progress in Text-to-Image models, many real-world applications require generating coherent image sets with diverse consistency requirements. Existing consistent methods often focus on a specific domain with specific…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Chengyou Jia , Xin Shen , Zhuohang Dang , Zhuohang Dang , Changliang Xia , Weijia Wu , Xinyu Zhang , Hangwei Qian , Ivor W. Tsang , Minnan Luo

The rapid evolution of lightweight consumer augmented reality (AR) smart glasses (a.k.a. optical see-through head-mounted displays) offers novel opportunities for learning, particularly through their unique capability to deliver multimodal…

人机交互 · 计算机科学 2025-07-22 Nuwan Janaka , Shengdong Zhao , Ashwin Ram , Ruoxin Sun , Sherisse Tan Jing Wen , Danae Li , David Hsu

We introduce NeuralOS, a neural framework that simulates graphical user interfaces (GUIs) of operating systems by directly predicting screen frames in response to user inputs such as mouse movements, clicks, and keyboard events. NeuralOS…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Luke Rivard , Sun Sun , Hongyu Guo , Wenhu Chen , Yuntian Deng

This paper introduces UI-TARS, a native GUI agent model that solely perceives the screenshots as input and performs human-like interactions (e.g., keyboard and mouse operations). Unlike prevailing agent frameworks that depend on heavily…

Screen user interfaces (UIs) and infographics, sharing similar visual language and design principles, play important roles in human communication and human-machine interaction. We introduce ScreenAI, a vision-language model that specializes…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Gilles Baechler , Srinivas Sunkara , Maria Wang , Fedir Zubach , Hassan Mansoor , Vincent Etter , Victor Cărbune , Jason Lin , Jindong Chen , Abhanshu Sharma

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM2) has shown remarkable performance in various…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Fu Rong , Meng Lan , Qian Zhang , Lefei Zhang

Most existing pre-trained language representation models (PLMs) are sub-optimal in sentiment analysis tasks, as they capture the sentiment information from word-level while under-considering sentence-level information. In this paper, we…

计算与语言 · 计算机科学 2022-10-20 Shuai Fan , Chen Lin , Haonan Li , Zhenghao Lin , Jinsong Su , Hang Zhang , Yeyun Gong , Jian Guo , Nan Duan

Screen recordings are becoming increasingly important as rich software artifacts that inform mobile application development processes. However, the amount of manual effort required to extract information from these graphical artifacts can…

Screen recordings of mobile applications are easy to obtain and capture a wealth of information pertinent to software developers (e.g., bugs or feature requests), making them a popular mechanism for crowdsourced app feedback. Thus, these…

Semantic segmentation, which aims to classify every pixel in an image, is a key task in machine perception, with many applications across robotics and autonomous driving. Due to the high dimensionality of this task, most existing approaches…

计算机视觉与模式识别 · 计算机科学 2023-10-04 Alex Zihao Zhu , Jieru Mei , Siyuan Qiao , Hang Yan , Yukun Zhu , Liang-Chieh Chen , Henrik Kretzschmar

With the development of multimodal reasoning models, Computer Use Agents (CUAs), akin to Jarvis from \textit{"Iron Man"}, are becoming a reality. GUI grounding is a core component for CUAs to execute actual actions, similar to mechanical…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Miaosen Zhang , Ziqiang Xu , Jialiang Zhu , Qi Dai , Kai Qiu , Yifan Yang , Chong Luo , Tianyi Chen , Justin Wagle , Tim Franklin , Baining Guo

Infographics are widely used to communicate information with a combination of text, icons, and data visualizations, but once exported as images their content is locked into pixels, making updates, localization, and reuse expensive. We…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Leonardo Gonzalez