中文
相关论文

相关论文: From Web to Pixels: Bringing Agentic Search into V…

200 篇论文

All instance perception tasks aim at finding certain objects specified by some queries such as category names, language expressions, and target annotations, but this complete field has been split into multiple independent subtasks. In this…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Bin Yan , Yi Jiang , Jiannan Wu , Dong Wang , Ping Luo , Zehuan Yuan , Huchuan Lu

Imagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely explanation for the person's discomfort, thereby alleviating the…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Nitzan Bitton-Guetta , Aviv Slobodkin , Aviya Maimon , Eliya Habba , Royi Rassin , Yonatan Bitton , Idan Szpektor , Amir Globerson , Yuval Elovici

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images,…

计算机视觉与模式识别 · 计算机科学 2017-08-09 Linjie Yang , Kevin Tang , Jianchao Yang , Li-Jia Li

Invisible watermarking is essential for tracing the provenance of digital content. However, training state-of-the-art models remains notoriously difficult, with current approaches often struggling to balance robustness against true…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Tomáš Souček , Pierre Fernandez , Hady Elsahar , Sylvestre-Alvise Rebuffi , Valeriu Lacatusu , Tuan Tran , Tom Sander , Alexandre Mourachko

Object recognition in humans depends primarily on shape cues. We have developed a new approach to measuring the shape recognition performance of a vision system based on nearest neighbor view matching within the system's embedding space.…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Jong Woo Nam , Amanda S. Rios , Bartlett W. Mel

Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation. In this work, we introduce…

计算与语言 · 计算机科学 2022-03-29 Yingshan Chang , Mridu Narang , Hisami Suzuki , Guihong Cao , Jianfeng Gao , Yonatan Bisk

Following the gaze of other people and analyzing the target they are looking at can help us understand what they are thinking, and doing, and predict the actions that may follow. Existing methods for gaze following struggle to perform well…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Feiyang Liu , Dan Guo , Jingyuan Xu , Zihao He , Shengeng Tang , Kun Li , Meng Wang

While large language models have become the prevailing approach for agentic reasoning and planning, their success in symbolic domains does not readily translate to the physical world. Spatial intelligence, the ability to perceive 3D…

机器学习 · 计算机科学 2026-02-03 Gloria Felicia , Nolan Bryant , Handi Putra , Ayaan Gazali , Eliel Lobo , Esteban Rojas

The eye-tracking video saliency prediction (VSP) task and video salient object detection (VSOD) task both focus on the most attractive objects in video and show the result in the form of predictive heatmaps and pixel-level saliency masks,…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Qi Qin , Runmin Cong , Gen Zhan , Yiting Liao , Sam Kwong

The ImageNet Large Scale Visual Recognition Challenge is a benchmark in object category classification and detection on hundreds of object categories and millions of images. The challenge has been run annually from 2010 to present,…

For web agents to be practically useful, they must adapt to the continuously evolving web environment characterized by frequent updates to user interfaces and content. However, most existing benchmarks only capture the static aspects of the…

计算与语言 · 计算机科学 2024-07-17 Yichen Pan , Dehan Kong , Sida Zhou , Cheng Cui , Yifei Leng , Bing Jiang , Hangyu Liu , Yanyi Shang , Shuyan Zhou , Tongshuang Wu , Zhengyang Wu

While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they often fail to grasp implicit user intentions. Although…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Jun He , Junyan Ye , Zilong Huang , Dongzhi Jiang , Chenjue Zhang , Leqi Zhu , Renrui Zhang , Xiang Zhang , Weijia Li

Graphical User Interface (GUI) agents have the potential to assist users in interacting with complex software (e.g., PowerPoint, Photoshop). While prior research has primarily focused on automating user actions through clicks and…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Saelyne Yang , Jaesang Yu , Yi-Hao Peng , Kevin Qinghong Lin , Jae Won Cho , Yale Song , Juho Kim

With the rapid development of generative models, discerning AI-generated content has evoked increasing attention from both industry and academia. In this paper, we conduct a sanity check on "whether the task of AI-generated image detection…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Shilin Yan , Ouxiang Li , Jiayin Cai , Yanbin Hao , Xiaolong Jiang , Yao Hu , Weidi Xie

We present the 2017 WebVision Challenge, a public image recognition challenge designed for deep learning based on web images without instance-level human annotation. Following the spirit of previous vision challenges, such as ILSVRC,…

计算机视觉与模式识别 · 计算机科学 2017-05-17 Wen Li , Limin Wang , Wei Li , Eirikur Agustsson , Jesse Berent , Abhinav Gupta , Rahul Sukthankar , Luc Van Gool

With the emergence of Virtual and Mixed Reality (XR) devices, eye tracking has received significant attention in the computer vision community. Eye gaze estimation is a crucial component in XR -- enabling energy efficient rendering,…

计算机视觉与模式识别 · 计算机科学 2020-03-20 Zhengyang Wu , Srivignesh Rajendran , Tarrence van As , Joelle Zimmermann , Vijay Badrinarayanan , Andrew Rabinovich

Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agentic models. In this work, we revisit the geolocalization…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Yikun Wang , Zuyan Liu , Ziyi Wang , Han Hu , Pengfei Liu , Yongming Rao

This paper introduces the first two pixel retrieval benchmarks. Pixel retrieval is segmented instance retrieval. Like semantic segmentation extends classification to the pixel level, pixel retrieval is an extension of image retrieval and…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Guoyuan An , Woo Jae Kim , Saelyne Yang , Rong Li , Yuchi Huo , Sung-Eui Yoon

Visual object localization is the key step in a series of object detection tasks. In the literature, high localization accuracy is achieved with the mainstream strongly supervised frameworks. However, such methods require object-level…

计算机视觉与模式识别 · 计算机科学 2021-08-10 Yi-Geng Hong , Hui-Chu Xiao , Wan-Lei Zhao

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Ali Athar , Xueqing Deng , Liang-Chieh Chen