中文
相关论文

相关论文: Using Keypoint Matching and Interactive Self Atten…

200 篇论文

Product images are the most impressing medium of customer interaction on the product detail pages of e-commerce websites. Millions of products are onboarded on to webstore catalogues daily and maintaining a high quality bar for a product's…

计算机视觉与模式识别 · 计算机科学 2023-01-20 Saurabh Sharma , Faizan Ahemad

Visual pointing, which aims to localize a target by predicting its coordinates on an image, has emerged as an important problem in the realm of vision-language models (VLMs). Despite its broad applicability, recent benchmarks show that…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Wenjie Yang , Zengfeng Huang

Computer vision algorithms are being implemented across a breadth of industries to enable technological innovations. In this paper, we study the problem of computer vision based customer tracking in retail industry. To this end, we…

计算机视觉与模式识别 · 计算机科学 2020-07-03 Aibek Musaev , Jiangping Wang , Liang Zhu , Cheng Li , Yi Chen , Jialin Liu , Wanqi Zhang , Juan Mei , De Wang

Set-of-Mark (SoM) Prompting unleashes the visual grounding capability of GPT-4V, by enabling the model to associate visual objects with tags inserted on the image. These tags, marked with alphanumerics, can be indexed via text tokens for…

计算机视觉与模式识别 · 计算机科学 2025-01-22 An Yan , Zhengyuan Yang , Junda Wu , Wanrong Zhu , Jianwei Yang , Linjie Li , Kevin Lin , Jianfeng Wang , Julian McAuley , Jianfeng Gao , Lijuan Wang

Assistive solutions for a better shopping experience can improve the quality of life of people, in particular also of visually impaired shoppers. We present a system that visually recognizes the fine-grained product classes of items on a…

计算机视觉与模式识别 · 计算机科学 2015-10-15 Marian George , Dejan Mircic , Gábor Sörös , Christian Floerkemeier , Friedemann Mattern

The construction industry represents a major sector in terms of resource consumption. Recycled construction material has high reuse potential, but quality monitoring of the aggregates is typically still performed with manual methods.…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Yu Zhou , Pelle Thielmann , Ayush Chamoli , Bruno Mirbach , Didier Stricker , Jason Rambach

An examination of object recognition challenge leaderboards (ILSVRC, PASCAL-VOC) reveals that the top-performing classifiers typically exhibit small differences amongst themselves in terms of error rate/mAP. To better differentiate the top…

计算机视觉与模式识别 · 计算机科学 2016-11-28 Ravi Kiran Sarvadevabhatla , Shanthakumar Venkatraman , R. Venkatesh Babu

We present methods to estimate the physical properties of household containers and their fillings manipulated by humans. We use a lightweight, pre-trained convolutional neural network with coordinate attention as a backbone model of the…

计算机视觉与模式识别 · 计算机科学 2022-03-03 Hengyi Wang , Chaoran Zhu , Ziyin Ma , Changjae Oh

We propose a novel solution for the cashier problem. Current cashier system/Point of Sale (POS) terminals can be inefficient, cumbersome and time-consuming for the users. There is a need for a solution dependent on modern technology and…

机器学习 · 计算机科学 2020-11-13 Farid Khan

The growing ubiquity of Extended Reality (XR) is driving Conversational Recommendation Systems (CRS) toward visually immersive experiences. We formalize this paradigm as Immersive CRS (ICRS), where recommended items are highlighted directly…

信息检索 · 计算机科学 2026-04-14 Jiazhou Liang , Yifan Simon Liu , David Guo , Minqi Sun , Yilun Jiang , Scott Sanner

With the rapid advancement of deep learning technologies, computer vision has shown immense potential in retail automation. This paper presents a novel self-checkout system for retail based on an improved YOLOv10 network, aimed at enhancing…

计算机视觉与模式识别 · 计算机科学 2024-08-19 Lianghao Tan , Shubing Liu , Jing Gao , Xiaoyi Liu , Linyue Chu , Huangqi Jiang

Retail Product Image Classification is an important Computer Vision and Machine Learning problem for building real world systems like self-checkout stores and automated retail execution evaluation. In this work, we present various tricks to…

计算机视觉与模式识别 · 计算机科学 2020-01-14 Muktabh Mayank Srivastava

Energy disaggregation (a.k.a nonintrusive load monitoring, NILM), a single-channel blind source separation problem, aims to decompose the mains which records the whole house electricity consumption into appliance-wise readings. This problem…

应用统计 · 统计学 2018-01-19 Chaoyun Zhang , Mingjun Zhong , Zongzuo Wang , Nigel Goddard , Charles Sutton

In the realm of point cloud scene understanding, particularly in indoor scenes, objects are arranged following human habits, resulting in objects of certain semantics being closely positioned and displaying notable inter-object…

计算机视觉与模式识别 · 计算机科学 2024-04-12 Yanhao Wu , Tong Zhang , Wei Ke , Congpei Qiu , Sabine Susstrunk , Mathieu Salzmann

We propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our knowledge, our system is among the first to utilize the camera…

机器人学 · 计算机科学 2022-07-14 Nathaniel Merrill , Yuliang Guo , Xingxing Zuo , Xinyu Huang , Stefan Leutenegger , Xi Peng , Liu Ren , Guoquan Huang

Face anti-spoofing (FAS) plays a vital role in securing the face recognition systems from presentation attacks. Most existing FAS methods capture various cues (e.g., texture, depth and reflection) to distinguish the live faces from the…

计算机视觉与模式识别 · 计算机科学 2020-07-07 Zitong Yu , Xiaobai Li , Xuesong Niu , Jingang Shi , Guoying Zhao

Non-maximum suppression is an integral part of the object detection pipeline. First, it sorts all detection boxes on the basis of their scores. The detection box M with the maximum score is selected and all other detection boxes with a…

计算机视觉与模式识别 · 计算机科学 2017-08-09 Navaneeth Bodla , Bharat Singh , Rama Chellappa , Larry S. Davis

We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Boyuan Sun , Bowen Yin , Yuanming Li , Xihan Wei , Qibin Hou

This paper presents a comprehensive pipeline for recognizing objects targeted by human pointing gestures using RGB images. As human-robot interaction moves toward more intuitive interfaces, the ability to identify targets of non-verbal…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Lukáš Hajdúch , Viktor Kocur

Salient object detection has achieved great improvement by using the Fully Convolution Network (FCN). However, the FCN-based U-shape architecture may cause the dilution problem in the high-level semantic information during the up-sample…

计算机视觉与模式识别 · 计算机科学 2020-05-01 Guangyu Ren , Tianhong Dai , Panagiotis Barmpoutis , Tania Stathaki