中文
相关论文

相关论文: GestFormer: Multiscale Wavelet Pooling Transformer…

200 篇论文

Transformers combined with convolutional encoders have been recently used for hand gesture recognition (HGR) using micro-Doppler signatures. We propose a vision-transformer-based architecture for HGR with multi-antenna continuous-wave…

Mesh denoising, aimed at removing noise from input meshes while preserving their feature structures, is a practical yet challenging task. Despite the remarkable progress in learning-based mesh denoising methodologies in recent years, their…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Wenbo Zhao , Xianming Liu , Deming Zhai , Junjun Jiang , Xiangyang Ji

Hand gesture is one of the most important means of touchless communication between human and machines. There is a great interest for commanding electronic equipment in surgery rooms by hand gesture for reducing the time of surgery and the…

计算机视觉与模式识别 · 计算机科学 2017-10-24 Ebrahim Nasr-Esfahani , Nader Karimi , S. M. Reza Soroushmehr , M. Hossein Jafari , M. Amin Khorsandi , Shadrokh Samavi , Kayvan Najarian

Most current anthropomorphic robotic hands can realize part of the human hand functions, particularly for object grasping. However, due to the complexity of the human hand, few current designs target at daily object manipulations, even for…

机器人学 · 计算机科学 2021-04-22 Li Tian , Hanhui Li , Qifa Wang , Xuezeng Du , Jialin Tao , Jordan Sia Chong , Nadia Magnenat Thalmann , Jianmin Zheng

3D object detectors for point clouds often rely on a pooling-based PointNet to encode sparse points into grid-like voxels or pillars. In this paper, we identify that the common PointNet design introduces an information bottleneck that…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Zhaoqi Leng , Pei Sun , Tong He , Dragomir Anguelov , Mingxing Tan

The generalization of the Transformer architecture via MetaFormer has reshaped our understanding of its success in computer vision. By replacing self-attention with simpler token mixers, MetaFormer provides strong baselines for vision…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Ron Keuth , Paul Kaftan , Mattias P. Heinrich

Traditional vision-based material perception methods often experience substantial performance degradation under visually impaired conditions, thereby motivating the shift toward non-visual multimodal material perception. Despite this,…

机器学习 · 计算机科学 2025-11-26 Kailin Lyu , Long Xiao , Jianing Zeng , Junhao Dong , Xuexin Liu , Zhuojun Zou , Haoyue Yang , Lin Shu , Jie Hao

The recent introduction of depth cameras like Leap Motion Controller allows researchers to exploit the depth information to recognize hand gesture more robustly. This paper proposes a novel hand gesture recognition system with Leap Motion…

计算机视觉与模式识别 · 计算机科学 2017-11-15 Youchen Du , Shenglan Liu , Lin Feng , Menghui Chen , Jie Wu

Free-form gesture understanding is highly appealing for human-computer interaction, as it liberates users from the constraints of predefined gesture categories. However, the sole existing solution GestureGPT suffers from limited recognition…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Zhuoming Li , Aitong Liu , Mengxi Jia , Yubi Lu , Tengxiang Zhang , Changzhi Sun , Dell Zhang , Xuelong Li

We propose a fully automatic method for learning gestures on big touch devices in a potentially multi-user context. The goal is to learn general models capable of adapting to different gestures, user styles and hardware variations (e.g.…

机器学习 · 计算机科学 2018-02-28 Quentin Debard , Christian Wolf , Stéphane Canu , Julien Arné

Hand gestures form an intuitive means of interaction in Mixed Reality (MR) applications. However, accurate gesture recognition can be achieved only through state-of-the-art deep learning models or with the use of expensive sensors. Despite…

计算机视觉与模式识别 · 计算机科学 2019-04-23 Varun Jain , Gaurav Garg , Ramakrishna Perla , Ramya Hebbalaguppe

Gesture recognition and hand motion tracking are important tasks in advanced gesture based interaction systems. In this paper, we propose to apply a sliding windows filtering approach to sample the incoming streams of data from data gloves…

机器学习 · 计算机科学 2019-12-02 Sara Masoud , Bijoy Chowdhury , Young-Jun Son , Chieri Kubota , Russell Tronstad

Transformer-based models have achieved great success in various NLP, vision, and speech tasks. However, the core of Transformer, the self-attention mechanism, has a quadratic time and memory complexity with respect to the sequence length,…

计算与语言 · 计算机科学 2023-05-23 Chao-Hong Tan , Qian Chen , Wen Wang , Qinglin Zhang , Siqi Zheng , Zhen-Hua Ling

Many animals and humans process the visual field with a varying spatial resolution (foveated vision) and use peripheral processing to make eye movements and point the fovea to acquire high-resolution information about objects of interest.…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Aditya Jonnalagadda , William Yang Wang , B. S. Manjunath , Miguel P. Eckstein

Existing gesture interfaces only work with a fixed set of gestures defined either by interface designers or by users themselves, which introduces learning or demonstration efforts that diminish their naturalness. Humans, on the other hand,…

计算与语言 · 计算机科学 2024-11-05 Xin Zeng , Xiaoyu Wang , Tengxiang Zhang , Chun Yu , Shengdong Zhao , Yiqiang Chen

In human computer interaction, real-time detection and classification of dynamic hand gestures is challenging as: 1) the system must run in a real-time video stream and there is no noticeable lag in response after performing a gesture; 2)…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yinghao Qin , Tijana Timotijevic

Most existing ultra-high resolution (UHR) segmentation methods always struggle in the dilemma of balancing memory cost and local characterization accuracy, which are both taken into account in our proposed Guided Patch-Grouping Wavelet…

计算机视觉与模式识别 · 计算机科学 2023-07-07 Deyi Ji , Feng Zhao , Hongtao Lu

Acquiring spatio-temporal states of an action is the most crucial step for action classification. In this paper, we propose a data level fusion strategy, Motion Fused Frames (MFFs), designed to fuse motion information into static images as…

计算机视觉与模式识别 · 计算机科学 2018-04-27 Okan Köpüklü , Neslihan Köse , Gerhard Rigoll

Skeleton-based human action recognition leverages sequences of human joint coordinates to identify actions performed in videos. Owing to the intrinsic spatiotemporal structure of skeleton data, Graph Convolutional Networks (GCNs) have been…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yusen Peng , Alper Yilmaz

We present a method for gesture detection and localisation based on multi-scale and multi-modal deep learning. Each visual modality captures spatial information at a particular spatial scale (such as motion of the upper body or a hand), and…

计算机视觉与模式识别 · 计算机科学 2015-07-21 Natalia Neverova , Christian Wolf , Graham W. Taylor , Florian Nebout