中文
相关论文

相关论文: QORT-Former: Query-optimized Real-time Transformer…

200 篇论文

We propose a new dataset and a novel approach to learning hand-object interaction priors for hand and articulated object pose estimation. We first collect a dataset using visual teleoperation, where the human operator can directly play…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Zehao Zhu , Jiashun Wang , Yuzhe Qin , Deqing Sun , Varun Jampani , Xiaolong Wang

Facial expression recognition (FER) is a challenging topic in artificial intelligence. Recently, many researchers have attempted to introduce Vision Transformer (ViT) to the FER task. However, ViT cannot fully utilize emotional features…

计算机视觉与模式识别 · 计算机科学 2023-03-15 Yu Zhou , Liyuan Guo , Lianghai Jin

Interaction intention anticipation aims to jointly predict future hand trajectories and interaction hotspots. Existing research often treated trajectory forecasting and interaction hotspots prediction as separate tasks or solely considered…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Zichen Zhang , Hongchen Luo , Wei Zhai , Yang Cao , Yu Kang

Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of…

计算机视觉与模式识别 · 计算机科学 2022-10-14 Jian Wang , Chenhui Gou , Qiman Wu , Haocheng Feng , Junyu Han , Errui Ding , Jingdong Wang

Multi-agent collaborative perception enhances each agent perceptual capabilities by sharing sensing information to cooperatively perform robot perception tasks. This approach has proven effective in addressing challenges such as sensor…

机器学习 · 计算机科学 2025-07-02 Rujia Wang , Xiangbo Gao , Hao Xiang , Runsheng Xu , Zhengzhong Tu

Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Wenhao Li , Mengyuan Liu , Hong Liu , Pichao Wang , Jialun Cai , Nicu Sebe

Recently, fully-transformer architectures have replaced the defacto convolutional architecture for the 3D human pose estimation task. In this paper we propose \textbf{\textit{ConvFormer}}, a novel convolutional transformer that leverages a…

计算机视觉与模式识别 · 计算机科学 2023-04-06 Alec Diaz-Arias , Dmitriy Shin

Robotic arms require precise, task-aware trajectory planning, yet sequence models that ignore motion structure often yield invalid or inefficient executions. We present a Path-based Transformer that encodes robot motion with a 3-grid…

机器人学 · 计算机科学 2025-10-24 Ahmed Alanazi , Duy Ho , Yugyung Lee

We construct the first markerless deformable interaction dataset recording interactive motions of the hands and deformable objects, called HMDO (Hand Manipulation with Deformable Objects). With our built multi-view capture system, it…

计算机视觉与模式识别 · 计算机科学 2023-01-19 Wei Xie , Zhipeng Yu , Zimeng Zhao , Binghui Zuo , Yangang Wang

Transformer-based approaches have been successfully proposed for 3D human pose estimation (HPE) from 2D pose sequence and achieved state-of-the-art (SOTA) performance. However, current SOTAs have difficulties in modeling spatial-temporal…

计算机视觉与模式识别 · 计算机科学 2023-01-19 Xiaoye Qian , Youbao Tang , Ning Zhang , Mei Han , Jing Xiao , Ming-Chun Huang , Ruei-Sung Lin

Transformers are among the state of the art for many tasks in speech, vision, and natural language processing, among others. Self-attentions, which are crucial contributors to this performance have quadratic computational complexity, which…

计算与语言 · 计算机科学 2022-12-21 Roshan Sharma , Bhiksha Raj

Deformable objects manipulation can benefit from representations that seamlessly integrate vision and touch while handling occlusions. In this work, we present a novel approach for, and real-world demonstration of, multimodal visuo-tactile…

机器人学 · 计算机科学 2022-10-10 Youngsun Wi , Andy Zeng , Pete Florence , Nima Fazeli

For 3D object manipulation, methods that build an explicit 3D representation perform better than those relying only on camera images. But using explicit 3D representations like voxels comes at large computing cost, adversely affecting…

机器人学 · 计算机科学 2023-06-27 Ankit Goyal , Jie Xu , Yijie Guo , Valts Blukis , Yu-Wei Chao , Dieter Fox

3D hand-object pose estimation is the key to the success of many computer vision applications. The main focus of this task is to effectively model the interaction between the hand and an object. To this end, existing works either rely on…

计算机视觉与模式识别 · 计算机科学 2023-01-09 Rong Wang , Wei Mao , Hongdong Li

Human action recognition has recently become one of the popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in the task of video…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Thanh-Dat Truong , Quoc-Huy Bui , Chi Nhan Duong , Han-Seok Seo , Son Lam Phung , Xin Li , Khoa Luu

Tracking the pose of an object while it is being held and manipulated by a robot hand is difficult for vision-based methods due to significant occlusions. Prior works have explored using contact feedback and particle filters to localize…

机器人学 · 计算机科学 2020-11-09 Jacky Liang , Ankur Handa , Karl Van Wyk , Viktor Makoviychuk , Oliver Kroemer , Dieter Fox

Real-time recognition and prediction of surgical activities are fundamental to advancing safety and autonomy in robot-assisted surgery. This paper presents a multimodal transformer architecture for real-time recognition and prediction of…

机器人学 · 计算机科学 2024-10-27 Keshara Weerasinghe , Seyed Hamid Reza Roodabeh , Kay Hutchinson , Homa Alemzadeh

Sequential DeepFake detection is an emerging task that predicts the manipulation sequence in order. Existing methods typically formulate it as an image-to-sequence problem, employing conventional Transformer architectures. However, these…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Yunfei Li , Yuezun Li , Baoyuan Wu , Junyu Dong , Guopu Zhu , Siwei Lyu

Transformer-based models have achieved top performance on major video recognition benchmarks. Benefiting from the self-attention mechanism, these models show stronger ability of modeling long-range dependencies compared to CNN-based models.…

计算机视觉与模式识别 · 计算机科学 2022-08-26 Rui Wang , Zuxuan Wu , Dongdong Chen , Yinpeng Chen , Xiyang Dai , Mengchen Liu , Luowei Zhou , Lu Yuan , Yu-Gang Jiang

Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Yilin Wen , Hao Pan , Lei Yang , Jia Pan , Taku Komura , Wenping Wang