中文
相关论文

相关论文: Exploring State Change Capture of Heterogeneous Ba…

200 篇论文

We study the 3D object understanding task for manipulating everyday objects with different material properties (diffuse, specular, transparent and mixed). Existing monocular and RGB-D methods suffer from scale ambiguity due to missing or…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Chuanrui Zhang , Yonggen Ling , Minglei Lu , Minghan Qin , Haoqian Wang

Finding and localizing the conceptual changes in two scenes in terms of the presence or removal of objects in two images belonging to the same scene at different times in special care applications is of great significance. This is mainly…

计算机视觉与模式识别 · 计算机科学 2023-01-02 Ali Atghaei , Ehsan Rahnama , Kiavash Azimi , Hassan Shahbazi

Ego hand gestures can be used as an interface in AR and VR environments. While the context of an image is important for tasks like scene understanding, object recognition, image caption generation and activity recognition, it plays a…

计算机视觉与模式识别 · 计算机科学 2019-09-19 Tejo Chalasani , Aljosa Smolic

Vision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a particular issue in video recognition, where models are…

计算机视觉与模式识别 · 计算机科学 2023-08-28 Matthew Dutson , Yin Li , Mohit Gupta

Outside-in multi-camera perception is increasingly important in indoor environments, where networks of static cameras must support multi-target tracking under occlusion and heterogeneous viewpoints. We evaluate Sparse4D, a query-based…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Ethan Anderson , Justin Silva , Kyle Zheng , Sameer Pusegaonkar , Yizhou Wang , Zheng Tang , Sujit Biswas

There is a recent growing interest in applying Deep Learning techniques to tabular data, in order to replicate the success of other Artificial Intelligence areas in this structured domain. Specifically interesting is the case in which…

机器学习 · 计算机科学 2025-05-06 Simone Luetto , Fabrizio Garuti , Enver Sangineto , Lorenzo Forni , Rita Cucchiara

Existing datasets for 3D hand-object interaction are limited either in the data cardinality, data variations in interaction scenarios, or the quality of annotations. In this work, we present a comprehensive new training dataset for…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Woojin Cho , Jihyun Lee , Minjae Yi , Minje Kim , Taeyun Woo , Donghwan Kim , Taewook Ha , Hyokeun Lee , Je-Hwan Ryu , Woontack Woo , Tae-Kyun Kim

The crucial step for localization is to match the current observation to the map. When the two sensor modalities are significantly different, matching becomes challenging. In this paper, we present an end-to-end deep phase correlation…

计算机视觉与模式识别 · 计算机科学 2020-11-03 Zexi Chen , Xuecheng Xu , Yue Wang , Rong Xiong

Recent convolutional neural network (CNN) development continues to advance the state-of-the-art model accuracy for various applications. However, the enhanced accuracy comes at the cost of substantial memory bandwidth and storage…

计算机视觉与模式识别 · 计算机科学 2021-07-20 Hsu-Hsun Chin , Ren-Song Tsay , Hsin-I Wu

Real-world robots localize objects from natural-language instructions while scenes around them keep changing. Yet most of the existing 3D visual grounding (3DVG) method still assumes a reconstructed and up-to-date point cloud, an assumption…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Miao Hu , Zhiwei Huang , Tai Wang , Jiangmiao Pang , Dahua Lin , Nanning Zheng , Runsen Xu

We study the task of establishing object-level visual correspondence across different viewpoints in videos, focusing on the challenging egocentric-to-exocentric and exocentric-to-egocentric scenarios. We propose a simple yet effective…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Shannan Yan , Leqi Zheng , Keyu Lv , Jingchen Ni , Hongyang Wei , Jiajun Zhang , Guangting Wang , Jing Lyu , Chun Yuan , Fengyun Rao

Monocular egocentric 3D human motion capture is a challenging and actively researched problem. Existing methods use synchronously operating visual sensors (e.g. RGB cameras) and often fail under low lighting and fast motions, which can be…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Christen Millerdurai , Hiroyasu Akada , Jian Wang , Diogo Luvizon , Christian Theobalt , Vladislav Golyanik

Ground Terrain Recognition is a difficult task as the context information varies significantly over the regions of a ground terrain image. In this paper, we propose a novel approach towards ground-terrain recognition via modeling the…

计算机视觉与模式识别 · 计算机科学 2020-10-28 Shuvozit Ghose , Pinaki Nath Chowdhury , Partha Pratim Roy , Umapada Pal

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

When a robot encounters a novel object, how should it respond$\unicode{x2014}$what data should it collect$\unicode{x2014}$so that it can find the object in the future? In this work, we present a method for learning image features of an…

机器人学 · 计算机科学 2024-10-16 Allison Pinosky , Todd D. Murphey

To support the modern machine-type communications, a crucial task during the random access phase is device activity detection, which is to detect the active devices from a large number of potential devices based on the received signal at…

信息论 · 计算机科学 2021-12-21 Yang Li , Zhilin Chen , Yunqi Wang , Chenyang Yang , Yik-Chung Wu

Touch contact and pressure are essential for understanding how humans interact with and manipulate objects, insights which can significantly benefit applications in mixed reality and robotics. However, estimating these interactions from an…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Yiming Zhao , Taein Kwon , Paul Streli , Marc Pollefeys , Christian Holz

Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Yilin Wen , Hao Pan , Lei Yang , Jia Pan , Taku Komura , Wenping Wang

Simultaneous object recognition and pose estimation are two key functionalities for robots to safely interact with humans as well as environments. Although both object recognition and pose estimation use visual input, most state-of-the-art…

机器人学 · 计算机科学 2023-04-10 Tommaso Parisotto , Subhaditya Mukherjee , Hamidreza Kasaei

Manual annotations of temporal bounds for object interactions (i.e. start and end times) are typical training input to recognition, localization and detection algorithms. For three publicly available egocentric datasets, we uncover…

计算机视觉与模式识别 · 计算机科学 2017-07-27 Davide Moltisanti , Michael Wray , Walterio Mayol-Cuevas , Dima Damen