中文
相关论文

相关论文: Versatile User Identification in Extended Reality …

200 篇论文

While accessibility (a11y) guidelines exist for 3D games and virtual worlds, their applicability to extended reality (XR)'s unique interaction paradigms (e.g., spatial tracking, kinesthetic interactions) remains unexplored. XR practitioners…

人机交互 · 计算机科学 2026-02-23 Daniel Killough , Tiger F. Ji , Kexin Zhang , Yaxin Hu , Yu Huang , Ruofei Du , Yuhang Zhao

Causal representation learning (CRL) offers the promise of uncovering the underlying causal model by which observed data was generated, but the practical applicability of existing methods remains limited by the strong assumptions required…

Extended reality (XR) systems, which consist of virtual reality (VR), augmented reality (AR), and mixed reality (XR), offer a transformative interface for immersive, multi-modal, and embodied human-computer interaction. In this paper, we…

We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introduces two lightweight adapters to ensure identity fidelity.…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Shen Sang , Tiancheng Zhi , Tianpei Gu , Jing Liu , Linjie Luo

Video-based person re-identification (video re-ID) has lately fascinated growing attention due to its broad practical applications in various areas, such as surveillance, smart city, and public safety. Nevertheless, video re-ID is quite…

计算机视觉与模式识别 · 计算机科学 2024-10-25 Khawar Islam

Persona prompting has been widely adopted to steer large language models (LLMs) behavior and improve their instruction performance by assigning specific characters. However, identifying an optimal persona is time-consuming, and its impact…

计算与语言 · 计算机科学 2026-04-13 Jihwan Oh , Soowon Oh , Murad Aghazada , Minchan Jeong , Sungnyun Kim , Se-Young Yun

Long-Term Person Re-Identification (LT-ReID) has become increasingly crucial in computer vision and biometrics. In this work, we aim to extend LT-ReID beyond pedestrian recognition to include a wider range of real-world human activities…

计算机视觉与模式识别 · 计算机科学 2023-09-22 Feng Liu , Minchul Kim , ZiAng Gu , Anil Jain , Xiaoming Liu

Recent studies show that vision models pre-trained in generic visual learning tasks with large-scale data can provide useful feature representations for a wide range of visual perception problems. However, few attempts have been made to…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Feng Lu , Lijun Zhang , Xiangyuan Lan , Shuting Dong , Yaowei Wang , Chun Yuan

Deep reinforcement learning (deep RL) has achieved superior performance in complex sequential tasks by learning directly from image input. A deep neural network is used as a function approximator and requires no specific state information.…

机器学习 · 计算机科学 2018-12-27 Xi Chen , Caylin Hickey

Pretraining is a common technique in deep learning for increasing performance and reducing training time, with promising experimental results in deep reinforcement learning (RL). However, pretraining requires a relevant dataset for…

机器学习 · 计算机科学 2021-10-07 Saurav Kadavath , Samuel Paradis , Brian Yao

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot…

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets contain images of people and hence are accompanied with…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Howard Zhong , Samarth Mishra , Donghyun Kim , SouYoung Jin , Rameswar Panda , Hilde Kuehne , Leonid Karlinsky , Venkatesh Saligrama , Aude Oliva , Rogerio Feris

As immersive technologies enable unique, multimodal interaction methods, developers must also use tailored methods to support user accessibility, distinct from traditional software practices. We interviewed 25 industry extended reality (XR)…

人机交互 · 计算机科学 2025-08-11 Daniel Killough , Tiger F. Ji , Kexin Zhang , Yaxin Hu , Yu Huang , Ruofei Du , Yuhang Zhao

Most sports visualizations rely on a combination of spatial, highly temporal, and user-centric data, making sports a challenging target for visualization. Emerging technologies, such as augmented and mixed reality (AR/XR), have brought…

人机交互 · 计算机科学 2022-11-16 Tica Lin , Zhutian Chen , Johanna Beyer , Yincai Wu , Hanspeter Pfister , Yalong Yang

Extreme far-distance video person re-identification (ReID) is particularly challenging due to scale compression, resolution degradation, motion blur, and aerial-ground viewpoint mismatch. As camera altitude and subject distance increase,…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Ashwat Rajbhandari , Bharatesh Chakravarthi

Building on existing approaches, we revisit Human-in-the-Loop Object Retrieval, a task that consists of iteratively retrieving images containing objects of a class-of-interest, specified by a user-provided query. Starting from a large…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Kawtar Zaher , Olivier Buisson , Alexis Joly

Recent advances in both machine learning and Internet-of-Things have attracted attention to automatic Activity Recognition, where users wear a device with sensors and their outputs are mapped to a predefined set of activities. However, few…

机器学习 · 计算机科学 2019-08-20 Taku Yamagata , Raúl Santos-Rodríguez , Ryan McConville , Atis Elsts

We consider a sequence of related multivariate time series learning tasks, such as predicting failures for different instances of a machine from time series of multi-sensor data, or activity recognition tasks over different individuals from…

机器学习 · 计算机科学 2022-03-15 Vibhor Gupta , Jyoti Narwariya , Pankaj Malhotra , Lovekesh Vig , Gautam Shroff

This paper presents a novel approach to solve simultaneously the problems of human activity recognition and whole-body motion and dynamics prediction for real-time applications. Starting from the dynamics of human motion and motor system…

机器人学 · 计算机科学 2023-03-15 Kourosh Darvish , Serena Ivaldi , Daniele Pucci

Human behavior expression and experience are inherently multi-modal, and characterized by vast individual and contextual heterogeneity. To achieve meaningful human-computer and human-robot interactions, multi-modal models of the users…

机器学习 · 计算机科学 2019-06-10 Ognjen Rudovic , Meiru Zhang , Bjorn Schuller , Rosalind W. Picard