English
Related papers

Related papers: How to Tune Autofocals: A Comparative Study of Adv…

200 papers

Adaptive Virtual Reality (VR) systems have the potential to enhance training and learning experiences by dynamically responding to users' cognitive states. This research investigates how eye tracking and heart rate variability (HRV) can be…

Human-Computer Interaction · Computer Science 2025-04-10 Mahsa Nasri

We study how to transfer representations pretrained on source tasks to target tasks in visual percept based RL. We analyze two popular approaches: freezing or finetuning the pretrained representations. Empirical studies on a set of popular…

Machine Learning · Computer Science 2023-02-14 Sébastien M. R. Arnold , Fei Sha

Despite advances in Vision-Language-Action (VLA) models, robotic manipulation struggles with fine-grained tasks because current models lack mechanisms for active visual attention allocation. Human gaze naturally encodes intent, planning,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Anupam Pani , Yanchao Yang

Robust trajectory planning under camera viewpoint changes is important for scalable end-to-end autonomous driving. However, existing models often depend heavily on the camera viewpoints seen during training. We investigate an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Hiroki Hashimoto , Hiromichi Goto , Hiroyuki Sugai , Hiroshi Kera , Kazuhiko Kawamoto

Recent advancements in Vision-Language (VL) models have sparked interest in their deployment on edge devices, yet challenges in handling diverse visual modalities, manual annotation, and computational constraints remain. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Kaiwen Cai , Zhekai Duan , Gaowen Liu , Charles Fleming , Chris Xiaoxuan Lu

It has already been observed that audio-visual embedding is more robust than uni-modality embedding for person verification. Here, we proposed a novel audio-visual strategy that considers aggregators from a fusion perspective. First, we…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 Peiwen Sun , Shanshan Zhang , Zishan Liu , Yougen Yuan , Taotao Zhang , Honggang Zhang , Pengfei Hu

To mitigate the pain of manually tuning hyperparameters of deep neural networks, automated machine learning (AutoML) methods have been developed to search for an optimal set of hyperparameters in large combinatorial search spaces. However,…

Human-Computer Interaction · Computer Science 2020-10-13 Heungseok Park , Yoonsoo Nam , Ji-Hoon Kim , Jaegul Choo

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between audio and visual…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-19 Abhinav Shukla , Stavros Petridis , Maja Pantic

Weakly supervised object localization (WSOL) aims at predicting object locations in an image using only image-level category labels. Common challenges that image classification models encounter when localizing objects are, (a) they tend to…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Saurav Gupta , Sourav Lakhotia , Abhay Rawat , Rahul Tallamraju

Eye tracking is an important tool with a wide range of applications in Virtual, Augmented, and Mixed Reality (VR/AR/MR) technologies. State-of-the-art eye tracking methods are either reflection-based and track reflections of sparse point…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Tianfu Wang , Jiazhang Wang , Oliver Cossairt , Florian Willomitzer

Bionic vision uses neuroprostheses to restore useful vision to people living with incurable blindness. However, a major outstanding challenge is predicting what people 'see' when they use their devices. The limited field of view of current…

Human-Computer Interaction · Computer Science 2022-03-14 Justin Kasowski , Michael Beyeler

This study evaluates the usage of virtual reality (VR) technologies as a teaching tool in oral placement therapy, a subset of speech therapy. The researcher distributed instructional videos using traditional lecture and modified…

Human-Computer Interaction · Computer Science 2024-04-25 Daniel E. Killough

The visual focus of attention (VFOA) has been recognized as a prominent conversational cue. We are interested in estimating and tracking the VFOAs associated with multi-party social interactions. We note that in this type of situations the…

Computer Vision and Pattern Recognition · Computer Science 2018-12-21 Benoît Massé , Silèye Ba , Radu Horaud

Vision sensors are versatile and can capture a wide range of visual cues, such as color, texture, shape, and depth. This versatility, along with the relatively inexpensive availability of machine vision cameras, played an important role in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Muhammad Z. Alam , Zeeshan Kaleem , Sousso Kelouwani

AutoFocus-IL is a simple yet effective method to improve data efficiency and generalization in visual imitation learning by guiding policies to attend to task-relevant features rather than distractors and spurious correlations. Although…

Robotics · Computer Science 2025-11-26 Litian Gong , Fatemeh Bahrani , Yutai Zhou , Amin Banayeeanzade , Jiachen Li , Erdem Bıyık

We introduce to VR a novel imperceptible gaze guidance technique from a recent discovery that human gaze can be attracted to a cue that contrasts from the background in its perceptually non-distinctive ocularity, defined as the relative…

Human-Computer Interaction · Computer Science 2024-12-13 Virmarie Maquiling , Li Zhaoping , Enkelejda Kasneci

The integration of medical imaging and clinical text has enabled the emergence of generalist artificial intelligence (AI) systems for healthcare. However, pervasive biases, such as imbalanced disease prevalence, skewed anatomical region…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Cheng Li , Weijian Huang , Jiarun Liu , Hao Yang , Qi Yang , Song Wu , Ye Li , Hairong Zheng , Shanshan Wang

Robotic manipulators are increasingly used to assist individuals with mobility impairments in object retrieval. However, the predominant joystick-based control interfaces can be challenging due to high precision requirements and unintuitive…

Robotics · Computer Science 2025-10-28 Zitiantao Lin , Yongpeng Sang , Yang Ye

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

Sound · Computer Science 2023-09-29 R. Gnana Praveen , Jahangir Alam

Real-world vision-language applications demand varying levels of perceptual granularity. However, most existing visual large language models (VLLMs), such as LLaVA, pre-assume a fixed resolution for downstream tasks, which leads to subpar…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Weiqing Luo , Zhen Tan , Yifan Li , Xinyu Zhao , Kwonjoon Lee , Behzad Dariush , Tianlong Chen