English
Related papers

Related papers: Spatio-Temporal Mixed and Augmented Reality Experi…

200 papers

Projected augmented reality, also called projection mapping or video mapping, is a form of augmented reality that uses projected light to directly augment 3D surfaces, as opposed to using pass-through screens or headsets. The value of…

Graphics · Computer Science 2020-01-03 Brittany Factura , Laura LaPerche , Phil Reyneri , Brett Jones , Kevin Karsch

A fundamental aspect for building intelligent autonomous robots that can assist humans in their daily lives is the construction of rich environmental representations. While advances in semantic scene representations have enriched robotic…

Robotics · Computer Science 2026-02-17 Phuoc Nguyen , Francesco Verdoja , Ville Kyrki

This paper introduces a novel framework for image and video demoir\'eing by integrating Maximum A Posteriori (MAP) estimation with advanced deep learning techniques. Demoir\'eing addresses inherently nonlinear degradation processes, which…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Liangyan Li , Yimo Ning , Kevin Le , Wei Dong , Yunzhe Li , Jun Chen , Xiaohong Liu

The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either…

Machine Learning · Computer Science 2020-03-17 Zhixuan Lin , Yi-Fu Wu , Skand Vishwanath Peri , Weihao Sun , Gautam Singh , Fei Deng , Jindong Jiang , Sungjin Ahn

Deep learning models have enjoyed great success for image related computer vision tasks like image classification and object detection. For video related tasks like human action recognition, however, the advancements are not as significant…

Computer Vision and Pattern Recognition · Computer Science 2018-09-12 Xiaolin Song , Cuiling Lan , Wenjun Zeng , Junliang Xing , Jingyu Yang , Xiaoyan Sun

Multi-frame story illustration requires long-horizon coherence beyond single-image text-to-image generation, including narrative decomposition and persistent character identity, layout, and affect across frames. We propose…

Artificial Intelligence · Computer Science 2026-05-22 Sijing Yin , Jiamou Liu , Xiao Tang , Yaser Shakib , Qian Liu

Accurate 3D human pose estimation from monocular videos requires effective modelling of complex spatial and temporal dependencies. However, existing methods often face challenges in efficiency and adaptability when modelling spatial and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Ruochen Li , Shuang Chen , Wenke E , Farshad Arvin , Amir Atapour-Abarghouei

In this work, we develop a multi-modal rendering framework comprising of hapto-visual and auditory data. The prime focus is to haptically render point cloud data representing virtual 3-D models of cultural significance and also to handle…

Event cameras asynchronously capture brightness changes with low latency, high temporal resolution, and high dynamic range. However, annotation of event data is a costly and laborious process, which limits the use of deep learning methods…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Simon Klenk , David Bonello , Lukas Koestler , Nikita Araslanov , Daniel Cremers

Typical techniques for video captioning follow the encoder-decoder framework, which can only focus on one source video being processed. A potential disadvantage of such design is that it cannot capture the multiple visual context…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Wenjie Pei , Jiyuan Zhang , Xiangrong Wang , Lei Ke , Xiaoyong Shen , Yu-Wing Tai

Multimodal diegetic narrative tools, as applied in multimedia arts practices, possess the ability to cross the spaces that exist between the physical world and the imaginary. Within this paper we present the findings of a multidiscipline…

Human-Computer Interaction · Computer Science 2020-11-13 Gareth W. Young , Siobhán Mannion , Sara Wentworth

Temporal Information and Event Markup Language (TIE-ML) is a markup strategy and annotation schema to improve the productivity and accuracy of temporal and event related annotation of corpora to facilitate machine learning based model…

Computation and Language · Computer Science 2021-09-29 Damir Cavar , Billy Dickson , Ali Aljubailan , Soyoung Kim

Short-form digital storytelling has become a popular medium for millions of people to express themselves. Traditionally, this medium uses primarily 2D media such as text (e.g., memes), images (e.g., Instagram), gifs (e.g., Giphy), and…

Human-Computer Interaction · Computer Science 2021-08-31 Mengyu Chen , Andrés Monroy-Hernández , Misha Sra

Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potential to outperform the classical solutions developed for this…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Zachary Seymour , Kowshik Thopalli , Niluthpol Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Despite the growing adoption of mixed reality and interactive AI agents, it remains challenging for these systems to generate high quality 2D/3D scenes in unseen environments. The common practice requires deploying an AI agent to collect…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Qiuyuan Huang , Jae Sung Park , Abhinav Gupta , Paul Bennett , Ran Gong , Subhojit Som , Baolin Peng , Owais Khan Mohammed , Chris Pal , Yejin Choi , Jianfeng Gao

Interaction with the physical environment and different users is essential to foster a collaborative experience. For this, we propose an interaction based on a central point represented by an Augmented Reality marker in which several users…

Human-Computer Interaction · Computer Science 2023-01-09 Bianca Marques , Rui Nóbrega , Carmen Morgado

This paper presents a novel approach to processing multimodal data for dynamic emotion recognition, named as the Multimodal Masked Autoencoder for Dynamic Emotion Recognition (MultiMAE-DER). The MultiMAE-DER leverages the closely correlated…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Peihao Xiang , Chaohao Lin , Kaida Wu , Ou Bai

This work introduces a novel Augmented Reality (AR) approach to visualize material data alongside real objects in order to facilitate detailed material analyses based on spatial non-destructive testing (NDT) data as generated in X-ray…

Human-Computer Interaction · Computer Science 2024-04-22 Alexander Gall , Anja Heim , Patrick Weinberger , Bernhard Fröhler , Johann Kastner , Christoph Heinzl

Neurosurgery requires exceptional precision and comprehensive preoperative planning to ensure optimal patient outcomes. Despite technological advancements, there remains a need for intuitive, accessible tools to enhance surgical preparation…

Robotics · Computer Science 2024-11-06 Hon Lung Ho , Yupeng Wang , An Wang , Long Bai , Hongliang Ren

We propose an online spatiotemporal articulation model estimation framework that estimates both articulated structure as well as a temporal prediction model solely using passive observations. The resulting model can predict future mo- tions…

Robotics · Computer Science 2016-04-13 Suren Kumar , Vikas Dhiman , Madan Ravi Ganesh , Jason J. Corso