中文
相关论文

相关论文: Attention based Occlusion Removal for Hybrid Telep…

200 篇论文

Visual place recognition is a challenging task for applications such as autonomous driving navigation and mobile robot localization. Distracting elements presenting in complex scenes often lead to deviations in the perception of visual…

计算机视觉与模式识别 · 计算机科学 2022-04-14 Ruotong Wang , Yanqing Shen , Weiliang Zuo , Sanping Zhou , Nanning Zheng

Responsive and accurate facial expression recognition is crucial to human-robot interaction for daily service robots. Nowadays, event cameras are becoming more widely adopted as they surpass RGB cameras in capturing facial expression…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Zhe Wang , Qijin Song , Yucen Peng , Weibang Bai

This paper presents Nana-HDR, a new non-attentive non-autoregressive model with hybrid Transformer-based Dense-fuse encoder and RNN-based decoder for TTS. It mainly consists of three parts: Firstly, a novel Dense-fuse encoder with dense…

计算与语言 · 计算机科学 2021-09-29 Shilun Lin , Wenchao Su , Li Meng , Fenglong Xie , Xinhui Li , Li Lu

Purpose: Image guidance is crucial for the success of many interventions. Images are displayed on designated monitors that cannot be positioned optimally due to sterility and spatial constraints. This indirect visualization causes potential…

其他计算机科学 · 计算机科学 2017-10-03 Long Qian , Mathias Unberath , Kevin Yu , Bernhard Fuerst , Alex Johnson , Nassir Navab , Greg Osgood

Face inpainting requires the model to have a precise global understanding of the facial position structure. Benefiting from the powerful capabilities of deep learning backbones, recent works in face inpainting have achieved decent…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Bo Zhao , Huan Yang , Jianlong Fu

Head-mounted displays (HMDs) are popular immersive tools in general, not limited to entertainment but also for education, military, and serious games for health. While these displays have strong popularity, they still have user experience…

人机交互 · 计算机科学 2022-07-15 Thiago Porcino , Derek Reilly , Esteban Clua , Daniela Trevisan

Face hallucination is a domain-specific super-resolution problem that aims to generate a high-resolution (HR) face image from a low-resolution~(LR) input. In contrast to the existing patch-wise super-resolution models that divide a face…

计算机视觉与模式识别 · 计算机科学 2019-05-07 Yukai Shi , Guanbin Li , Qingxing Cao , Keze Wang , Liang Lin

Because of affected by weather conditions, camera pose and range, etc. Objects are usually small, blur, occluded and diverse pose in the images gathered from outdoor surveillance cameras or access control system. It is challenging and…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Zexun Zhou , Zhongshi He , Ziyu Chen , Yuanyuan Jia , Haiyan Wang , Jinglong Du , Dingding Chen

Recent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such…

计算机视觉与模式识别 · 计算机科学 2023-08-24 Xiang Zhang , Huiyuan Yang , Taoyue Wang , Xiaotian Li , Lijun Yin

User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture,…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Zhengxian Yang , Shi Pan , Shengqi Wang , Haoxiang Wang , Li Lin , Guanjun Li , Zhengqi Wen , Borong Lin , Jianhua Tao , Tao Yu

Gaze interaction presents a promising avenue in Virtual Reality (VR) due to its intuitive and efficient user experience. Yet, the depth control inherent in our visual system remains underutilized in current methods. In this study, we…

人机交互 · 计算机科学 2024-05-08 Chenyang Zhang , Tiansu Chen , Eric Shaffer , Elahe Soltanaghai

Transducer and Attention based Encoder-Decoder (AED) are two widely used frameworks for speech-to-text tasks. They are designed for different purposes and each has its own benefits and drawbacks for speech-to-text tasks. In order to…

计算与语言 · 计算机科学 2023-05-08 Yun Tang , Anna Y. Sun , Hirofumi Inaguma , Xinyue Chen , Ning Dong , Xutai Ma , Paden D. Tomasello , Juan Pino

Many real-world applications today like video surveillance and urban governance need to address the recognition of masked faces, where content replacement by diverse masks often brings in incomplete appearance and ambiguous representation,…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Chenyu Li , Shiming Ge , Daichi Zhang , Jia Li

Frame prediction based on AutoEncoder plays a significant role in unsupervised video anomaly detection. Ideally, the models trained on the normal data could generate larger prediction errors of anomalies. However, the correlation between…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Xiangyu Huang , Caidan Zhao , Jinghui Yu , Chenxing Gao , Zhiqiang Wu

Transformer architecture has emerged to be successful in a number of natural language processing tasks. However, its applications to medical vision remain largely unexplored. In this study, we present UTNet, a simple yet powerful hybrid…

计算机视觉与模式识别 · 计算机科学 2021-09-29 Yunhe Gao , Mu Zhou , Dimitris Metaxas

Virtual reality (VR) head-mounted displays (HMD) have recently been used to provide an immersive, first-person vision/view in real-time for manipulating remotely-controlled unmanned ground vehicles (UGV). The teleoperation of UGV can be…

人机交互 · 计算机科学 2021-07-13 Yiming Luo , Jialin Wang , Hai-Ning Liang , Shan Luo , Eng Gee Lim

Relightable portrait animation aims to animate a static reference portrait to match the head movements and expressions of a driving video while adapting to user-specified or reference lighting conditions. Existing portrait animation methods…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Mingtao Guo , Guanyu Xing , Yanli Liu

Intense bandwidth depletion within consumer and constrained networks has the potential to undermine the stability of real-time video conferencing: encoder rate management becomes saturated, packet loss escalates, frame rates deteriorate,…

图像与视频处理 · 电气工程与系统科学 2026-02-16 Vineet Kumar Rakesh , Soumya Mazumdar , Tapas Samanta , Hemendra Kumar Pandey , Amitabha Das , Sarbajit Pal

We propose a neural talking-head video synthesis model and demonstrate its application to video conferencing. Our model learns to synthesize a talking-head video using a source image containing the target person's appearance and a driving…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Ting-Chun Wang , Arun Mallya , Ming-Yu Liu

Despite recent advances in facial recognition, there remains a fundamental issue concerning degradations in performance due to substantial perspective (pose) differences between enrollment and query (probe) imagery. Therefore, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-05-15 J. Brennan Peace , Shuowen Hu , Benjamin S. Riggan
‹ 上一页 1 8 9 10 下一页 ›