中文
相关论文

相关论文: Continuous interaction with a smart speaker via lo…

200 篇论文

While current talking head models are capable of generating photorealistic talking head videos, they provide limited pose controllability. Most methods require specific video sequences that should exactly contain the head pose desired,…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Kwangho Lee , Patrick Kwon , Myung Ki Lee , Namhyuk Ahn , Junsoo Lee

Accurate and responsive myoelectric prosthesis control typically relies on complex, dense multi-sensor arrays, which limits consumer accessibility. This paper presents a novel, data-efficient deep learning framework designed to achieve…

机器学习 · 计算机科学 2026-02-04 Blagoj Hristov , Hristijan Gjoreski , Vesna Ojleska Latkoska , Gorjan Nadzinski

We describe a novel metric-based learning approach that introduces a multimodal framework and uses deep audio and geophone encoders in siamese configuration to design an adaptable and lightweight supervised model. This framework eliminates…

声音 · 计算机科学 2021-11-16 Muhammad Shakeel , Katsutoshi Itoyama , Kenji Nishida , Kazuhiro Nakadai

While deep-learning-based speaker localization has shown advantages in challenging acoustic environments, it often yields only direction-of-arrival (DOA) cues rather than precise two-dimensional (2D) coordinates. To address this, we propose…

音频与语音处理 · 电气工程与系统科学 2024-04-02 Shupei Liu , Linfeng Feng , Yijun Gong , Chengdong Liang , Chen Zhang , Xiao-Lei Zhang , Xuelong Li

Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Jisu Nam , Yicong Hong , Chun-Hao Paul Huang , Feng Liu , JoungBin Lee , Jiyoung Kim , Siyoon Jin , Yunsung Lee , Jaeyoon Jung , Suhwan Choi , Seungryong Kim , Yang Zhou

Facial expression recognition is a crucial component in enhancing human-computer interaction and developing emotion-aware systems. Real-time detection and interpretation of facial expressions have become increasingly important for various…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Talha Enes Koksal , Abdurrahman Gumus

Mobile imitation learning on portable demonstration interfaces faces two coupled bottlenecks: locomotion-contaminated action labels and inference-induced execution latency on a continuously moving base. Recent wrist-mounted interfaces lower…

机器人学 · 计算机科学 2026-05-21 Haoran Huang , Haonan Dong , Huixu Dong

Tactile sensing is a crucial perception mode for robots and human amputees in need of controlling a prosthetic device. Today robotic and prosthetic systems are still missing the important feature of accurate tactile sensing. This lack is…

机器人学 · 计算机科学 2022-03-30 Xiaying Wang , Fabian Geiger , Vlad Niculescu , Michele Magno , Luca Benini

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

Compositing human figures into scene images has broad applications in areas such as entertainment and advertising. However, existing methods often cannot handle occlusion of the inserted person by foreground objects and unnaturally place…

图形学 · 计算机科学 2025-05-08 Shun Masuda , Yuki Endo , Yoshihiro Kanamori

We are concerned with a novel sensor-based gesture input/instruction technology which enables human beings to interact with computers conveniently. The human being wears an emitter on the finger or holds a digital pen that generates a time…

经典物理 · 物理学 2017-05-23 Yukun Guo , Jingzhi Li , Hongyu Liu , Xianchao Wang

Grasping is natural for humans. However, it involves complex hand configurations and soft tissue deformation that can result in complicated regions of contact between the hand and the object. Understanding and modeling this contact can…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Samarth Brahmbhatt , Chengcheng Tang , Christopher D. Twigg , Charles C. Kemp , James Hays

Human pose estimation in two-dimensional images videos has been a hot topic in the computer vision problem recently due to its vast benefits and potential applications for improving human life, such as behaviors recognition, motion capture…

计算机视觉与模式识别 · 计算机科学 2022-02-08 Thong Duy Nguyen , Milan Kresovic

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces…

音频与语音处理 · 电气工程与系统科学 2020-05-05 Arsha Nagrani , Joon Son Chung , Samuel Albanie , Andrew Zisserman

We present Implicit Two Hands (Im2Hands), the first neural implicit representation of two interacting hands. Unlike existing methods on two-hand reconstruction that rely on a parametric hand model and/or low-resolution meshes, Im2Hands can…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Jihyun Lee , Minhyuk Sung , Honggyu Choi , Tae-Kyun Kim

Hand-over-face gestures can provide important implicit interactions during conversations, such as frustration or excitement. However, in situations where interlocutors are not visible, such as phone calls or textual communication, the…

人机交互 · 计算机科学 2024-03-28 Mengxi Liu , Hymalai Bello , Bo Zhou , Paul Lukowicz , Jakob Karolus

Pose Estimation techniques rely on visual cues available through observations represented in the form of pixels. But the performance is bounded by the frame rate of the video and struggles from motion blur, occlusions, and temporal…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Snehesh Shrestha , Cornelia Fermüller , Tianyu Huang , Pyone Thant Win , Adam Zukerman , Chethan M. Parameshwara , Yiannis Aloimonos

Hand pose estimation from 3D depth images, has been explored widely using various kinds of techniques in the field of computer vision. Though, deep learning based method improve the performance greatly recently, however, this problem still…

计算机视觉与模式识别 · 计算机科学 2020-01-24 Zhaohui Zhang , Shipeng Xie , Mingxiu Chen , Haichao Zhu

Around-device interaction promises to extend the input space of mobile and wearable devices beyond the common but restricted touchscreen. So far, most around-device interaction approaches rely on instrumenting the device or the environment…

人机交互 · 计算机科学 2017-01-17 Jens Grubert , Eyal Ofek , Michel Pahud , Matthias Kranz , Dieter Schmalstieg

Reliable control of myoelectric prostheses is often hindered by high inter-subject variability and the clinical impracticality of high-density sensor arrays. This study proposes a deep learning framework for accurate gesture recognition…