English
Related papers

Related papers: Continuous interaction with a smart speaker via lo…

200 papers

Active headrests can reduce low-frequency noise around ears based on active noise control (ANC) system. Both the control system using fixed control filters and the remote microphone-based adaptive control system provide good noise reduction…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Yuteng Liu , Haowen Li , Haishan Zou , Jing Lu , Zhibin Lin

The acoustic response of an object can reveal a lot about its global state, for example its material properties or the extrinsic contacts it is making with the world. In this work, we build an active acoustic sensing gripper equipped with…

Recent pose-to-video models can translate 2D pose sequences into photorealistic, identity-preserving dance videos, so the key challenge is to generate temporally coherent, rhythm-aligned 2D poses from music, especially under complex,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Yan Zhang , Han Zou , Lincong Feng , Cong Xie , Ruiqi Yu , Zhenpeng Zhan

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Wenqing Wang , Yun Fu

This paper focuses on enhancing human-agent communication by integrating spatial context into virtual agents' non-verbal behaviors, specifically gestures. Recent advances in co-speech gesture generation have primarily utilized data-driven…

Human-Computer Interaction · Computer Science 2024-08-09 Anna Deichler , Simon Alexanderson , Jonas Beskow

Manual assembly workers face increasing complexity in their work. Human-centered assistance systems could help, but object recognition as an enabling technology hinders sophisticated human-centered design of these systems. At the same time,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Christian Jauch , Timo Leitritz , Marco F. Huber

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the…

Multimedia · Computer Science 2025-02-11 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

In this paper, we present an efficient method to incrementally learn to classify static hand gestures. This method allows users to teach a robot to recognize new symbols in an incremental manner. Contrary to other works which use special…

Robotics · Computer Science 2023-04-14 Xavier Cucurull , Anaís Garrell

Emotion recognition,as a step toward mind reading,seeks to infer internal states from external cues.Most existing methods rely on explicit signals-such as facial expressions,speech,or gestures-that reflect only bodily responses and overlook…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Mengke Song , Yuge Xie , Qi Cui , Luming Li , Xinyu Liu , Guotao Wang , Chenglizhao Chen , Shanchen Pang

Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Kefan Chen , Chaerin Min , Linguang Zhang , Shreyas Hampali , Cem Keskin , Srinath Sridhar

Speaker tracking methods often rely on spatial observations to assign coherent track identities over time. This raises limits in scenarios with intermittent and moving speakers, i.e., speakers that may change position when they are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-26 Taous Iatariene , Can Cui , Alexandre Guérin , Romain Serizel

Event camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges:…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Shihao Zou , Chuan Guo , Xinxin Zuo , Sen Wang , Pengyu Wang , Xiaoqin Hu , Shoushun Chen , Minglun Gong , Li Cheng

Static and dynamic hand movements are basic way for human-machine interactions. To recognize and classify these movements, first these movements are captured by the cameras mounted on the augmented reality (AR) or virtual reality (VR)…

Human-Computer Interaction · Computer Science 2020-04-03 Nizamuddin Maitlo , Yanbo Wang , Chao Ping Chen , Lantian Mi , Wenbo Zhang

Force estimation in human-object interactions is crucial for various fields like ergonomics, physical therapy, and sports science. Traditional methods depend on specialized equipment such as force plates and sensors, which makes accurate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Nandakishor M , Vrinda Govind , Anuradha Puthalath , Anzy L , Swathi P S , Aswathi R , Devaprabha A R , Varsha Raj , Midhuna Krishnan K , Akhila Anilkumar T , Yamuna P

Audio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress, they often struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Weizhi Zhong , Junfan Lin , Peixin Chen , Liang Lin , Guanbin Li

Dynamic and dexterous manipulation of objects presents a complex challenge, requiring the synchronization of hand motions with the trajectories of objects to achieve seamless and physically plausible interactions. In this work, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Jiajun Zhang , Yuxiang Zhang , Liang An , Mengcheng Li , Hongwen Zhang , Zonghai Hu , Yebin Liu

To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural…

Graphics · Computer Science 2021-09-27 Yuanxun Lu , Jinxiang Chai , Xun Cao

This paper presents an experimental study on deep speaker embedding with an attention mechanism that has been found to be a powerful representation learning technique in speaker recognition. In this framework, an attention model works as a…

Sound · Computer Science 2018-09-26 Qiongqiong Wang , Koji Okabe , Kong Aik Lee , Hitoshi Yamamoto , Takafumi Koshinaka

Recovering world space 4D motion of two interacting hands from egocentric video is a fundamental capability for supervising robot policy learning, where wrist trajectories track the end-effector and finger articulations specify the grasp…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Huajian Zeng , Chaohua Yao , Yuantai Zhang , Jiaqi Yang , Rolandos Alexandros Potamias , Xingxing Zuo

A key component of dyadic spoken interactions is the contextually relevant non-verbal gestures, such as head movements that reflect a listener's response to the interlocutor's speech. Although significant progress has been made in the…

Robotics · Computer Science 2024-10-01 Bishal Ghosh , Emma Li , Tanaya Guha
‹ Prev 1 4 5 6 7 8 10 Next ›