中文
相关论文

相关论文: Continuous interaction with a smart speaker via lo…

200 篇论文

Active headrests can reduce low-frequency noise around ears based on active noise control (ANC) system. Both the control system using fixed control filters and the remote microphone-based adaptive control system provide good noise reduction…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Yuteng Liu , Haowen Li , Haishan Zou , Jing Lu , Zhibin Lin

The acoustic response of an object can reveal a lot about its global state, for example its material properties or the extrinsic contacts it is making with the world. In this work, we build an active acoustic sensing gripper equipped with…

Recent pose-to-video models can translate 2D pose sequences into photorealistic, identity-preserving dance videos, so the key challenge is to generate temporally coherent, rhythm-aligned 2D poses from music, especially under complex,…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Yan Zhang , Han Zou , Lincong Feng , Cong Xie , Ruiqi Yu , Zhenpeng Zhan

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Wenqing Wang , Yun Fu

This paper focuses on enhancing human-agent communication by integrating spatial context into virtual agents' non-verbal behaviors, specifically gestures. Recent advances in co-speech gesture generation have primarily utilized data-driven…

人机交互 · 计算机科学 2024-08-09 Anna Deichler , Simon Alexanderson , Jonas Beskow

Manual assembly workers face increasing complexity in their work. Human-centered assistance systems could help, but object recognition as an enabling technology hinders sophisticated human-centered design of these systems. At the same time,…

计算机视觉与模式识别 · 计算机科学 2024-02-15 Christian Jauch , Timo Leitritz , Marco F. Huber

Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the temporal correspondence of audiovisual cues, such as the…

多媒体 · 计算机科学 2025-02-11 Jason Clarke , Yoshihiko Gotoh , Stefan Goetze

In this paper, we present an efficient method to incrementally learn to classify static hand gestures. This method allows users to teach a robot to recognize new symbols in an incremental manner. Contrary to other works which use special…

机器人学 · 计算机科学 2023-04-14 Xavier Cucurull , Anaís Garrell

Emotion recognition,as a step toward mind reading,seeks to infer internal states from external cues.Most existing methods rely on explicit signals-such as facial expressions,speech,or gestures-that reflect only bodily responses and overlook…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Mengke Song , Yuge Xie , Qi Cui , Luming Li , Xinyu Liu , Guotao Wang , Chenglizhao Chen , Shanchen Pang

Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Kefan Chen , Chaerin Min , Linguang Zhang , Shreyas Hampali , Cem Keskin , Srinath Sridhar

Speaker tracking methods often rely on spatial observations to assign coherent track identities over time. This raises limits in scenarios with intermittent and moving speakers, i.e., speakers that may change position when they are…

音频与语音处理 · 电气工程与系统科学 2025-06-26 Taous Iatariene , Can Cui , Alexandre Guérin , Romain Serizel

Event camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges:…

计算机视觉与模式识别 · 计算机科学 2021-08-17 Shihao Zou , Chuan Guo , Xinxin Zuo , Sen Wang , Pengyu Wang , Xiaoqin Hu , Shoushun Chen , Minglun Gong , Li Cheng

Static and dynamic hand movements are basic way for human-machine interactions. To recognize and classify these movements, first these movements are captured by the cameras mounted on the augmented reality (AR) or virtual reality (VR)…

人机交互 · 计算机科学 2020-04-03 Nizamuddin Maitlo , Yanbo Wang , Chao Ping Chen , Lantian Mi , Wenbo Zhang

Force estimation in human-object interactions is crucial for various fields like ergonomics, physical therapy, and sports science. Traditional methods depend on specialized equipment such as force plates and sensors, which makes accurate…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Nandakishor M , Vrinda Govind , Anuradha Puthalath , Anzy L , Swathi P S , Aswathi R , Devaprabha A R , Varsha Raj , Midhuna Krishnan K , Akhila Anilkumar T , Yamuna P

Audio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress, they often struggle…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Weizhi Zhong , Junfan Lin , Peixin Chen , Liang Lin , Guanbin Li

Dynamic and dexterous manipulation of objects presents a complex challenge, requiring the synchronization of hand motions with the trajectories of objects to achieve seamless and physically plausible interactions. In this work, we introduce…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Jiajun Zhang , Yuxiang Zhang , Liang An , Mengcheng Li , Hongwen Zhang , Zonghai Hu , Yebin Liu

To the best of our knowledge, we first present a live system that generates personalized photorealistic talking-head animation only driven by audio signals at over 30 fps. Our system contains three stages. The first stage is a deep neural…

图形学 · 计算机科学 2021-09-27 Yuanxun Lu , Jinxiang Chai , Xun Cao

This paper presents an experimental study on deep speaker embedding with an attention mechanism that has been found to be a powerful representation learning technique in speaker recognition. In this framework, an attention model works as a…

声音 · 计算机科学 2018-09-26 Qiongqiong Wang , Koji Okabe , Kong Aik Lee , Hitoshi Yamamoto , Takafumi Koshinaka

Recovering world space 4D motion of two interacting hands from egocentric video is a fundamental capability for supervising robot policy learning, where wrist trajectories track the end-effector and finger articulations specify the grasp…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Huajian Zeng , Chaohua Yao , Yuantai Zhang , Jiaqi Yang , Rolandos Alexandros Potamias , Xingxing Zuo

A key component of dyadic spoken interactions is the contextually relevant non-verbal gestures, such as head movements that reflect a listener's response to the interlocutor's speech. Although significant progress has been made in the…

机器人学 · 计算机科学 2024-10-01 Bishal Ghosh , Emma Li , Tanaya Guha