中文
相关论文

相关论文: Designing, Playing, and Performing with a Vision-b…

200 篇论文

The tongue is a crucial organ for performing basic biological functions, such as chewing, swallowing and phonation. Understanding how it behaves, its motor control and involvement in the execution of these different tasks is therefore an…

医学物理 · 物理学 2024-07-19 Maxime Calka , Pascal Perrier , Michel Rochette , Yohan Payan

We aim to edit the lip movements in talking video according to the given speech while preserving the personal identity and visual details. The task can be decomposed into two sub-problems: (1) speech-driven lip motion generation and (2)…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Runyi Yu , Tianyu He , Ailing Zhang , Yuchi Wang , Junliang Guo , Xu Tan , Chang Liu , Jie Chen , Jiang Bian

Gestures performed accompanying the voice are essential for voice interaction to convey complementary semantics for interaction purposes such as wake-up state and input modality. In this paper, we investigated voice-accompanying…

人机交互 · 计算机科学 2023-03-21 Zisu Li , Cheng Liang , Yuntao Wang , Yue Qin , Chun Yu , Yukang Yan , Mingming Fan , Yuanchun Shi

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial…

计算机视觉与模式识别 · 计算机科学 2023-06-14 Ricong Huang , Peiwen Lai , Yipeng Qin , Guanbin Li

Musical expression requires control of both what notes are played, and how they are performed. Conventional audio synthesizers provide detailed expressive controls, but at the cost of realism. Black-box neural audio synthesis and…

Although significant progress has been made in the field of speech-driven 3D facial animation recently, the speech-driven animation of an indispensable facial component, eye gaze, has been overlooked by recent research. This is primarily…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Yixiang Zhuang , Chunshan Ma , Yao Cheng , Xuan Cheng , Jing Liao , Juncong Lin

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This often requires…

音频与语音处理 · 电气工程与系统科学 2018-07-20 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

This paper reviews the existing literature on input device evaluation and design in human-computer interaction (HCI) and discusses possible applications of this knowledge to the design and evaluation of new interfaces for musical…

人机交互 · 计算机科学 2020-10-06 Nicola Orio , Norbert Schnell , Marcelo M. Wanderley

This paper presents a new continuous interaction strategy with visual feedback of hand pose and mid-air gesture recognition and control for a smart music speaker, which utilizes only 2 video frames to recognize gestures. Frame-based hand…

人机交互 · 计算机科学 2023-03-01 Songpei Xu , Chaitanya Kaul , Xuri Ge , Roderick Murray-Smith

In this study, we explore the representation mapping from the domain of visual arts to the domain of music, with which we can use visual arts as an effective handle to control music generation. Unlike most studies in multimodal…

声音 · 计算机科学 2022-11-11 Runbang Zhang , Yixiao Zhang , Kai Shao , Ying Shan , Gus Xia

Interaction methods based on computer-vision hold the potential to become the next powerful technology to support breakthroughs in the field of human-computer interaction. Non-invasive vision-based techniques permit unconventional…

人机交互 · 计算机科学 2017-07-27 Chamin Morikawa , Michael J. Lyons

For realistic talking head generation, creating natural head motion while maintaining accurate lip synchronization is essential. To fulfill this challenging task, we propose DisCoHead, a novel method to disentangle and control head pose and…

计算机视觉与模式识别 · 计算机科学 2023-03-15 Geumbyeol Hwang , Sunwon Hong , Seunghyun Lee , Sungwoo Park , Gyeongsu Chae

Speech has been a widely used modality in the field of affective computing. Recently however, there has been a growing interest in the use of multi-modal affective computing systems. These multi-modal systems incorporate both verbal and…

人机交互 · 计算机科学 2018-05-18 Jonny O'Dwyer , Niall Murray , Ronan Flynn

Recently audio-driven talking face video generation has attracted considerable attention. However, very few researches address the issue of emotional editing of these talking face videos with continuously controllable expressions, which is…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Zhiyao Sun , Yu-Hui Wen , Tian Lv , Yanan Sun , Ziyang Zhang , Yaoyuan Wang , Yong-Jin Liu

In recent years, audio-driven 3D facial animation has gained significant attention, particularly in applications such as virtual reality, gaming, and video conferencing. However, accurately modeling the intricate and subtle dynamics of…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Guinan Su , Yanwu Yang , Zhifeng Li

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

One of the many tasks facing the typically-developing child language learner is learning to discriminate between the distinctive sounds that make up words in their native language. Here we investigate whether multimodal…

计算与语言 · 计算机科学 2024-07-24 Sophia Zhi , Roger P. Levy , Stephan C. Meylan

Speech-driven 3D facial animation is challenging due to the diversity in speaking styles and the limited availability of 3D audio-visual data. Speech predominantly dictates the coarse motion trends of the lip region, while specific styles…

多媒体 · 计算机科学 2025-03-14 An Yang , Chenyu Liu , Pengcheng Xia , Jun Du