English
Related papers

Related papers: Semantics-Aware Human Motion Generation from Audio…

200 papers

When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential cross-modal alignment by modelling the image description generation process computationally. We take as our…

Computation and Language · Computer Science 2020-11-10 Ece Takmaz , Sandro Pezzelle , Lisa Beinborn , Raquel Fernández

Content-based music information retrieval has seen rapid progress with the adoption of deep learning. Current approaches to high-level music description typically make use of classification models, such as in auto-tagging or genre and mood…

Sound · Computer Science 2021-12-09 Ilaria Manco , Emmanouil Benetos , Elio Quinton , Gyorgy Fazekas

Animation elevates digital documents into immersive experiences, yet creating custom motion paths remains cumbersome, requiring designers to manually select presets, plot B\'ezier points, and configure timing properties. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Mannat Khurana , Sanyam Jain , Rishav Agarwal

Synthesizing natural interactions between virtual humans and their 3D environments is critical for numerous applications, such as computer games and AR/VR experiences. Our goal is to synthesize humans interacting with a given 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Kaifeng Zhao , Shaofei Wang , Yan Zhang , Thabo Beeler , Siyu Tang

This paper presents an end-to-end pipeline for generating character-specific, emotion-aware speech from comics. The proposed system takes full comic volumes as input and produces speech aligned with each character's dialogue and emotional…

Sound · Computer Science 2025-09-22 Zhiwen Qian , Jinhua Liang , Huan Zhang

Semantic communication aims to transmit information most relevant to a task rather than raw data, offering significant gains in communication efficiency for applications such as telepresence, augmented reality, and remote sensing. Recent…

Machine Learning · Computer Science 2025-12-18 Matin Mortaheb , Erciyes Karakaya , Sennur Ulukus

Investigating the relationship between internal tissue point motion of the tongue and oropharyngeal muscle deformation measured from tagged MRI and intelligible speech can aid in advancing speech motor control theories and developing novel…

Image and Video Processing · Electrical Eng. & Systems 2023-02-15 Xiaofeng Liu , Fangxu Xing , Jerry L. Prince , Maureen Stone , Georges El Fakhri , Jonghye Woo

In this work, we address the task of unconditional head motion generation to animate still human faces in a low-dimensional semantic space from a single reference pose. Different from traditional audio-conditioned talking head generation…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Louis Airale , Xavier Alameda-Pineda , Stéphane Lathuilière , Dominique Vaufreydaz

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the speaker's words with their timestamps, our approach…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Evonne Ng , Sanjay Subramanian , Dan Klein , Angjoo Kanazawa , Trevor Darrell , Shiry Ginosar

We present an end-to-end method for transforming audio from one style to another. For the case of speech, by conditioning on speaker identities, we can train a single model to transform words spoken by multiple people into multiple target…

Sound · Computer Science 2018-06-08 Albert Haque , Michelle Guo , Prateek Verma

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Xingyu Chen

We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal…

Multimedia · Computer Science 2025-07-28 Hyunwoo Oh , SeungJu Cha , Kwanyoung Lee , Si-Woo Kim , Dong-Jin Kim

We present a framework that can impose the audio effects and production style from one recording to another by example with the goal of simplifying the audio production process. We train a deep neural network to analyze an input recording…

Sound · Computer Science 2022-07-19 Christian J. Steinmetz , Nicholas J. Bryan , Joshua D. Reiss

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

The speech code is a vehicle of language: it defines a set of forms used by a community to carry information. Such a code is necessary to support the linguistic interactions that allow humans to communicate. How then may a speech code be…

Machine Learning · Computer Science 2007-05-23 Pierre-Yves Oudeyer

Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This…

Computation and Language · Computer Science 2025-01-07 Ariel Shaulov , Tal Shaharabany , Eitan Shaar , Gal Chechik , Lior Wolf

We study the problem of making 3D scene reconstructions interactive by asking the following question: can we predict the sounds of human hands physically interacting with a scene? First, we record a video of a human manipulating objects…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yiming Dou , Wonseok Oh , Yuqing Luo , Antonio Loquercio , Andrew Owens

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we…

Sound · Computer Science 2025-09-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal…

Sound · Computer Science 2024-12-25 Yaoyun Zhang , Xuenan Xu , Mengyue Wu