English
Related papers

Related papers: Estimating Visual Information From Audio Through M…

200 papers

Human motion generation involves creating natural sequences of human body poses, widely used in gaming, virtual reality, and human-computer interaction. It aims to produce lifelike virtual characters with realistic movements, enhancing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Jiayi Zhao , Dongdong Weng , Qiuxin Du , Zeyu Tian

Finding sound effects or environmental sounds that match a creator's intended impression remains a largely manual process in multimedia production. This is especially relevant for comics and other visual media, where visually stylized…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Keisuke Imoto , Yamato Kojima , Takao Tsuchiya

Automatically learning features, especially robust features, has attracted much attention in the machine learning community. In this paper, we propose a new method to learn non-linear robust features by taking advantage of the data manifold…

Machine Learning · Computer Science 2017-05-29 Yanan Li , Donghui Wang

In this work, we present a method for learning interpretable music signal representations directly from waveform signals. Our method can be trained using unsupervised objectives and relies on the denoising auto-encoder model that uses a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-02 Stylianos I. Mimilakis , Konstantinos Drossos , Gerald Schuller

Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the design of robust models for information retrieval.…

Sound · Computer Science 2022-10-31 Kleanthis Avramidis , Shanti Stewart , Shrikanth Narayanan

Neural networks used for image classification tasks in critical applications must be tested with sufficient realistic data to assure their correctness. To effectively test an image classification neural network, one must obtain realistic…

Machine Learning · Computer Science 2020-02-18 Taejoon Byun , Abhishek Vijayakumar , Sanjai Rayadurgam , Darren Cofer

This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multiple European cities, using the same equipment, in multiple…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-12 Shanshan Wang , Annamaria Mesaros , Toni Heittola , Tuomas Virtanen

Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Jinyang Yuan , Bin Li , Xiangyang Xue

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang

How to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual segmentation (AVS) task has been proposed, aiming to segment the sounding objects in…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Tianxiang Chen , Zhentao Tan , Tao Gong , Qi Chu , Yue Wu , Bin Liu , Le Lu , Jieping Ye , Nenghai Yu

Manipulated videos often contain subtle inconsistencies between their visual and audio signals. We propose a video forensics method, based on anomaly detection, that can identify these inconsistencies, and that can be trained solely using…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Chao Feng , Ziyang Chen , Andrew Owens

Previous methods for audio-image matching generally fall into one of two categories: pipeline models or End-to-End models. Pipeline models first transcribe speech and then encode the resulting text; End-to-End models encode speech directly.…

Sound · Computer Science 2024-08-21 Zhenyu Lu , Lakshay Sethi

Stereophonic audio, especially binaural audio, plays an essential role in immersive viewing environments. Recent research has explored generating visually guided stereophonic audios supervised by multi-channel audio collections. However,…

Sound · Computer Science 2021-04-14 Xudong Xu , Hang Zhou , Ziwei Liu , Bo Dai , Xiaogang Wang , Dahua Lin

Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs become stronger, the key challenge shifts from enabling tool use…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-28 Yucheng Wang , Jing Peng , Hanqi Li , Chenghao Wang , Wenming Tu , Yu Xi , Zhaokai Sun , Kai Yu , Shuai Wang

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy…

Graphics · Computer Science 2021-12-02 Seung Hyun Lee , Wonseok Roh , Wonmin Byeon , Sang Ho Yoon , Chan Young Kim , Jinkyu Kim , Sangpil Kim

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

Sound · Computer Science 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

This dissertation proposes the study of multimodal learning in the context of musical signals. Throughout, we focus on the interaction between audio signals and text information. Among the many text sources related to music that can be used…

Sound · Computer Science 2021-11-01 Gabriel Meseguer-Brocal

In this paper we propose a novel learning framework called Supervised and Weakly Supervised Learning where the goal is to learn simultaneously from weakly and strongly labeled data. Strongly labeled data can be simply understood as fully…

Machine Learning · Computer Science 2017-02-21 Anurag Kumar , Bhiksha Raj

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on…

The development of models for quality prediction of both audio and video signals is a fairly mature field. But, although several multimodal models have been proposed, the area of audio-visual quality prediction is still an emerging area. In…

Multimedia · Computer Science 2020-02-06 Helard Martinez , M. C. Farias , A. Hines