English
Related papers

Related papers: A Proposal for Foley Sound Synthesis Challenge

200 papers

We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective,…

Artificial intelligence (AI) has been widely applied to music generation topics such as continuation, melody/harmony generation, genre transfer and music infilling application. Although with the burst interest to apply AI to music, there…

Sound · Computer Science 2022-03-25 Rui Guo

This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The proposed system is a score-level fusion of several light…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-14 Yuxiang Zhang , Jingze Lu , Xingming Wang , Zhuo Li , Runqiu Xiao , Wenchao Wang , Ming Li , Pengyuan Zhang

Audio textures are a subset of environmental sounds, often defined as having stable statistical characteristics within an adequately large window of time but may be unstructured locally. They include common everyday sounds such as from…

Sound · Computer Science 2020-11-26 M. Huzaifah , L. Wyse

Current audio generation conditioned by text or video focuses on aligning audio with text/video modalities. Despite excellent alignment results, these multimodal frameworks still cannot be directly applied to compelling movie storytelling…

Sound · Computer Science 2025-06-03 Zixuan Wang , Chi-Keung Tang , Yu-Wing Tai

Acoustic radiation force and torque arising from wave scattering are commonly used to manipulate micro-objects without contact. We applied the partial wave expansion series and the conformal transformation approach to estimate the radiation…

Computational Engineering, Finance, and Science · Computer Science 2022-11-16 Tianquan Tang , Lixi Huang

A number of recent advances in neural audio synthesis rely on upsampling layers, which can introduce undesired artifacts. In computer vision, upsampling artifacts have been studied and are known as checkerboard artifacts (due to their…

Sound · Computer Science 2021-02-10 Jordi Pons , Santiago Pascual , Giulio Cengarle , Joan Serrà

We present Synthio, a novel approach for augmenting small-scale audio classification datasets with synthetic data. Our goal is to improve audio classification accuracy with limited labeled data. Traditional data augmentation techniques,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-13 Sreyan Ghosh , Sonal Kumar , Zhifeng Kong , Rafael Valle , Bryan Catanzaro , Dinesh Manocha

Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media, ranging from movies to social media posts. Machine learning models that can synthesize music are…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Sanjoy Chowdhury , Sayan Nag , K J Joseph , Balaji Vasan Srinivasan , Dinesh Manocha

This demo is about automatic authoring of various motion effects that are provided with audiovisual content to improve user experiences. Traditionally, motion effects have been used for simulators, e.g., flight simulators for pilots and…

Human-Computer Interaction · Computer Science 2024-11-11 Jiwan Lee , Seungmoon Choi

Recent advancements in AI, especially deep learning, have contributed to a significant increase in the creation of new realistic-looking synthetic media (video, image, and audio) and manipulation of existing media, which has led to the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Enes Altuncu , Virginia N. L. Franqueira , Shujun Li

Inspired by a concrete industry problem we consider the input synthesis problem for hybrid systems: given a hybrid system that is subject to input from outside (also called disturbance or noise), find an input sequence that steers the…

Systems and Control · Computer Science 2016-02-22 Takumi Akazaki , Ichiro Hasuo , Kohei Suenaga

The growing prominence of the field of audio deepfake detection is driven by its wide range of applications, notably in protecting the public from potential fraud and other malicious activities, prompting the need for greater attention and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-12 Jiangyan Yi , Chu Yuan Zhang , Jianhua Tao , Chenglong Wang , Xinrui Yan , Yong Ren , Hao Gu , Junzuo Zhou

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

Sound · Computer Science 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

Systems for synthesizer sound matching, which automatically set the parameters of a synthesizer to emulate an input sound, have the potential to make the process of synthesizer programming faster and easier for novice and experienced…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-24 Fred Bruford , Frederik Blang , Shahan Nercessian

The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused…

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

Inspired by the role sound and friction play in interactions with everyday objects, this work aims to identify some of the ways in which kinetic surface friction rendering can complement interactive sonification controlled by movable…

Human-Computer Interaction · Computer Science 2021-07-30 Staas de Jong

Despite surveillance systems are becoming increasingly ubiquitous in our living environment, automated surveillance, currently based on video sensory modality and machine intelligence, lacks most of the time the robustness and reliability…

Sound · Computer Science 2014-09-30 Marco Crocco , Marco Cristani , Andrea Trucco , Vittorio Murino

The ability to localize and track acoustic events is a fundamental prerequisite for equipping machines with the ability to be aware of and engage with humans in their surrounding environment. However, in realistic scenarios, audio signals…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-22 Christine Evers , Heinrich Loellmann , Heinrich Mellmann , Alexander Schmidt , Hendrik Barfuss , Patrick Naylor , Walter Kellermann
‹ Prev 1 8 9 10 Next ›