中文
相关论文

相关论文: Text2FX: Harnessing CLAP Embeddings for Text-Guide…

200 篇论文

Music recommender systems frequently utilize network-based models to capture relationships between music pieces, artists, and users. Although these relationships provide valuable insights for predictions, new music pieces or artists often…

声音 · 计算机科学 2024-09-16 Florian Grötschla , Luca Strässle , Luca A. Lanzendörfer , Roger Wattenhofer

Recent years have seen an increased interest in establishing association between faces and voices of celebrities leveraging audio-visual information from YouTube. Prior works adopt metric learning methods to learn an embedding space that is…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Muhammad Saad Saeed , Shah Nawaz , Muhammad Haris Khan , Sajid Javed , Muhammad Haroon Yousaf , Alessio Del Bue

Digital audio effects are widely used by audio engineers to alter the acoustic and temporal qualities of audio data. However, these effects can have a large number of parameters which can make them difficult to learn for beginners and…

机器学习 · 计算机科学 2023-10-02 Kieran Grant

Audio texture manipulation involves modifying the perceptual characteristics of a sound to achieve specific transformations, such as adding, removing, or replacing auditory elements. In this paper, we propose an exemplar-based analogy model…

声音 · 计算机科学 2025-01-22 Kan Jen Cheng , Tingle Li , Gopala Anumanchipalli

Language-queried Audio Separation (LASS) employs linguistic queries to isolate target sounds based on semantic descriptions. However, existing methods face challenges in aligning complex auditory features with linguistic context while…

声音 · 计算机科学 2025-06-23 Jianyuan Feng , Guangzheng Li , Yangfei Xu

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are…

音频与语音处理 · 电气工程与系统科学 2025-03-24 Hao Ma , Zhiyuan Peng , Xu Li , Yukai Li , Mingjie Shao , Qiuqiang Kong , Ju Liu

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic…

计算与语言 · 计算机科学 2021-02-08 Yanpei Shi , Thomas Hain

Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio, different people may perceive the same audio differently, resulting in caption disparities (i.e., one audio may correlate to…

声音 · 计算机科学 2022-04-19 Yiming Zhang , Hong Yu , Ruoyi Du , Zhanyu Ma , Yuan Dong

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression,…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Baiqin Wang , Xiangyu Zhu , Fan Shen , Hao Xu , Zhen Lei

The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Yicheng Zhong , Huawei Wei , Peiji Yang , Zhisheng Wang

Speech encodes multiple simultaneous attributes -- linguistic content, speaker identity, dialect, gender --that conventional single-vector embeddings conflate. We present a factor-partitioned embedding framework that maps each utterance…

音频与语音处理 · 电气工程与系统科学 2026-05-11 Jim O'Regan , Jens Edlund

Adversarial input attacks can cause a significant shift of CLIP embeddings. This can affect the downstream robustness of models incorporating CLIP in the pipeline, such as text-to-image generative models or large vision language models.…

In recent years, foundation models have significantly advanced data-driven systems across various domains. Yet, their underlying properties, especially when functioning as feature extractors, remain under-explored. In this paper, we…

机器学习 · 计算机科学 2025-01-28 Victor Deng , Changhong Wang , Gael Richard , Brian McFee

Contrastively pretrained audio-language models (e.g., CLAP) excel at clip-level understanding but struggle with frame-level tasks. Existing extensions fail to exploit the varying granularity of real-world audio-text data, where massive…

声音 · 计算机科学 2026-04-02 Xiquan Li , Xuenan Xu , Ziyang Ma , Wenxi Chen , Haolin He , Qiuqiang Kong , Xie Chen

Automated Audio Captioning (AAC) aims to describe the semantic contexts of general sounds, including acoustic events and scenes, by leveraging effective acoustic features. To enhance performance, an AAC method, EnCLAP, employed discrete…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Daiki Takeuchi , Binh Thien Nguyen , Masahiro Yasuda , Yasunori Ohishi , Daisuke Niizumi , Noboru Harada

The task of converting text input into video content is becoming an important topic for synthetic media generation. Several methods have been proposed with some of them reaching close-to-natural performances in constrained tasks. In this…

音频与语音处理 · 电气工程与系统科学 2022-06-08 Dan Oneata , Beata Lorincz , Adriana Stan , Horia Cucu

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal…

声音 · 计算机科学 2020-11-05 Soo-Whan Chung , Hong Goo Kang , Joon Son Chung

Perceptual similarity representations enable music retrieval systems to determine which songs sound most similar to listeners. State-of-the-art approaches based on task-specific training via self-supervised metric learning show promising…

声音 · 计算机科学 2026-01-28 Arhan Vohra , Taketo Akama

Hair editing is an interesting and challenging problem in computer vision and graphics. Many existing methods require well-drawn sketches or masks as conditional inputs for editing, however these interactions are neither straightforward nor…

计算机视觉与模式识别 · 计算机科学 2022-03-03 Tianyi Wei , Dongdong Chen , Wenbo Zhou , Jing Liao , Zhentao Tan , Lu Yuan , Weiming Zhang , Nenghai Yu

Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it can be desirable to design systems that answer not only…