中文
相关论文

相关论文: Audio-JEPA: Joint-Embedding Predictive Architectur…

200 篇论文

We study whether 3D self-supervised pretraining with Point--JEPA enables label-efficient grasp joint-angle prediction. Meshes are sampled to point clouds and tokenized; a ShapeNet-pretrained Point--JEPA encoder feeds a $K{=}5$…

机器人学 · 计算机科学 2025-09-26 Jed Guzelkabaagac , Boris Petrović

The proliferation of audio deepfakes poses a growing threat to trust in digital communications. While detection methods have advanced, attributing audio deepfakes to their source models remains an underexplored yet crucial challenge. In…

声音 · 计算机科学 2025-10-13 Andrea Di Pierno , Luca Guarnera , Dario Allegra , Sebastiano Battiato

In this paper, we propose AUREXA-SE (Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement), a progressive bimodal framework tailored for audio-visual speech enhancement…

Recently, autoregressive models have demonstrated remarkable performance in class-conditional image generation. However, the application of next-token prediction to high-resolution text-to-image generation remains largely unexplored. In…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Dengsheng Chen , Jie Hu , Tiezhu Yue , Xiaoming Wei , Enhua Wu

We present MeFEm, a vision model based on a modified Joint Embedding Predictive Architecture (JEPA) for biometric and medical analysis from facial images. Key modifications include an axial stripe masking strategy to focus learning on…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Yury Borets , Stepan Botman

This paper proposes joint speaker feature learning methods for zero-shot adaptation of audio-visual multichannel speech separation and recognition systems. xVector and ECAPA-TDNN speaker encoders are connected using purpose-built fusion…

Methods for extracting audio and speech features have been studied since pioneering work on spectrum analysis decades ago. Recent efforts are guided by the ambition to develop general-purpose audio representations. For example, deep neural…

Vision-based pose estimation of articulated robots with unknown joint angles has applications in collaborative robotics and human-robot interaction tasks. Current frameworks use neural network encoders to extract image features and…

机器人学 · 计算机科学 2025-05-05 Raktim Gautam Goswami , Prashanth Krishnamurthy , Yann LeCun , Farshad Khorrami

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provided textual sentences,…

图形学 · 计算机科学 2024-01-19 Jeongsoo Choi , Minsu Kim , Se Jin Park , Yong Man Ro

In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked…

声音 · 计算机科学 2024-10-18 Ashish Seth , Ramaneswaran Selvakumar , S Sakshi , Sonal Kumar , Sreyan Ghosh , Dinesh Manocha

While self-supervised learning (SSL) has revolutionized audio representation, the excessive parameterization and quadratic computational cost of standard Transformers limit their deployment on resource-constrained devices. To address this…

声音 · 计算机科学 2026-03-30 Harunori Kawano , Takeshi Sasaki

Recent advances in machine learning (ML) have shown promise in accelerating the discovery of polymers with desired properties by aiding in tasks such as virtual screening via property prediction. However, progress in polymer ML is hampered…

机器学习 · 计算机科学 2025-06-25 Francesco Piccoli , Gabriel Vogel , Jana M. Weber

Sound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges. However, they still rely on an RNN-based context network to model temporal…

声音 · 计算机科学 2024-08-20 Pengfei Cai , Yan Song , Kang Li , Haoyu Song , Ian McLoughlin

In this paper, we present MixRep, a simple and effective data augmentation strategy based on mixup for low-resource ASR. MixRep interpolates the feature dimensions of hidden representations in the neural network that can be applied to both…

音频与语音处理 · 电气工程与系统科学 2025-06-19 Jiamin Xie , John H. L. Hansen

With the advent of Joint Embedding Predictive Architectures (JEPAs), which appear to be more capable than reconstruction-based methods, this paper introduces a novel technique for creating world models using continuous-time dynamic systems…

机器学习 · 计算机科学 2025-08-15 Jonas Ulmen , Ganesh Sundaram , Daniel Görges

Autonomous driving, as an agent operating in the physical world, requires the fundamental capability to build \textit{world models} that capture how the environment evolves spatiotemporally in order to support long-term planning. At the…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Haoran Zhu , Anna Choromanska

Given the strong results of self-supervised models on various tasks, there have been surprisingly few studies exploring self-supervised representations for acoustic word embeddings (AWE), fixed-dimensional vectors representing…

计算与语言 · 计算机科学 2023-03-16 Ramon Sanabria , Hao Tang , Sharon Goldwater

Spoken language understanding is typically based on pipeline architectures including speech recognition and natural language understanding steps. These components are optimized independently to allow usage of available data, but the overall…

音频与语音处理 · 电气工程与系统科学 2020-08-13 Pavel Denisov , Ngoc Thang Vu

We introduce COLA, a self-supervised pre-training approach for learning a general-purpose representation of audio. Our approach is based on contrastive learning: it learns a representation which assigns high similarity to audio segments…

声音 · 计算机科学 2020-10-22 Aaqib Saeed , David Grangier , Neil Zeghidour

Self-supervised learning has become an incredibly successful method for feature learning, widely applied to many downstream tasks. It has proven especially effective for discriminative tasks, surpassing the trending generative models.…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Yuping Qiu , Rui Zhu , Ying-cong Chen