中文
相关论文

相关论文: DSA-Tokenizer: Disentangled Semantic-Acoustic Toke…

200 篇论文

The rapid advancement of speech generation technologies in the era of large language models (LLMs) has established discrete speech tokens as a foundational paradigm for speech representation. These tokens, characterized by their discrete,…

音频与语音处理 · 电气工程与系统科学 2025-12-15 Yiwei Guo , Zhihan Li , Hankun Wang , Bohan Li , Chongtian Shao , Hanglei Zhang , Chenpeng Du , Xie Chen , Shujie Liu , Kai Yu

Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semantic tokens, with the latter being more advantageous for…

音频与语音处理 · 电气工程与系统科学 2025-12-04 Mohan Shi , Natarajan Balaji Shankar , Kaiyuan Zhang , Zilai Wang , Abeer Alwan

Neural speech models build deeply entangled internal representations, which capture a variety of features (e.g., fundamental frequency, loudness, syntactic category, or semantic content of a word) in a distributed encoding. This complexity…

计算与语言 · 计算机科学 2024-10-07 Hosein Mohebbi , Grzegorz Chrupała , Willem Zuidema , Afra Alishahi , Ivan Titov

Diffusion-based generative models have exhibited powerful generative performance in recent years. However, as many attributes exist in the data distribution and owing to several limitations of sharing the model parameters across all levels…

音频与语音处理 · 电气工程与系统科学 2023-05-26 Ha-Yeong Choi , Sang-Hoon Lee , Seong-Whan Lee

Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcription. However, these models are not designed for streaming…

计算与语言 · 计算机科学 2026-04-07 Tomer Krichli , Bhiksha Raj , Joseph Keshet

Most existing audio-text retrieval (ATR) approaches typically rely on a single-level interaction to associate audio and text, limiting their ability to align different modalities and leading to suboptimal matches. In this work, we present a…

声音 · 计算机科学 2025-05-06 Yifei Xin , Zhihong Zhu , Xuxin Cheng , Xusheng Yang , Yuexian Zou

When working with textual data, a natural application of disentangled representations is fair classification where the goal is to make predictions without being biased (or influenced) by sensitive attributes that may be present in the data…

计算与语言 · 计算机科学 2022-10-10 Pierre Colombo , Guillaume Staerman , Nathan Noiry , Pablo Piantanida

ASR systems often struggle with maintaining syntactic and semantic accuracy in long audio transcripts, impacting tasks like Named Entity Recognition (NER), capitalization, and punctuation. We propose a novel approach that enhances ASR by…

计算与语言 · 计算机科学 2025-08-20 Duygu Altinok

Speech dysfluency modeling is a task to detect dysfluencies in speech, such as repetition, block, insertion, replacement, and deletion. Most recent advancements treat this problem as a time-based object detection problem. In this work, we…

Domain mismatch problem caused by speaker-unrelated feature has been a major topic in speaker recognition. In this paper, we propose an explicit disentanglement framework to unravel speaker-relevant features from speaker-unrelated features…

音频与语音处理 · 电气工程与系统科学 2022-10-13 Sung Hwan Mun , Min Hyun Han , Minchan Kim , Dongjune Lee , Nam Soo Kim

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these representations, such…

声音 · 计算机科学 2026-02-05 Takanori Ashihara , Shota Horiguchi , Kohei Matsuura , Tsubasa Ochiai , Marc Delcroix

Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility…

声音 · 计算机科学 2025-06-03 Xueyuan Chen , Dongchao Yang , Wenxuan Wu , Minglin Wu , Jing Xu , Xixin Wu , Zhiyong Wu , Helen Meng

Data augmentation is vital to the generalization ability and robustness of deep neural networks (DNNs) models. Existing augmentation methods for speaker verification manipulate the raw signal, which are time-consuming and the augmented…

音频与语音处理 · 电气工程与系统科学 2023-10-19 Yuanyuan Wang , Yang Zhang , Zhiyong Wu , Zhihan Yang , Tao Wei , Kun Zou , Helen Meng

Recently, discrete tokens derived from self-supervised learning (SSL) models via k-means clustering have been actively studied as pseudo-text in speech language models and as efficient intermediate representations for various tasks.…

声音 · 计算机科学 2025-08-18 Kentaro Onda , Satoru Fukayama , Daisuke Saito , Nobuaki Minematsu

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Hui Fu , Zeqing Wang , Ke Gong , Keze Wang , Tianshui Chen , Haojie Li , Haifeng Zeng , Wenxiong Kang

Machine Speech Chain, simulating the human perception-production loop, proves effective in jointly improving ASR and TTS. We propose TokenChain, a fully discrete speech chain coupling semantic-token ASR with a two-stage TTS: an…

音频与语音处理 · 电气工程与系统科学 2026-05-05 Mingxuan Wang , Satoshi Nakamura

Despite speaker verification has achieved significant performance improvement with the development of deep neural networks, domain mismatch is still a challenging problem in this field. In this study, we propose a novel framework to…

音频与语音处理 · 电气工程与系统科学 2021-02-24 Mufan Sang , Wei Xia , John H. L. Hansen

Disentangling the encodings of neural models is a fundamental aspect for improving interpretability, semantic control and downstream task performance in Natural Language Processing. Currently, most disentanglement methods are unsupervised…

计算与语言 · 计算机科学 2023-02-17 Danilo S. Carvalho , Giangiacomo Mercatali , Yingji Zhang , Andre Freitas

Speech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate…

计算与语言 · 计算机科学 2019-07-12 Laurent El Shafey , Hagen Soltau , Izhak Shafran

Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplify this process by introducing a semantic knowledge…

声音 · 计算机科学 2024-09-18 Gerard I. Gállego , Roy Fejgin , Chunghsin Yeh , Xiaoyu Liu , Gautam Bhattacharya