中文
相关论文

相关论文: Unifying Speech Editing Detection and Content Loca…

200 篇论文

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. While audio-driven…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Changpeng Cai , Guinan Guo , Jiao Li , Junhao Su , Fei Shen , Chenghao He , Jing Xiao , Yuanxu Chen , Lei Dai , Feiyu Zhu

Comparing spoken segments is a central operation to speech processing. Traditional approaches in this area have favored frame-level dynamic programming algorithms, such as dynamic time warping, because they require no supervision, but they…

计算与语言 · 计算机科学 2023-08-30 Shane Settle

This research introduces a transformative framework for integrating Vision-Enhanced Large Language Models (LLMs) with advanced transformer-based architectures to tackle challenges in high-resolution image synthesis and multimodal data…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Karthikeya KV

In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to segment the audio and then transcribe each segmentation. These two stages are addressed separately by speaker change detection…

声音 · 计算机科学 2022-11-18 Zhiyun Fan , Zhenlin Liang , Linhao Dong , Yi Liu , Shiyu Zhou , Meng Cai , Jun Zhang , Zejun Ma , Bo Xu

Large language models (LLMs) acquire knowledge during pre-training, but over time, this knowledge may become incorrect or outdated, necessitating updates after training. Knowledge editing techniques address this issue without the need for…

计算与语言 · 计算机科学 2024-10-16 Yuchen Cai , Ding Cao

Deepfake content on social networks is increasingly produced through multiple \emph{sequential} edits to biometric data such as facial imagery. Consequently, the final appearance of an image often reflects a latent chain of operations…

密码学与安全 · 计算机科学 2026-04-14 Mengieong Hoi , Zhedong Zheng , Ping Liu , Wei Liu

Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained…

声音 · 计算机科学 2026-05-28 Yuyue Wang , Xihua Wang , Xin Cheng , Yijing Chen , Ruihua Song

Recent work shows promising results in expanding the capabilities of large language models (LLM) to directly understand and synthesize speech. However, an LLM-based strategy for modeling spoken dialogs remains elusive, calling for further…

Large language model (LLM) decoding involves generating a sequence of tokens based on a given context, where each token is predicted one at a time using the model's learned probabilities. The typical autoregressive decoding method requires…

计算与语言 · 计算机科学 2024-08-20 Xukun Liu , Bowen Lei , Ruqi Zhang , Dongkuan Xu

This research introduces a novel approach to textual and multimodal Hate Speech Detection (HSD), using Large Language Models (LLMs) as dynamic knowledge bases to generate background context and incorporate it into the input of HSD…

计算与语言 · 计算机科学 2025-10-20 Joshua Wolfe Brook , Ilia Markov

Cued Speech (CS) is a visual communication system that combines lip-reading with hand coding to facilitate communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) aims to convert CS hand gestures and lip…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Guanjie Huang , Danny H. K. Tsang , Shan Yang , Guangzhi Lei , Li Liu

With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approaches aim to isolate…

计算与语言 · 计算机科学 2025-06-23 Daejin Jo , Jeeyoung Yun , Byungseok Roh , Sungwoong Kim

Modern Text-to-Speech (TTS) systems increasingly leverage Large Language Model (LLM) architectures to achieve scalable, high-fidelity, zero-shot generation. However, these systems typically rely on fixed-frame-rate acoustic tokenization,…

Lip synchronization and audio-visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, virtual avatars, and telepresence. Despite recent progress,…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Lixiang Lin , Siyuan Jin , Jinshan Zhang

Polyphonic Sound Event Detection (SED) in real-world recordings is a challenging task because of the dynamic polyphony level, intensity, and duration of sound events. Current polyphonic SED systems fail to model the temporal structure of…

音频与语音处理 · 电气工程与系统科学 2019-08-02 Arjun Pankajakshan , Helen L. Bear , Emmanouil Benetos

Conventional speech enhancement (SE) aims to improve speech perception and intelligibility by suppressing noise without requiring enrollment speech as reference, whereas personalized SE (PSE) addresses the cocktail party problem by…

音频与语音处理 · 电气工程与系统科学 2025-05-20 Ziling Huang , Haixin Guan , Yanhua Long

Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Jiarui Hai , Mounya Elhilali

Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend on explicit labels or reference audio, which limits both…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Yun Chen , Qi Chen , Zheqi Dai , Arshdeep Singh , Philip J. B. Jackson , Mark D. Plumbley

Speech emotion recognition systems have high prediction latency because of the high computational requirements for deep learning models and low generalizability mainly because of the poor reliability of emotional measurements across…

声音 · 计算机科学 2023-02-23 Abdul Rehman , Zhen-Tao Liu , Min Wu , Wei-Hua Cao , Cheng-Shan Jiang

Lifelong learning enables large language models (LLMs) to adapt to evolving information by continually updating their internal knowledge. An ideal system should support efficient, wide-ranging updates while preserving existing capabilities…

计算与语言 · 计算机科学 2026-03-11 Xiaojie Gu , Ziying Huang , Jia-Chen Gu , Kai Zhang