中文
相关论文

相关论文: Finnish Dialect Identification: The Effect of Audi…

200 篇论文

Recent advances in speech synthesis have enabled many useful applications like audio directions in Google Maps, screen readers, and automated content generation on platforms like TikTok. However, these systems are mostly dominated by voices…

音频与语音处理 · 电气工程与系统科学 2024-06-28 Sewade Ogun , Abraham T. Owodunni , Tobi Olatunji , Eniola Alese , Babatunde Oladimeji , Tejumade Afonja , Kayode Olaleye , Naome A. Etori , Tosin Adewumi

The historical and geographical spread from older to more modern languages has long been studied by examining textual changes and in terms of changes in phonetic transcriptions. However, it is more difficult to analyze language change from…

应用统计 · 统计学 2017-05-19 Davide Pigoli , Pantelis Z. Hadjipantelis , John S. Coleman , John A. D. Aston

A vast majority of the world's 7,000 spoken languages are predicted to become extinct within this century, including the endangered language of Ladin from the Italian Alps. Linguists who work to preserve a language's phonetic and…

音频与语音处理 · 电气工程与系统科学 2021-08-31 Zane Durante , Leena Mathur , Eric Ye , Sichong Zhao , Tejas Ramdas , Khalil Iskarous

Here we present a work aimed on efficiently creating textual language dialects and supporting tools for them (e.g. compiler front-ends, IDE support, pretty-printers, etc.). A dialect is a language which may be described with a (relatively…

编程语言 · 计算机科学 2009-06-10 Andrey Breslav

The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what extent can low error rates on academic benchmarks translate…

音频与语音处理 · 电气工程与系统科学 2025-05-23 Ashi Garg , Zexin Cai , Lin Zhang , Henry Li Xinyuan , Leibny Paola García-Perera , Kevin Duh , Sanjeev Khudanpur , Matthew Wiesner , Nicholas Andrews

Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data…

Existing fake audio detection systems perform well in in-domain testing, but still face many challenges in out-of-domain testing. This is due to the mismatch between the training and test data, as well as the poor generalizability of…

声音 · 计算机科学 2023-05-24 Chenglong Wang , Jiangyan Yi , Jianhua Tao , Chuyuan Zhang , Shuai Zhang , Xun Chen

Recent progress in generative AI has made it increasingly easy to create natural-sounding deepfake speech from just a few seconds of audio. While these tools support helpful applications, they also raise serious concerns by making it…

音频与语音处理 · 电气工程与系统科学 2025-08-07 Xi Xuan , Yang Xiao , Rohan Kumar Das , Tomi Kinnunen

Spoofed audio, i.e. audio that is manipulated or AI-generated deepfake audio, is difficult to detect when only using acoustic features. Some recent innovative work involving AI-spoofed audio detection models augmented with phonetic and…

声音 · 计算机科学 2024-10-22 Zahra Khanjani , Christine Mallinson , James Foulds , Vandana P Janeja

In this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of different languages together for multilingual generalization…

音频与语音处理 · 电气工程与系统科学 2021-06-17 Roza Chojnacka , Jason Pelecanos , Quan Wang , Ignacio Lopez Moreno

In recent years, using raw waveforms as input for deep networks has been widely explored for the speaker verification system. For example, RawNet and RawNet2 extracted speaker's feature embeddings from waveforms automatically for…

音频与语音处理 · 电气工程与系统科学 2021-10-08 Jin Li , Nan Yan , Lan Wang

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

计算与语言 · 计算机科学 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

Curating datasets that span multiple languages is challenging. To make the collection more scalable, researchers often incorporate one or more imperfect classifiers in the process, like language identification models. These models, however,…

计算与语言 · 计算机科学 2024-10-08 Farhan Samir , Emily P. Ahn , Shreya Prakash , Márton Soskuthy , Vered Shwartz , Jian Zhu

The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and…

声音 · 计算机科学 2025-05-22 Kunyang Huang , Bin Hu

Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored…

计算与语言 · 计算机科学 2025-03-13 Marco Giordano , Claudia Rinaldi

We introduce Faina, the first dataset for fallacy detection that embraces multiple plausible answers and natural disagreement. Faina includes over 11K span-level annotations with overlaps across 20 fallacy types on social media posts in…

计算与语言 · 计算机科学 2025-02-20 Alan Ramponi , Agnese Daffara , Sara Tonelli

Keyword spotting systems often struggle to generalize to a diverse population with various accents and age groups. To address this challenge, we propose a novel approach that integrates speaker information into keyword spotting using…

音频与语音处理 · 电气工程与系统科学 2023-11-08 Beltrán Labrador , Pai Zhu , Guanlong Zhao , Angelo Scorza Scarpati , Quan Wang , Alicia Lozano-Diez , Alex Park , Ignacio López Moreno

Despite impressive performance on many text classification tasks, deep neural networks tend to learn frequent superficial patterns that are specific to the training data and do not always generalize well. In this work, we observe this…

机器学习 · 计算机科学 2021-06-16 Sachin Kumar , Shuly Wintner , Noah A. Smith , Yulia Tsvetkov

This paper analyzes dysarthric speech datasets from three languages with different prosodic systems: English, Korean, and Tamil. We inspect 39 acoustic measurements which reflect three speech dimensions including voice quality,…

计算与语言 · 计算机科学 2022-11-04 Eun Jung Yeo , Sunhee Kim , Minhwa Chung

In conversational speech, the acoustic signal provides cues that help listeners disambiguate difficult parses. For automatically parsing spoken utterances, we introduce a model that integrates transcribed text and acoustic-prosodic features…

计算与语言 · 计算机科学 2018-04-17 Trang Tran , Shubham Toshniwal , Mohit Bansal , Kevin Gimpel , Karen Livescu , Mari Ostendorf