English
Related papers

Related papers: ViSpeechFormer: A Phonemic Approach for Vietnamese…

200 papers

In this paper, we aimed to develop a neural parser for Vietnamese based on simplified Head-Driven Phrase Structure Grammar (HPSG). The existing corpora, VietTreebank and VnDT, had around 15% of constituency and dependency tree pairs that…

Computation and Language · Computer Science 2025-04-29 Duc-Vu Nguyen , Thang Chau Phan , Quoc-Nam Nguyen , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

Vietnamese Speech Emotion Recognition (SER) remains challenging due to ambiguous acoustic patterns and the lack of reliable annotated data, especially in real-world conditions where emotional boundaries are not clearly separable. To address…

Computation and Language · Computer Science 2026-04-03 Truc Nguyen , Then Tran , Binh Truong , Phuoc Nguyen T. H

Training automatic speech recognition (ASR) systems requires large amounts of data in the target language in order to achieve good performance. Whereas large training corpora are readily available for languages like English, there exists a…

Audio and Speech Processing · Electrical Eng. & Systems 2017-11-15 Markus Müller , Sebastian Stüker , Alex Waibel

Text summarization is a challenging task within natural language processing that involves text generation from lengthy input sequences. While this task has been widely studied in English, there is very limited research on summarization for…

Computation and Language · Computer Science 2021-10-11 Hieu Nguyen , Long Phan , James Anibal , Alec Peltekian , Hieu Tran

We investigate what linguistic factors affect the performance of Automatic Speech Recognition (ASR) models. We hypothesize that orthographic and phonological complexities both degrade accuracy. To examine this, we fine-tune the multilingual…

Computation and Language · Computer Science 2024-06-14 Chihiro Taguchi , David Chiang

Transformers have recently become very popular for sequence-to-sequence applications such as machine translation and speech recognition. In this work, we propose a multi-task learning-based transformer model for low-resource multilingual…

Computation and Language · Computer Science 2021-09-13 Krishna D N

Recent research in speaker recognition aims to address vulnerabilities due to variations between enrolment and test utterances, particularly in the multi-genre phenomenon where the utterances are in different speech genres. Previous…

Sound · Computer Science 2025-01-03 Hoang Long Vu , Phuong Tuan Dat , Pham Thao Nhi , Nguyen Song Hao , Nguyen Thi Thu Trang

Audio-Visual Speech Recognition (AVSR) models have surpassed their audio-only counterparts in terms of performance. However, the interpretability of AVSR systems, particularly the role of the visual modality, remains under-explored. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-06 Aristeidis Papadopoulos , Naomi Harte

In this work, we propose a speaker anonymization pipeline that leverages high quality automatic speech recognition and synthesis systems to generate speech conditioned on phonetic transcriptions and anonymized speaker embeddings. Using…

Sound · Computer Science 2022-07-12 Sarina Meyer , Florian Lux , Pavel Denisov , Julia Koch , Pascal Tilli , Ngoc Thang Vu

Phoneme-based ASR factorizes recognition into speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G), enabling cross-lingual acoustic sharing while keeping language-specific orthography in a separate module. While large language models…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-01 Lukuang Dong , Ziwei Li , Saierdaer Yusuyin , Xianyu Zhao , Zhijian Ou

This paper presents a high quality Vietnamese speech corpus that can be used for analyzing Vietnamese speech characteristic as well as building speech synthesis models. The corpus consists of 5400 clean-speech utterances spoken by 12…

Computation and Language · Computer Science 2019-04-12 Pham Ngoc Phuong , Quoc Truong Do , Luong Chi Mai

Integrating pretrained speech encoders with large language models (LLMs) is promising for ASR, but performance and data efficiency depend on the speech-language interface. A common choice is a learned projector that maps encoder features…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-13 Ziwei Li , Lukuang Dong , Saierdaer Yusuyin , Xianyu Zhao , Zhijian Ou

This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g.,…

Speaker-independent VSR is a complex task that involves identifying spoken words or phrases from video recordings of a speaker's facial movements. Over the years, there has been a considerable amount of research in the field of VSR…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Praneeth Nemani , G. Sai Krishna , Supriya Kundrapu

Determining the difficulty of a text involves assessing various textual features that may impact the reader's text comprehension, yet current research in Vietnamese has only focused on statistical features. This paper introduces a new…

Computation and Language · Computer Science 2024-11-08 Hung Tuan Le , Long Truong To , Manh Trong Nguyen , Quyen Nguyen , Trong-Hop Do

Recent years have witnessed significant improvement in ASR systems to recognize spoken utterances. However, it is still a challenging task for noisy and out-of-domain data, where substitution and deletion errors are prevalent in the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Mukuntha Narayanan Sundararaman , Ayush Kumar , Jithendra Vepa

Phonemic or phonetic sub-word units are the most commonly used atomic elements to represent speech signals in modern ASRs. However they are not the optimal choice due to several reasons such as: large amount of effort required to handcraft…

Computation and Language · Computer Science 2016-06-17 Naoya Takahashi , Tofigh Naghibi , Beat Pfister

In automatic speech recognition (ASR), phoneme-based multilingual pre-training and crosslingual fine-tuning is attractive for its high data efficiency and competitive results compared to subword-based models. However, Weighted Finite State…

Sound · Computer Science 2025-06-06 Te Ma , Min Bi , Saierdaer Yusuyin , Hao Huang , Zhijian Ou

Recent advancements in machine learning have significantly improved speech recognition, but recognizing speech from non-fluent or accented speakers remains a challenge. Previous efforts, relying on rule-based pronunciation patterns, have…

Computation and Language · Computer Science 2025-06-04 Anna Seo Gyeong Choi , Jonghyeon Park , Myungwoo Oh

Automatic speech recognition (ASR) is a crucial tool for linguists aiming to perform a variety of language documentation tasks. However, modern ASR systems use data-hungry transformer architectures, rendering them generally unusable for…

Computation and Language · Computer Science 2025-10-09 Massimo Daul , Alessio Tosolini , Claire Bowern