English
Related papers

Related papers: VoxVietnam: a Large-Scale Multi-Genre Dataset for …

200 papers

Vietnamese, a low-resource language, is typically categorized into three primary dialect groups that belong to Northern, Central, and Southern Vietnam. However, each province within these regions exhibits its own distinct pronunciation…

Computation and Language · Computer Science 2024-10-07 Nguyen Van Dinh , Thanh Chi Dang , Luan Thanh Nguyen , Kiet Van Nguyen

Research on speaker recognition is extending to address the vulnerability in the wild conditions, among which genre mismatch is perhaps the most challenging, for instance, enrollment with reading speech while testing with conversational or…

Due to privacy restrictions, there's a shortage of publicly available speech recognition datasets in the medical domain. In this work, we present VietMed - a Vietnamese speech recognition dataset in the medical domain comprising 16h of…

Computation and Language · Computer Science 2025-04-07 Khai Le-Duc

The goal of this work is to investigate the performance of popular speaker recognition models on speech segments from movies, where often actors intentionally disguise their voice to play a character. We make the following three…

Sound · Computer Science 2021-02-12 Andrew Brown , Jaesung Huh , Arsha Nagrani , Joon Son Chung , Andrew Zisserman

The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media.…

Sound · Computer Science 2020-11-05 Joon Son Chung , Arsha Nagrani , Andrew Zisserman

This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on predicting mean opinion score of naturalness, leaving…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-29 Junseok Ahn , Youkyum Kim , Yeunju Choi , Doyeop Kwak , Ji-Hoon Kim , Seongkyu Mun , Joon Son Chung

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental in developing speaker…

Sound · Computer Science 2023-02-28 Saqlain Hussain Shah , Muhammad Saad Saeed , Shah Nawaz , Muhammad Haroon Yousaf

In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This dataset represents a significant expansion over the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-17 Yuke Lin , Ming Cheng , Fulin Zhang , Yingying Gao , Shilei Zhang , Ming Li

Vietnamese, the 20th most spoken language with over 102 million native speakers, lacks robust resources for key natural language processing tasks such as text segmentation and machine reading comprehension (MRC). To address this gap, we…

Computation and Language · Computer Science 2025-06-23 Toan Nguyen Hai , Ha Nguyen Viet , Truong Quan Xuan , Duc Do Minh

Vietnamese medical research has become an increasingly vital domain, particularly with the rise of intelligent technologies aimed at reducing time and resource burdens in clinical diagnosis. Recent advances in vision-language models (VLMs),…

Computation and Language · Computer Science 2026-03-17 Duy Vu Minh Nguyen , Chinh Thanh Truong , Phuc Hoang Tran , Hung Tuan Le , Nguyen Van-Thanh Dat , Trung Hieu Pham , Kiet Van Nguyen

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

The performance of speaker verification systems is adversely affected by speaker aging. However, due to challenges in data collection, particularly the lack of sustained and large-scale longitudinal data for individuals, research on speaker…

Sound · Computer Science 2025-05-28 Zhiqi Ai , Meixuan Bao , Zhiyong Chen , Zhi Yang , Xinnuo Li , Shugong Xu

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-06 Yue Fan , Jiawen Kang , Lantian Li , Kaicheng Li , Haolin Chen , Sitong Cheng , Pengyuan Zhang , Ziya Zhou , Yunqi Cai , Dong Wang

The rapid spread of information in the digital age highlights the critical need for effective fact-checking tools, particularly for languages with limited resources, such as Vietnamese. In response to this challenge, we introduce…

Computation and Language · Computer Science 2024-12-23 Tran Thai Hoa , Tran Quang Duy , Khanh Quoc Tran , Kiet Van Nguyen

The VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020 offers a challenging evaluation for speaker recognition systems, which includes celebrities playing different parts in movies. The goal of this work is robust speaker…

Sound · Computer Science 2020-10-30 Yoohwan Kwon , Hee-Soo Heo , Bong-Jin Lee , Joon Son Chung

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

Sound · Computer Science 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Over 97 million people speak Vietnamese as their native language in the world. However, there are few research studies on machine reading comprehension (MRC) for Vietnamese, the task of understanding a text and answering questions related…

Computation and Language · Computer Science 2020-11-10 Kiet Van Nguyen , Duc-Vu Nguyen , Anh Gia-Tuan Nguyen , Ngan Luu-Thuy Nguyen

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-14 Yuke Lin , Xiaoyi Qin , Guoqing Zhao , Ming Cheng , Ning Jiang , Haiyang Wu , Ming Li

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains,…

‹ Prev 1 2 3 10 Next ›