English
Related papers

Related papers: VoxLingua107: a Dataset for Spoken Language Recogn…

200 papers

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to…

Computer Vision and Pattern Recognition · Computer Science 2016-02-22 Hao Fang , Saurabh Gupta , Forrest Iandola , Rupesh Srivastava , Li Deng , Piotr Dollár , Jianfeng Gao , Xiaodong He , Margaret Mitchell , John C. Platt , C. Lawrence Zitnick , Geoffrey Zweig

The driving factors behind the development of large language models (LLMs) with impressive learning capabilities are their colossal model sizes and extensive training datasets. Along with the progress in natural language processing, LLMs…

Computation and Language · Computer Science 2023-09-19 Thuat Nguyen , Chien Van Nguyen , Viet Dac Lai , Hieu Man , Nghia Trung Ngo , Franck Dernoncourt , Ryan A. Rossi , Thien Huu Nguyen

This article introduces Mi-Go, a novel testing framework aimed at evaluating the performance and adaptability of general-purpose speech recognition machine learning models across diverse real-world scenarios. The framework leverages YouTube…

Sound · Computer Science 2023-09-04 Tomasz Wojnar , Jaroslaw Hryszko , Adam Roman

We learn rich natural sound representations by capitalizing on large amounts of unlabeled sound data collected in the wild. We leverage the natural synchronization between vision and sound to learn an acoustic representation using…

Computer Vision and Pattern Recognition · Computer Science 2016-10-31 Yusuf Aytar , Carl Vondrick , Antonio Torralba

Almost none of the 2,000+ languages spoken in Africa have widely available automatic speech recognition systems, and the required data is also only available for a few languages. We have experimented with two techniques which may provide…

Computation and Language · Computer Science 2022-10-05 Sandy Ritchie , You-Chi Cheng , Mingqing Chen , Rajiv Mathews , Daan van Esch , Bo Li , Khe Chai Sim

Keyphrases are useful for a variety of purposes, including summarizing, indexing, labeling, categorizing, clustering, highlighting, browsing, and searching. The task of automatic keyphrase extraction is to select keyphrases from within the…

Machine Learning · Computer Science 2007-05-23 Peter D. Turney

We conducted a data collection on the basis of the Google AudioSet database by selecting a subset of the samples annotated with \textit{laughter}. The selection criterion was to be present a communicative act with clear connotation of being…

Sound · Computer Science 2023-05-24 Aljoscha Düsterhöft , Felix Burkhardt , Björn W. Schuller

Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is…

Computation and Language · Computer Science 2024-10-10 Gautam Kishore Shahi , Tim A. Majchrzak

This work presents an extensive and detailed study on Audio-Visual Speech Recognition (AVSR) for five widely spoken languages: Chinese, Spanish, English, Arabic, and French. We have collected large-scale datasets for each language except…

Computation and Language · Computer Science 2024-06-04 Sanath Narayan , Yasser Abdelaziz Dahou Djilali , Ankit Singh , Eustache Le Bihan , Hakim Hacid

Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack generalization to the full spectrum of human diversity in ethnicity, language, and age…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Shunian Chen , Hejin Huang , Yexin Liu , Zihan Ye , Pengcheng Chen , Chenghao Zhu , Michael Guan , Rongsheng Wang , Junying Chen , Guanbin Li , Ser-Nam Lim , Harry Yang , Benyou Wang

Zero-shot learning aims to recognize unseen objects using their semantic representations. Most existing works use visual attributes labeled by humans, not suitable for large-scale applications. In this paper, we revisit the use of documents…

Computer Vision and Pattern Recognition · Computer Science 2021-04-22 Jihyung Kil , Wei-Lun Chao

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 418 thousand hours of…

Computation and Language · Computer Science 2022-11-10 Paul-Ambroise Duquenne , Hongyu Gong , Ning Dong , Jingfei Du , Ann Lee , Vedanuj Goswani , Changhan Wang , Juan Pino , Benoît Sagot , Holger Schwenk

Large Language Models (LLMs) are trained on vast amounts of data, most of which is automatically scraped from the internet. This data includes encyclopedic documents that harbor a vast amount of general knowledge (e.g., Wikipedia) but also…

Large datasets as required for deep learning of lip reading do not exist in many languages. In this paper we present the dataset GLips (German Lips) consisting of 250,000 publicly available videos of the faces of speakers of the Hessian…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Gerald Schwiebert , Cornelius Weber , Leyuan Qu , Henrique Siqueira , Stefan Wermter

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

Random Indexing is a simple implementation of Random Projections with a wide range of applications. It can solve a variety of problems with good accuracy without introducing much complexity. Here we use it for identifying the language of…

Computation and Language · Computer Science 2015-03-02 Aditya Joshi , Johan Halseth , Pentti Kanerva

Recent advancements in speech generation have been driven by large-scale training datasets. However, current models struggle to capture the spontaneity and variability inherent in real-world human speech, as they are primarily trained on…

This paper describes the WiLI-2018 benchmark dataset for monolingual written natural language identification. WiLI-2018 is a publicly available, free of charge dataset of short text extracts from Wikipedia. It contains 1000 paragraphs of…

Computer Vision and Pattern Recognition · Computer Science 2018-01-25 Martin Thoma

Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as…

Machine Learning · Computer Science 2022-06-24 Bowen Baker , Ilge Akkaya , Peter Zhokhov , Joost Huizinga , Jie Tang , Adrien Ecoffet , Brandon Houghton , Raul Sampedro , Jeff Clune
‹ Prev 1 8 9 10 Next ›