English
Related papers

Related papers: Mi-Go: Test Framework which uses YouTube as Data S…

200 papers

When evaluating the performance of automatic speech recognition models, usually word error rate within a certain dataset is used. Special care must be taken in understanding the dataset in order to report realistic performance numbers. We…

Computation and Language · Computer Science 2021-05-21 Aashish Agarwal , Torsten Zesch

One of the most challenging scenarios for smart speakers is multi-talker, when target speech from the desired speaker is mixed with interfering speech from one or more speakers. A smart assistant needs to determine which voice to recognize…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-19 Joe Caroselli , Arun Narayanan , Yiteng Huang

Large Language Models are increasingly being deployed to extract structured data from unstructured and semi-structured sources: parsing invoices, medical records, and converting PDF documents to database entries. Yet existing benchmarks for…

Computation and Language · Computer Science 2026-04-29 Abhinav Kumar Singh , Harsha Vardhan Khurdula , Yoeven D Khemlani , Vineet Agarwal

Multi-lingual speech recognition aims to distinguish linguistic expressions in different languages and integrate acoustic processing simultaneously. In contrast, current multi-lingual speech recognition research follows a language-aware…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-28 Yoohwan Kwon , Soo-Whan Chung

Subtitles are essential for video accessibility and audience engagement. Modern Automatic Speech Recognition (ASR) systems, built upon Encoder-Decoder neural network architectures and trained on massive amounts of data, have progressively…

Computation and Language · Computer Science 2025-12-23 Alessandro Lucca , Francesco Pierri

Audio-language pretraining holds promise for general-purpose audio understanding, yet remains underexplored compared to its vision counterpart. While vision-language models like CLIP serve as widely adopted foundations, existing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-24 Wei-Cheng Tseng , Xuanru Zhou , Mingyue Huo , Yiwen Shao , Hao Zhang , Dong Yu

Recent advancement in large language models (LLMs) has offered a strong potential for natural language systems to process informal language. A representative form of informal language is slang, used commonly in daily conversations and…

Computation and Language · Computer Science 2024-04-16 Zhewei Sun , Qian Hu , Rahul Gupta , Richard Zemel , Yang Xu

Misinformation poses a significant threat in today's digital world, often spreading rapidly through platforms like YouTube. This paper introduces a novel approach to combating misinformation by developing an AI-powered system that not only…

Computation and Language · Computer Science 2025-07-17 Cécile Logé , Rehan Ghori

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Alexander Polok , Dominik Klement , Samuele Cornell , Matthew Wiesner , Jan Černocký , Sanjeev Khudanpur , Lukáš Burget

This paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits responsiveness in real-time…

YouTube presents an unprecedented opportunity to explore how machine learning methods can improve healthcare information dissemination. We propose an interdisciplinary lens that synthesizes machine learning methods with healthcare…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Xiao Liu , Bin Zhang , Anjana Susarla , Rema Padman

Exploiting unlabeled data through semi-supervised learning (SSL) or leveraging pre-trained models via fine-tuning are two prevailing paradigms for addressing label-scarce scenarios. Recently, growing attention has been given to combining…

Machine Learning · Computer Science 2026-05-26 Rui Zhu , Song-Lin Lv , Zi-Kang Wang , Lan-Zhe Guo

Establishing retrieval-based dialogue systems that can select appropriate responses from the pre-built index has gained increasing attention from researchers. For this task, the adoption of pre-trained language models (such as BERT) has led…

Computation and Language · Computer Science 2021-10-04 Chongyang Tao , Jiazhan Feng , Chang Liu , Juntao Li , Xiubo Geng , Daxin Jiang

Recognizing and localizing student confusion from video is an important yet challenging problem in educational AI. Existing confusion datasets suffer from noisy labels, coarse temporal annotations, and limited expert validation, which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Lu Dong , Xiao Wang , Mark Frank , Srirangaraj Setlur , Venu Govindaraju , Ifeoma Nwogu

The field of audio captioning has seen significant advancements in recent years, driven by the availability of large-scale audio datasets and advancements in deep learning techniques. In this technical report, we present our approach to…

Sound · Computer Science 2023-05-18 Marek Kadlčík , Adam Hájek , Jürgen Kieslich , Radosław Winiecki

The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI's Whisper model. As…

Sound · Computer Science 2025-02-03 Falguni Sharma , Priyanka Gupta

YouTube is the leading social media platform for sharing videos. As a result, it is plagued with misleading content that includes staged videos presented as real footages from an incident, videos with misrepresented context and videos where…

Computation and Language · Computer Science 2019-01-28 Priyank Palod , Ayush Patwari , Sudhanshu Bahety , Saurabh Bagchi , Pawan Goyal

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

Video grounding aims to localize a spatio-temporal section in a video corresponding to an input text query. This paper addresses a critical limitation in current video grounding methodologies by introducing an Open-Vocabulary…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Syed Talal Wasim , Muzammal Naseer , Salman Khan , Ming-Hsuan Yang , Fahad Shahbaz Khan

Action recognition is so far mainly focusing on the problem of classification of hand selected preclipped actions and reaching impressive results in this field. But with the performance even ceiling on current datasets, it also appears that…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Hilde Kuehne , Ahsan Iqbal , Alexander Richard , Juergen Gall