English
Related papers

Related papers: SoccerNet-Echoes: A Soccer Game Audio Commentary D…

200 papers

In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR). We explore novel VC-ASR approaches to leverage video and text…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-10 Shahram Ghorbani , Yashesh Gaur , Yu Shi , Jinyu Li

Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational…

Computation and Language · Computer Science 2024-09-19 Gaurav Maheshwari , Dmitry Ivanov , Théo Johannet , Kevin El Haddad

The SoccerNet 2025 Challenges mark the fifth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in football video understanding. This year's challenges span four vision-based tasks: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Silvio Giancola , Anthony Cioppa , Marc Gutiérrez-Pérez , Jan Held , Carlos Hinojosa , Victor Joos , Arnaud Leduc , Floriane Magera , Karen Sanchez , Vladimir Somers , Artur Xarles , Antonio Agudo , Alexandre Alahi , Olivier Barnich , Albert Clapés , Christophe De Vleeschouwer , Sergio Escalera , Bernard Ghanem , Thomas B. Moeslund , Marc Van Droogenbroeck , Tomoki Abe , Saad Alotaibi , Faisal Altawijri , Steven Araujo , Xiang Bai , Xiaoyang Bi , Jiawang Cao , Vanyi Chao , Kamil Czarnogórski , Fabian Deuser , Mingyang Du , Tianrui Feng , Patrick Frenzel , Mirco Fuchs , Jorge García , Konrad Habel , Takaya Hashiguchi , Sadao Hirose , Xinting Hu , Yewon Hwang , Ririko Inoue , Riku Itsuji , Kazuto Iwai , Hongwei Ji , Yangguang Ji , Licheng Jiao , Yuto Kageyama , Yuta Kamikawa , Yuuki Kanasugi , Hyungjung Kim , Jinwook Kim , Takuya Kurihara , Bozheng Li , Lingling Li , Xian Li , Youxing Lian , Dingkang Liang , Hongkai Lin , Jiadong Lin , Jian Liu , Liang Liu , Shuaikun Liu , Zhaohong Liu , Yi Lu , Federico Méndez , Huadong Ma , Wenping Ma , Jacek Maksymiuk , Henry Mantilla , Ismail Mathkour , Daniel Matthes , Ayaha Motomochi , Amrulloh Robbani Muhammad , Haruto Nakayama , Joohyung Oh , Yin May Oo , Marcelo Ortega , Norbert Oswald , Rintaro Otsubo , Fabian Perez , Mengshi Qi , Cristian Rey , Abel Reyes-Angulo , Oliver Rose , Hoover Rueda-Chacón , Hideo Saito , Jose Sarmiento , Kanta Sawafuji , Atom Scott , Xi Shen , Pragyan Shrestha , Jae-Young Sim , Long Sun , Yuyang Sun , Tomohiro Suzuki , Licheng Tang , Masato Tonouchi , Ikuma Uchida , Henry O. Velesaca , Tiancheng Wang , Rio Watanabe , Jay Wu , Yongliang Wu , Shunzo Yamagishi , Di Yang , Xu Yang , Yuxin Yang , Hao Ye , Xinyu Ye , Calvin Yeung , Xuanlong Yu , Chao Zhang , Dingyuan Zhang , Kexing Zhang , Zhe Zhao , Xin Zhou , Wenbo Zhu , Julian Ziegler

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to…

In-game win probability models, which provide a sports team's likelihood of winning at each point in a game based on historical observations, are becoming increasingly popular. In baseball, basketball and American football, they have become…

Machine Learning · Computer Science 2021-08-16 Pieter Robberechts , Jan Van Haaren , Jesse Davis

The proposed system aims at the retrieval of the summarized information from the documents collected from web based search engine as per the user query related to cricket and hockey domain. The system is designed in a manner that it takes…

Information Retrieval · Computer Science 2010-04-27 S. Saraswathi , Narasimha Sravan. , Sai Vamsi Krishna. B. , Suresh Reddy. S

Automatic speech recognition (ASR) systems can suffer from poor recall for various reasons, such as noisy audio, lack of sufficient training data, etc. Previous work has shown that recall can be improved by retrieving rewrite candidates…

Automatic speech recognition (ASR) system is becoming a ubiquitous technology. Although its accuracy is closing the gap with that of human level under certain settings, one area that can further improve is to incorporate user-specific…

Computation and Language · Computer Science 2020-05-05 Young Mo Kang , Yingbo Zhou

We propose a semi-supervised learning method for building end-to-end rich transcription-style automatic speech recognition (RT-ASR) systems from small-scale rich transcription-style and large-scale common transcription-style datasets. In…

Computation and Language · Computer Science 2021-07-13 Tomohiro Tanaka , Ryo Masumura , Mana Ihori , Akihiko Takashima , Shota Orihashi , Naoki Makishima

Automatic Speech Recognition (ASR) for air traffic control is generally trained by pooling Air Traffic Controller (ATCO) and pilot data into one set. This is motivated by the fact that pilot's voice communications are more scarce than…

Computation and Language · Computer Science 2022-12-15 Amrutha Prasad , Juan Zuluaga-Gomez , Petr Motlicek , Saeed Sarfjoo , Iuliia Nigmatulina , Oliver Ohneiser , Hartmut Helmke

Code-switching automatic speech recognition (CS-ASR) presents unique challenges due to language confusion introduced by spontaneous intra-sentence switching and accent bias that blurs the phonetic boundaries. Although the constituent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-18 Hexin Liu , Haoyang Zhang , Qiquan Zhang , Xiangyu Zhang , Dongyuan Shi , Eng Siong Chng , Haizhou Li

In the FAME! project, we aim to develop an automatic speech recognition (ASR) system for Frisian-Dutch code-switching (CS) speech extracted from the archives of a local broadcaster with the ultimate goal of building a spoken document…

Computation and Language · Computer Science 2018-10-24 Emre Yılmaz , Mitchell McLaren , Henk van den Heuvel , David A. van Leeuwen

In this paper, I introduce RisingBALLER, the first publicly available approach that leverages a transformer model trained on football match data to learn match-specific player representations. Drawing inspiration from advances in language…

Machine Learning · Computer Science 2024-10-03 Akedjou Achraff Adjileye

Automatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of…

Computation and Language · Computer Science 2025-04-14 Xabier de Zuazo , Eva Navas , Ibon Saratxaga , Inma Hernáez Rioja

Air traffic management and specifically air-traffic control (ATC) rely mostly on voice communications between Air Traffic Controllers (ATCos) and pilots. In most cases, these voice communications follow a well-defined grammar that could be…

Computation and Language · Computer Science 2021-08-30 Juan Zuluaga-Gomez , Iuliia Nigmatulina , Amrutha Prasad , Petr Motlicek , Karel Veselý , Martin Kocour , Igor Szöke

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Vladimir Iashin , Esa Rahtu

Artificial intelligence has revolutionized the way we analyze sports videos, whether to understand the actions of games in long untrimmed videos or to anticipate the player's motion in future frames. Despite these efforts, little attention…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Mohamad Dalal , Artur Xarles , Anthony Cioppa , Silvio Giancola , Marc Van Droogenbroeck , Bernard Ghanem , Albert Clapés , Sergio Escalera , Thomas B. Moeslund

Segmenting audio into homogeneous sections such as music and speech helps us understand the content of audio. It is useful as a pre-processing step to index, store, and modify audio recordings, radio broadcasts and TV programmes. Deep…

With the huge technological advances introduced by deep learning in audio & speech processing, many novel synthetic speech techniques achieved incredible realistic results. As these methods generate realistic fake human voices, they can be…

Contextual biasing improves automatic speech recognition (ASR) by integrating external knowledge, such as user-specific phrases or entities, during decoding. In this work, we use an attention-based biasing decoder to produce scores for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-29 Wanting Huang , Weiran Wang