中文
相关论文

相关论文: Kaggle Competition: Cantonese Audio-Visual Speech …

200 篇论文

This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an…

声音 · 计算机科学 2025-05-30 Yuhang Dai , He Wang , Xingchen Li , Zihan Zhang , Shuiyuan Wang , Lei Xie , Xin Xu , Hongxiao Guo , Shaoji Zhang , Hui Bu , Wei Chen

The task of visual grounding requires locating the most relevant region or object in an image, given a natural language query. So far, progress on this task was mostly measured on curated datasets, which are not always representative of…

计算机视觉与模式识别 · 计算机科学 2020-09-21 Thierry Deruyttere , Simon Vandenhende , Dusan Grujicic , Yu Liu , Luc Van Gool , Matthew Blaschko , Tinne Tuytelaars , Marie-Francine Moens

This paper presents our system submission for the In-Car Multi-Channel Automatic Speech Recognition (ICMC-ASR) Challenge, which focuses on speaker diarization and speech recognition in complex multi-speaker scenarios. To address these…

声音 · 计算机科学 2024-05-10 Jingguang Tian , Shuaishuai Ye , Shunfei Chen , Yang Xiang , Zhaohui Yin , Xinhui Hu , Xinkang Xu

In this paper, we introduce Context-Aware Priority Sampling (CAPS), a novel method designed to enhance data efficiency in learning-based autonomous driving systems. CAPS addresses the challenge of imbalanced datasets in imitation learning…

In this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research. We define the task at…

In this paper we propose a novel virtual simulation-pilot engine for speeding up air traffic controller (ATCo) training by integrating different state-of-the-art artificial intelligence (AI) based tools. The virtual simulation-pilot engine…

音频与语音处理 · 电气工程与系统科学 2023-04-18 Juan Zuluaga-Gomez , Amrutha Prasad , Iuliia Nigmatulina , Petr Motlicek , Matthias Kleinert

The recently proposed audio-visual scene-aware dialog task paves the way to a more data-driven way of learning virtual assistants, smart speakers and car navigation systems. However, very little is known to date about how to effectively…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Idan Schwartz , Alexander Schwing , Tamir Hazan

Computer-generated imagery of car models has become an indispensable part of car manufacturers' advertisement concepts. They are for instance used in car configurators to offer customers the possibility to configure their car online…

机器学习 · 计算机科学 2021-10-19 Patrick Hemmer , Niklas Kühl , Jakob Schöffer

Autonomous driving is a multi-task problem requiring a deep understanding of the visual environment. End-to-end autonomous systems have attracted increasing interest as a method of learning to drive without exhaustively programming…

计算机视觉与模式识别 · 计算机科学 2019-09-12 Alexander Makrigiorgos , Ali Shafti , Alex Harston , Julien Gerard , A. Aldo Faisal

Expected to provide higher transportation efficiency and security, autonomous driving has attracted substantial attentions from both industry and academia. Meanwhile, the emergence of edge intelligence has further introduced significant…

信号处理 · 电气工程与系统科学 2024-09-25 Yunqi Feng , Hesheng Shen , Zhendong Shan , Qianqian Yang , Xiufang Shi

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

声音 · 计算机科学 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support only a small subset. Prior research indicates LLMs can learn new languages for certain tasks without supervised data. We extend this…

计算与语言 · 计算机科学 2026-01-29 Zhaolin Li , Jan Niehues

This paper addresses the problem of building a speech recognition system attuned to the control of unmanned aerial vehicles (UAVs). Even though UAVs are becoming widespread, the task of creating voice interfaces for them is largely…

声音 · 计算机科学 2019-07-03 Dan Oneata , Horia Cucu

A context-aware recommender system (CARS) applies sensing and analysis of user context to provide personalized services. The contextual information can be driven from sensors in order to improve the accuracy of the recommendations. Yet,…

机器学习 · 计算机科学 2022-08-10 Amit Livne , Eliad Shem Tov , Adir Solomon , Achiya Elyasaf , Bracha Shapira , Lior Rokach

Air traffic management and specifically air-traffic control (ATC) rely mostly on voice communications between Air Traffic Controllers (ATCos) and pilots. In most cases, these voice communications follow a well-defined grammar that could be…

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

声音 · 计算机科学 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

Despite recent advances in end-to-end speech recognition methods, the output tends to be biased to the training data's vocabulary, resulting in inaccurate recognition of proper nouns and other unknown terms. To address this issue, we…

计算与语言 · 计算机科学 2025-06-03 Yu Nakagome , Michael Hentschel

Automatic speech recognition (ASR) plays a pivotal role in our daily lives, offering utility not only for interacting with machines but also for facilitating communication for individuals with partial or profound hearing impairments. The…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Billel Essaid , Hamza Kheddar , Noureddine Batel , Muhammad E. H. Chowdhury , Abderrahmane Lakas

Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora…

音频与语音处理 · 电气工程与系统科学 2026-03-10 Pol Buitrago , Pol Gàlvez , Oriol Pareras , Javier Hernando

Single-word Automatic Speech Recognition (ASR) is a challenging task due to the lack of linguistic context and sensitivity to noise, pronunciation variation, and channel artifacts, especially in low-resource, communication-critical domains…

声音 · 计算机科学 2026-01-30 Manali Sharma , Riya Naik , Buvaneshwari G