中文
相关论文

相关论文: Action Sensitivity Learning for the Ego4D Episodic…

200 篇论文

We developed an American Sign Language (ASL) learning platform in a Virtual Reality (VR) environment to facilitate immersive interaction and real-time feedback for ASL learners. We describe the first game to use an interactive teaching…

This paper investigates the problem of understanding dynamic 3D scenes from egocentric observations, a key challenge in robotics and embodied AI. Unlike prior studies that explored this as long-form video understanding and utilized…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Yue Fan , Xiaojian Ma , Rongpeng Su , Jun Guo , Rujie Wu , Xi Chen , Qing Li

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with noise-invariant visual cues and improve the system's…

声音 · 计算机科学 2024-04-09 He Wang , Pengcheng Guo , Pan Zhou , Lei Xie

This report presents the technical details of our submission on the EGO4D Audio-Visual (AV) Automatic Speech Recognition Challenge 2023 from the OxfordVGG team. We present WhisperX, a system for efficient speech transcription of long-form…

声音 · 计算机科学 2023-07-19 Jaesung Huh , Max Bain , Andrew Zisserman

This paper presents a transfer learning method in speech emotion recognition based on a Time-Delay Neural Network (TDNN) architecture. A major challenge in the current speech-based emotion detection research is data scarcity. The proposed…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Sitong Zhou , Homayoon Beigi

In this paper, we propose a method for incremental learning of two distinct tasks over time: acoustic scene classification (ASC) and audio tagging (AT). We use a simple convolutional neural network (CNN) model as an incremental learner to…

音频与语音处理 · 电气工程与系统科学 2023-08-25 Manjunath Mulimani , Annamaria Mesaros

This paper presents a real-time American Sign Language (ASL) recognition system utilizing a hybrid deep learning architecture combining 3D Convolutional Neural Networks (3D CNN) with Long Short-Term Memory (LSTM) networks. The system…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Dawnena Key

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D dataset…

In this report, we describe the technical details of our submission to the EPIC-SOUNDS Audio-Based Interaction Recognition Challenge 2023, by Team "AcieLee" (username: Yuqi\_Li). The task is to classify the audio caused by interactions…

声音 · 计算机科学 2023-06-16 Yuqi Li , Yizhi Luo , Xiaoshuai Hao , Chuanguang Yang , Zhulin An , Dantong Song , Wei Yi

As image-based deep learning becomes pervasive on every device, from cell phones to smart watches, there is a growing need to develop methods that continually learn from data while minimizing memory footprint and power consumption. While…

机器学习 · 计算机科学 2021-03-24 Dongsub Shim , Zheda Mai , Jihwan Jeong , Scott Sanner , Hyunwoo Kim , Jongseong Jang

In this paper, we describe our approach for the OMG- Emotion Challenge 2018. The goal is to produce utterance-level valence and arousal estimations for videos of approximately 1 minute length. We tackle this problem by first extracting…

计算机视觉与模式识别 · 计算机科学 2018-05-23 Tianlin Liu , Arvid Kappas

This paper describes our NPU-ASLP system for the Audio-Visual Diarization and Recognition (AVDR) task in the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. Specifically, the weighted prediction error (WPE) and guided…

音频与语音处理 · 电气工程与系统科学 2023-03-14 Pengcheng Guo , He Wang , Bingshen Mu , Ao Zhang , Peikun Chen

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Changan Chen , Puyuan Peng , Ami Baid , Zihui Xue , Wei-Ning Hsu , David Harwath , Kristen Grauman

As emotions play a central role in human communication, automatic emotion recognition has attracted increasing attention in the last two decades. While multimodal systems enjoy high performances on lab-controlled data, they are still far…

机器学习 · 计算机科学 2024-03-20 Denis Dresvyanskiy , Maxim Markitantov , Jiawei Yu , Peitong Li , Heysem Kaya , Alexey Karpov

Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants and agents. However, most current ASR systems still follow a single-pass paradigm, which…

人工智能 · 计算机科学 2026-05-29 Zixuan Jiang , Yanqiao Zhu , Peng Wang , Qinyuan Chen , Xinjian Zhao , Xipeng Qiu , Wupeng Wang , Zhifu Gao , Xiangang Li , Kai Yu , Xie Chen

The domain of speech emotion recognition (SER) has persistently been a frontier within the landscape of machine learning. It is an active field that has been revolutionized in the last few decades and whose implementations are remarkable in…

音频与语音处理 · 电气工程与系统科学 2024-07-18 Marc Casals-Salvador , Federico Costa , Miquel India , Javier Hernando

This paper presents a light-weight and accurate deep neural model for audiovisual emotion recognition. To design this model, the authors followed a philosophy of simplicity, drastically limiting the number of parameters to learn from the…

人工智能 · 计算机科学 2018-08-09 Valentin Vielzeuf , Corentin Kervadec , Stéphane Pateux , Alexis Lechervy , Frédéric Jurie

Continual learning aims to provide intelligent agents capable of learning multiple tasks sequentially with neural networks. One of its main challenging, catastrophic forgetting, is caused by the neural networks non-optimal ability to learn…

机器学习 · 计算机科学 2021-01-29 Ghada Sokar , Decebal Constantin Mocanu , Mykola Pechenizkiy

This paper introduces an active learning (AL) framework for anomalous sound detection (ASD) in machine condition monitoring system. Typically, ASD models are trained solely on normal samples due to the scarcity of anomalous data, leading to…

声音 · 计算机科学 2024-08-13 Tuan Vu Ho , Kota Dohi , Yohei Kawaguchi

Speech Emotion Recognition (SER) plays a key role in advancing human-computer interaction. Attention mechanisms have become the dominant approach for modeling emotional speech due to their ability to capture long-range dependencies and…

音频与语音处理 · 电气工程与系统科学 2026-03-17 Marc Casals-Salvador , Federico Costa , Rodolfo Zevallos , Javier Hernando