English
Related papers

Related papers: Naver at ActivityNet Challenge 2019 -- Task B Acti…

200 papers

Active Noise Control (ANC) systems are challenged by nonlinear distortions, which degrade the performance of traditional adaptive filters. While deep learning-based ANC algorithms have emerged to address nonlinearity, existing approaches…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-05 Lu Bai , Mengtong Li , Siyuan Lian , Kai Chen , Jing Lu

This paper presents the method that underlies our submission to the untrimmed video classification task of ActivityNet Challenge 2016. We follow the basic pipeline of temporal segment networks and further raise the performance via a number…

Computer Vision and Pattern Recognition · Computer Science 2016-08-03 Yuanjun Xiong , Limin Wang , Zhe Wang , Bowen Zhang , Hang Song , Wei Li , Dahua Lin , Yu Qiao , Luc Van Gool , Xiaoou Tang

In this paper, a novel Convolutional Neural Network architecture has been developed for speaker verification in order to simultaneously capture and discard speaker and non-speaker information, respectively. In training phase, the network is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-08-13 Hossein Salehghaffari

The strong relation between face and voice can aid active speaker detection systems when faces are visible, even in difficult settings, when the face of a speaker is not clear or when there are several people in the same scene. By being…

Machine Learning · Computer Science 2021-09-07 Hugo Carneiro , Cornelius Weber , Stefan Wermter

Deep learning approaches are still not very common in the speaker verification field. We investigate the possibility of using deep residual convolutional neural network with spectrograms as an input features in the text-dependent speaker…

Sound · Computer Science 2017-05-31 Egor Malykh , Sergey Novoselov , Oleg Kudashev

The first voice timbre attribute detection challenge is featured in a special session at NCMMSC 2025. It focuses on the explainability of voice timbre and compares the intensity of two speech utterances in a specified timbre descriptor…

Sound · Computer Science 2025-09-09 Liping Chen , Jinghao He , Zhengyan Sheng , Kong Aik Lee , Zhen-Hua Ling

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the…

Speaker recognition on household devices, such as smart speakers, features several challenges: (i) robustness across a vast number of heterogeneous domains (households), (ii) short utterances, (iii) possibly absent speaker labels of the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-06 Alexey Sholokhov , Xuechen Liu , Md Sahidullah , Tomi Kinnunen

Recently, x-vector has been a successful and popular approach for speaker verification, which employs a time delay neural network (TDNN) and statistics pooling to extract speaker characterizing embedding from variable-length utterances.…

Sound · Computer Science 2022-01-02 Wentao Zhu , Tianlong Kong , Shun Lu , Jixiang Li , Dawei Zhang , Feng Deng , Xiaorui Wang , Sen Yang , Ji Liu

This paper introduces our approaches for the Mask and Breathing Sub-Challenge in the Interspeech COMPARE Challenge 2020. For the mask detection task, we train deep convolutional neural networks with filter-bank energies, gender-aware…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Haiwei Wu , Lin Zhang , Lin Yang , Xuyang Wang , Junjie Wang , Dong Zhang , Ming Li

This notebook paper presents an overview and comparative analysis of our system designed for activity detection in extended videos (ActEV-PC) in ActivityNet Challenge 2019. Specifically, we exploit person/vehicle detections in spatial level…

Computer Vision and Pattern Recognition · Computer Science 2019-06-21 Fuchen Long , Qi Cai , Zhaofan Qiu , Zhijian Hou , Yingwei Pan , Ting Yao , Chong-Wah Ngo

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

Sound · Computer Science 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Xiulong Liu , Sudipta Paul , Moitreya Chatterjee , Anoop Cherian

Most existing research on visual question answering (VQA) is limited to information explicitly present in an image or a video. In this paper, we take visual understanding to a higher level where systems are challenged to answer questions…

Computer Vision and Pattern Recognition · Computer Science 2021-04-14 Shailaja Keyur Sampat , Akshay Kumar , Yezhou Yang , Chitta Baral

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their adaptability in open-ended analytical scenarios. The recent…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yuxuan Yan , Shiqi Jiang , Ting Cao , Yifan Yang , Qianqian Yang , Yuanchao Shu , Yuqing Yang , Lili Qiu

Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Sungnyun Kim , Sungwoo Cho , Sangmin Bae , Kangwook Jang , Se-Young Yun

The Detection and Classification of Acoustic Scenes and Events (DCASE) 2019 challenge focuses on audio tagging, sound event detection and spatial localisation. DCASE 2019 consists of five tasks: 1) acoustic scene classification, 2) audio…

Sound · Computer Science 2019-06-11 Qiuqiang Kong , Yin Cao , Turab Iqbal , Yong Xu , Wenwu Wang , Mark D. Plumbley

Audio classification is considered as a challenging problem in pattern recognition. Recently, many algorithms have been proposed using deep neural networks. In this paper, we introduce a new attention-based neural network architecture…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-18 Haoye Lu , Haolong Zhang , Amit Nayak

Research in speaker recognition has recently seen significant progress due to the application of neural network models and the availability of new large-scale datasets. There has been a plethora of work in search for more powerful…

Sound · Computer Science 2020-02-04 Joon Son Chung , Jaesung Huh , Seongkyu Mun

Convolutional neural networks (CNNs), such as the time-delay neural network (TDNN), have shown their remarkable capability in learning speaker embedding. However, they meanwhile bring a huge computational cost in storage size, processing,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-22 Rui Wang , Zhihua Wei , Haoran Duan , Shouling Ji , Yang Long , Zhen Hong
‹ Prev 1 8 9 10 Next ›