English
Related papers

Related papers: Kernel-based Sensor Fusion with Application to Aud…

200 papers

As a fundamental problem in ubiquitous computing and machine learning, sensor-based human activity recognition (HAR) has drawn extensive attention and made great progress in recent years. HAR aims to recognize human activities based on the…

Signal Processing · Electrical Eng. & Systems 2022-03-01 Yimu Wang , Kun Yu , Yan Wang , Hui Xue

It is challenging to recognize facial action unit (AU) from spontaneous facial displays, especially when they are accompanied by speech. The major reason is that the information is extracted from a single source, i.e., the visual channel,…

Computer Vision and Pattern Recognition · Computer Science 2017-07-03 Zibo Meng , Shizhong Han , Ping Liu , Yan Tong

Machine learning and deep learning have been used extensively to classify physical surfaces through images and time-series contact data. However, these methods rely on human expertise and entail the time-consuming processes of data and…

Machine Learning · Computer Science 2023-08-10 Behnam Khojasteh , Friedrich Solowjow , Sebastian Trimpe , Katherine J. Kuchenbecker

Many adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Yikai Wang , Xinghao Chen , Lele Cao , Wenbing Huang , Fuchun Sun , Yunhe Wang

In the field of detection and ranging, multiple complementary sensing modalities may be used to enrich the information obtained from a dynamic scene. One application of this sensor fusion is in public security and surveillance, whose…

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

Voice Activity Detection (VAD) is an important pre-processing step in a wide variety of speech processing systems. VAD should in a practical application be able to detect speech in both noisy and noise-free environments, while not…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-06 Claus Meyer Larsen , Peter Koch , Zheng-Hua Tan

Action recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recognition accuracy by: 1) deriving more precise features for…

Computer Vision and Pattern Recognition · Computer Science 2017-11-21 Weiyao Lin , Yang Mi , Jianxin Wu , Ke Lu , Hongkai Xiong

Video fusion is a process that combines visual data from different sensors to obtain a single composite video preserving the information of the sources. The availability of a system, enhancing human ability to perceive the observed…

Multimedia · Computer Science 2010-04-27 Anjali Malviya , S. G. Bhirud

Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. The approach of AVR…

Computer Vision and Pattern Recognition · Computer Science 2017-11-01 Amirsina Torfi , Seyed Mehdi Iranmanesh , Nasser M. Nasrabadi , Jeremy Dawson

In this paper, we describe a conceptual design methodology to design distributed neural network architectures that can perform efficient inference within sensor networks with communication bandwidth constraints. The different sensor…

Machine Learning · Computer Science 2022-10-17 Thomas Strypsteen , Alexander Bertrand

This article presents a method for improving a keyword spotter (KWS) algorithm in noisy environments. Although beamforming (BF) and adaptive noise cancellation (ANC) techniques are robust in some conditions, they may degrade the performance…

Video action recognition has been partially addressed by the CNNs stacking of fixed-size 3D kernels. However, these methods may under-perform for only capturing rigid spatial-temporal patterns in single-scale spaces, while neglecting the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-20 Yuan Tian , Guangtao Zhai , Zhiyong Gao

This paper presents a technique that combines the occurrence of certain events, as observed by different sensors, in order to detect and classify objects. This technique explores the extent of dependence between features being observed by…

Signal Processing · Electrical Eng. & Systems 2018-10-02 Siddharth Roheda , Hamid Krim , Zhi-Quan Luo , Tianfu Wu

Object detection is an essential task for autonomous robots operating in dynamic and changing environments. A robot should be able to detect objects in the presence of sensor noise that can be induced by changing lighting conditions for…

Robotics · Computer Science 2019-11-20 Oier Mees , Andreas Eitel , Wolfram Burgard

In this paper, we propose regular vine copula based fusion of multiple deep neural network classifiers for the problem of multi-sensor based human activity recognition. We take the cross-modal dependence into account by employing regular…

Signal Processing · Electrical Eng. & Systems 2019-11-22 Shan Zhang , Baocheng Geng , Pramod K. Varshney , Muralidhar Rangaswamy

This paper presents a technique which exploits the occurrence of certain events as observed by different sensors, to detect and classify objects. This technique explores the extent of dependence between features being observed by the…

Signal Processing · Electrical Eng. & Systems 2021-03-08 Siddharth Roheda , Hamid Krim , Zhi-Quan Luo , Tianfu Wu

We propose a method for automated synchronization of vehicle sensors useful for the study of multi-modal driver behavior and for the design of advanced driver assistance systems. Multi-sensor decision fusion relies on synchronized data…

Robotics · Computer Science 2016-03-02 Lex Fridman , Daniel E Brown , William Angell , Irman Abdić , Bryan Reimer , Hae Young Noh

We consider the design of an image representation that embeds and aggregates a set of local descriptors into a single vector. Popular representations of this kind include the bag-of-visual-words, the Fisher vector and the VLAD. When two…

Computer Vision and Pattern Recognition · Computer Science 2016-11-28 Naila Murray , Hervé Jégou , Florent Perronnin , Andrew Zisserman

A successful class of image denoising methods is based on Bayesian approaches working in wavelet representations. However, analytical estimates can be obtained only for particular combinations of analytical models of signal and noise, thus…

Computer Vision and Pattern Recognition · Computer Science 2016-02-02 Valero Laparra , Juan Gutiérrez , Gustavo Camps-Valls , Jesús Malo
‹ Prev 1 8 9 10 Next ›