English
Related papers

Related papers: Solid State Bus-Comp: A Large-Scale and Diverse Da…

200 papers

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-14 Yuke Lin , Xiaoyi Qin , Guoqing Zhao , Ming Cheng , Ning Jiang , Haiyang Wu , Ming Li

While design of high performance lenses and image sensors has long been the focus of camera development, the size, weight and power of image data processing components is currently the primary barrier to radical improvements in camera…

Image and Video Processing · Electrical Eng. & Systems 2019-08-30 Xuefei Yan , David J. Brady , Jianqiang Wang , Chao Huang , Zian Li , Songsong Yan , Di Liu , Zhan Ma

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

For many Automatic Speech Recognition (ASR) tasks audio features as spectrograms show better results than Mel-frequency Cepstral Coefficients (MFCC), but in practice they are hard to use due to a complex dimensionality of a feature space.…

Sound · Computer Science 2024-10-07 Olga Iakovenko , Ivan Bondarenko

Video semantic segmentation(VSS) has been widely employed in lots of fields, such as simultaneous localization and mapping, autonomous driving and surveillance. Its core challenge is how to leverage temporal information to achieve better…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Zhigang Cen , Ningyan Guo , Wenjing Xu , Zhiyong Feng , Danlan Huang

Integrating audio and visual data for training multimodal foundational models remains a challenge. The Audio-Video Vector Alignment (AVVA) framework addresses this by considering AV scene alignment beyond mere temporal synchronization, and…

Multimedia · Computer Science 2025-11-12 Ali Vosoughi , Dimitra Emmanouilidou , Hannes Gamper

In this paper, we develop a new framework for sensing and recovering structured signals. In contrast to compressive sensing (CS) systems that employ linear measurements, sparse representations, and computationally complex convex/greedy…

Machine Learning · Computer Science 2016-09-01 Ali Mousavi , Ankit B. Patel , Richard G. Baraniuk

With the proliferation of IoT devices, the distributed control systems are now capturing and processing more sensors at higher frequency than ever before. These new data, due to their volume and novelty, cannot be effectively consumed…

Signal Processing · Electrical Eng. & Systems 2022-01-25 Chao Zhang , Sthitie Bom

Power quality monitoring has become a vital need in modern power systems owing to the need for agile operation and troubleshooting scheme. On the other hand, the nature of load in modern power system is changing in many ways. Digital loads,…

Systems and Control · Electrical Eng. & Systems 2023-08-15 Saeed Nasiri

Event cameras are gaining traction in traffic monitoring applications due to their low latency, high temporal resolution, and energy efficiency, which makes them well-suited for real-time object detection at traffic intersections. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Kaiyuan Tan , Pavan Kumar B N , Bharatesh Chakravarthi

Deep convolution neural network has attracted many attentions in large-scale visual classification task, and achieves significant performance improvement compared to traditional visual analysis methods. In this paper, we explore many kinds…

Computer Vision and Pattern Recognition · Computer Science 2020-07-06 Feifei Huang , Jie Li , Xuelin Zhu

This work presents a sum-of-squares (SOS) based framework to perform data-driven stabilization and robust control tasks on discrete-time linear systems where the full-state observations are corrupted by L-infinity bounded input,…

Optimization and Control · Mathematics 2023-03-31 Jared Miller , Tianyu Dai , Mario Sznaier

How to improve the ability of scene representation is a key issue in vision-oriented decision-making applications, and current approaches usually learn task-relevant state representations within visual reinforcement learning to address this…

Artificial Intelligence · Computer Science 2024-10-24 Dayang Liang , Jinyang Lai , Yunlong Liu

Getting the best performance from the ever-increasing number of hardware platforms has been a recurring challenge for data processing systems. In recent years, the advent of data science with its increasingly numerous and complex types of…

Most existing data-driven power system short-term voltage stability assessment (STVSA) approaches presume class-balanced input data. However, in practical applications, the occurrence of short-term voltage instability following a…

Systems and Control · Electrical Eng. & Systems 2023-11-09 Yang Li , Jiting Cao , Yan Xu , Lipeng Zhu , Zhao Yang Dong

In this paper, a siamese DNN model is proposed to learn the characteristics of the audio dynamic range compressor (DRC). This facilitates an intelligent control system that uses audio examples to configure the DRC, a widely used non-linear…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-06 Di Sheng , György Fazekas

We introduce a cutting-edge video compression framework tailored for the age of ubiquitous video data, uniquely designed to serve machine learning applications. Unlike traditional compression methods that prioritize human visual perception,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Huan Cui , Qing Li , Hanling Wang , Yong jiang

Motivated by the ever-increasing demands for limited communication bandwidth and low-power consumption, we propose a new methodology, named joint Variational Autoencoders with Bernoulli mixture models (VAB), for performing clustering in the…

Image and Video Processing · Electrical Eng. & Systems 2020-06-11 Suya Wu , Enmao Diao , Jie Ding , Vahid Tarokh

Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are scarce. Existing workarounds that use synthetic data often…

Sound · Computer Science 2026-01-15 Pierfrancesco Melucci , Paolo Merialdo , Taketo Akama

Versatile Video Coding (VVC) is the next generation video coding standard expected by the end of 2020. Compared to its predecessor, VVC introduces new coding tools to make compression more efficient at the expense of higher computational…

Image and Video Processing · Electrical Eng. & Systems 2020-02-19 I. Farhat , W. Hamidouche , A Grill , D. Ménard , O. Deforges
‹ Prev 1 8 9 10 Next ›