English
Related papers

Related papers: DCAR: A Discriminative and Compact Audio Represent…

200 papers

Optimizing video inference efficiency has become increasingly important with the growing demand for video analysis in various fields. Some existing methods achieve high efficiency by explicit discard of spatial or temporal information,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Rui Deng , Qian Wu , Yuke Li , Haoran Fu

Current autoencoder-based disentangled representation learning methods achieve disentanglement by penalizing the (aggregate) posterior to encourage statistical independence of the latent factors. This approach introduces a trade-off between…

Video anomaly detection aims to develop automated models capable of identifying abnormal events in surveillance videos. The benchmark setup for this task is extremely challenging due to: i) the limited size of the training sets, ii) weak…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Jash Dalvi , Ali Dabouei , Gunjan Dhanuka , Min Xu

This paper develops a semiparametric Bayesian instrumental variable analysis method for estimating the causal effect of an endogenous variable when dealing with unobserved confounders and measurement errors with partly interval-censored…

Methodology · Statistics 2025-01-28 Elvis Han Cui , Xuyang Lu , Jin Zhou , Hua Zhou , Gang Li

The goal of acoustic (or sound) events detection (AED or SED) is to predict the temporal position of target events in given audio segments. This task plays a significant role in safety monitoring, acoustic early warning and other scenarios.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-26 Wenhao Ding , Liang He

We present Dance2Music-GAN (D2M-GAN), a novel adversarial multi-modal framework that generates complex musical samples conditioned on dance videos. Our proposed framework takes dance video frames and human body motions as input, and learns…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Ye Zhu , Kyle Olszewski , Yu Wu , Panos Achlioptas , Menglei Chai , Yan Yan , Sergey Tulyakov

In speaker tracking research, integrating and complementing multi-modal data is a crucial strategy for improving the accuracy and robustness of tracking systems. However, tracking with incomplete modalities remains a challenging issue due…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Yidi Li , Yihan Li , Yixin Guo , Bin Ren , Zhenhuan Xu , Hao Guo , Hong Liu , Nicu Sebe

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-31 Nan Xu , Zhaolong Huang , Xiaonan Zhi

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Vladimir Iashin , Esa Rahtu

Vision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Bo Xu , Cheng Lu , Yandong Guo , Jacob Wang

Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by jointly leveraging auditory and visual information. However, existing methods often suffer from multi-source entanglement and audio-visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Jingqi Tian , Yiheng Du , Haoji Zhang , Yuji Wang , Isaac Ning Lee , Xulong Bai , Tianrui Zhu , Jingxuan Niu , Yansong Tang

The automatic speaker identification procedure is used to extract features that help to identify the components of the acoustic signal by discarding all the other stuff like background noise, emotion, hesitation, etc. The acoustic signal is…

Sound · Computer Science 2017-04-14 Soumen Kanrar

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

Sound · Computer Science 2021-02-11 Zeqian Li , Jacob Whitehill

Unsupervised anomalous sound detection (ASD) aims to identify anomalous sounds by learning the features of normal operational sounds and sensing their deviations. Recent approaches have focused on the self-supervised task utilizing the…

Sound · Computer Science 2023-10-11 Soonhyeon Choi , Jung-Woo Choi

Event detection improves when events are captured by two different modalities rather than just one. But to train detection systems on multiple modalities is challenging, in particular when there is abundance of unlabelled data but limited…

Sound · Computer Science 2022-11-18 Sumit Kumar , B. Anshuman , Linus Ruettimann , Richard H. R. Hahnloser , Vipul Arora

We propose discriminative neighborhood smoothing of generative anomaly scores for anomalous sound detection. While the discriminative approach is known to achieve better performance than generative approaches often, we have found that it…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-19 Takuya Fujimura , Keisuke Imoto , Tomoki Toda

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

We tackle the task of learning dynamic 3D semantic radiance fields given a single monocular video as input. Our learned semantic radiance field captures per-point semantics as well as color and geometric properties for a dynamic 3D scene,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Isaac Labe , Noam Issachar , Itai Lang , Sagie Benaim

The deep complex convolution recurrent network (DCCRN) achieves excellent speech enhancement performance by utilizing the audio spectrum's complex features. However, it has a large number of model parameters. We propose a smaller model,…

Sound · Computer Science 2024-08-09 Runduo Han , Weiming Xu , Zihan Zhang , Mingshuai Liu , Lei Xie

Reconstructing Dynamic 3D Gaussian Splatting (3DGS) from low-framerate RGB videos is challenging. This is because large inter-frame motions will increase the uncertainty of the solution space. For example, one pixel in the first frame might…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Junhao He , Jiaxu Wang , Jia Li , Mingyuan Sun , Qiang Zhang , Jiahang Cao , Ziyi Zhang , Yi Gu , Jingkai Sun , Renjing Xu
‹ Prev 1 4 5 6 7 8 10 Next ›