English
Related papers

Related papers: Unbiased Sliced Wasserstein Kernels for High-Quali…

200 papers

Realistic sound propagation is essential for immersion in a virtual scene, yet physically accurate wave-based simulations remain computationally prohibitive for real-time applications. Wave coding methods address this limitation by…

Sound · Computer Science 2026-02-09 Hugo Seuté , Pranai Vasudev , Etienne Richan , Louis-Xavier Buffoni

Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer from limitations, including unimodal supervision, disjoint…

Sound · Computer Science 2025-12-17 Ramesh Gundluru , Shubham Gupta , Sri Rama Murty K

Wasserstein distributionally robust optimization (\textsf{WDRO}) is a popular model to enhance the robustness of machine learning with ambiguous data. However, the complexity of \textsf{WDRO} can be prohibitive in practice since solving its…

Machine Learning · Computer Science 2023-05-10 Ruomin Huang , Jiawei Huang , Wenjie Liu , Hu Ding

It is encouraged to see that progress has been made to bridge videos and natural language. However, mainstream video captioning methods suffer from slow inference speed due to the sequential manner of autoregressive decoding, and prefer…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Bang Yang , Yuexian Zou , Fenglin Liu , Can Zhang

Recent developments have made it possible to overcome grid-based limitations of finite difference (FD) methods by adopting the kernel-based meshless framework using radial basis functions (RBFs). Such an approach provides a meshless…

Numerical Analysis · Mathematics 2019-01-07 Pankaj K Mishra , Gregory E Fasshauer , Mrinal K Sen , Leevan Ling

The advances in the Neural Radiance Fields (NeRF) research offer extensive applications in diverse domains, but protecting their copyrights has not yet been researched in depth. Recently, NeRF watermarking has been considered one of the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Youngdong Jang , Dong In Lee , MinHyuk Jang , Jong Wook Kim , Feng Yang , Sangpil Kim

Wasserstein gradient and Hamiltonian flows have emerged as essential tools for modeling complex dynamics in the natural sciences, with applications ranging from partial differential equations (PDEs) and optimal transport to quantum…

Numerical Analysis · Mathematics 2025-11-11 Jianyu Hu , Juan-Pablo Ortega , Daiying Yin

Radio frequency interference (RFI) mitigation and radar echo recovery are critically important for the proper functioning of ultra-wideband (UWB) radar systems using one-bit sampling techniques. We recently introduced a technique for…

Signal Processing · Electrical Eng. & Systems 2021-03-22 Tianyi Zhang , Jiaying Ren , Jian Li , Lam H. Nguyen , Petre Stoica

Image denoising is a fundamental challenge in computer vision, with applications in photography and medical imaging. While deep learning-based methods have shown remarkable success, their reliance on specific noise distributions limits…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Dongjin Kim , Jaekyun Ko , Muhammad Kashif Ali , Tae Hyun Kim

Distributionally robust optimization (DRO)-based robust adaptive beamforming (RAB) enables enhanced robustness against model uncertainties, such as steering vector mismatches and interference-plus-noise covariance matrix estimation errors.…

Signal Processing · Electrical Eng. & Systems 2025-06-03 Kiarash Hassas Irani , Sergiy A. Vorobyov , Yongwei Huang

One of the key challenges in learning joint embeddings of multiple modalities, e.g. of images and text, is to ensure coherent cross-modal semantics that generalize across datasets. We propose to address this through joint Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Shweta Mahajan , Teresa Botschen , Iryna Gurevych , Stefan Roth

We address prevailing challenges of the brain-powered research, departing from the observation that the literature hardly recover accurate spatial information and require subject-specific models. To address these challenges, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Weihao Xia , Raoul de Charette , Cengiz Öztireli , Jing-Hao Xue

Automatic video captioning aims for a holistic visual scene understanding. It requires a mechanism for capturing temporal context in video frames and the ability to comprehend the actions and associations of objects in a given timeframe.…

Computer Vision and Pattern Recognition · Computer Science 2022-12-22 Daniel Lukas Rothenpieler , Shahin Amiriparian

Speaker identification is the process of determining which registered speaker provides a given utterance. Speaker identification required to make a claim on the identity of speaker from the Ns trained speaker in its user database. In this…

Multimedia · Computer Science 2010-04-27 Ibrahim A. Albidewi , Yap Teck Ann

Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity, limiting their ability to model the diverse temporal and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Bing Han , Chushu Zhou , Yifan Yang , Wei Wang , Chenda Li , Wangyou Zhang , Yanmin Qian

Some response surface functions in complex engineering systems are usually highly nonlinear, unformed, and expensive-to-evaluate. To tackle this challenge, Bayesian optimization, which conducts sequential design via a posterior distribution…

Machine Learning · Statistics 2021-09-23 Areej AlBahar , Inyoung Kim , Xiaowei Yue

In this study, we propose the global context guided channel and time-frequency transformations to model the long-range, non-local time-frequency dependencies and channel variances in speaker representations. We use the global context…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Wei Xia , John H. L. Hansen

Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding…

Computer Vision and Pattern Recognition · Computer Science 2022-12-20 Xian Zhong , Zipeng Li , Shuqin Chen , Kui Jiang , Chen Chen , Mang Ye

Recently, vision-language models like CLIP have advanced the state of the art in a variety of multi-modal tasks including image captioning and caption evaluation. Many approaches leverage CLIP for cross-modal retrieval to condition…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Fabian Paischer , Markus Hofmarcher , Sepp Hochreiter , Thomas Adler

Current temporal forgery localization (TFL) approaches typically rely on temporal boundary regression or continuous frame-level anomaly detection paradigms to derive candidate forgery proposals. However, they suffer not only from feature…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Tianyi Wang , Xi Shao , Harry Cheng , Yinglong Wang , Mohan Kankanhalli
‹ Prev 1 3 4 5 6 7 10 Next ›