中文
相关论文

相关论文: Unbiased Sliced Wasserstein Kernels for High-Quali…

200 篇论文

Video caching can significantly improve backhaul traffic congestion by locally storing the popular content that users frequently request. A privacy-preserving method is desirable to learn how users' demands change over time. As such, this…

网络与互联网体系结构 · 计算机科学 2024-02-27 Md Ferdous Pervej , Andreas F Molisch

Image captioning is the generation of natural language descriptions of images which have increased immense popularity in the recent past. With this different deep-learning techniques are devised for the development of factual and stylized…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Dhruv Sharma , Chhavi Dhiman , Dinesh Kumar

Ultrasound (US) image segmentation embraced its significant improvement in deep learning era. However, the lack of sharp boundaries in US images still remains an inherent challenge for segmentation. Previous methods often resort to global…

图像与视频处理 · 电气工程与系统科学 2020-10-13 Haoming Li , Xin Yang , Jiamin Liang , Wenlong Shi , Chaoyu Chen , Haoran Dou , Rui Li , Rui Gao , Guangquan Zhou , Jinghui Fang , Xiaowen Liang , Ruobing Huang , Alejandro Frangi , Zhiyi Chen , Dong Ni

Wasserstein distributionally robust optimization (WDRO) optimizes against worst-case distributional shifts within a specified uncertainty set, leading to enhanced generalization on unseen adversarial examples, compared to standard…

机器学习 · 计算机科学 2025-03-07 Shuang Liu , Yihan Wang , Yifan Zhu , Yibo Miao , Xiao-Shan Gao

The Wasserstein barycenter problem is to compute the average of $m$ given probability measures, which has been widely studied in many different areas; however, real-world data sets are often noisy and huge, which impedes its applications in…

机器学习 · 计算机科学 2023-12-27 Xu Wang , Jiawei Huang , Qingyuan Yang , Jinpeng Zhang

The Residual Quantization (RQ) framework is revisited where the quantization distortion is being successively reduced in multi-layers. Inspired by the reverse-water-filling paradigm in rate-distortion theory, an efficient regularization on…

机器学习 · 计算机科学 2017-05-02 Sohrab Ferdowsi , Slava Voloshynovskiy , Dimche Kostadinov

The emergence of multimodal foundation models has revolutionized learning paradigms by enabling joint understanding across diverse data types. In the context of next-generation wireless networks, integrating sensing and communication…

网络与互联网体系结构 · 计算机科学 2026-01-01 Mohammad Farzanullah , Han Zhang , Akram Bin Sediq , Ali Afana , Melike Erol-Kantarci

Distributionally robust supervised learning (DRSL) is emerging as a key paradigm for building reliable machine learning systems for real-world applications -- reflecting the need for classifiers and predictive models that are robust to the…

机器学习 · 计算机科学 2022-01-26 Yaodong Yu , Tianyi Lin , Eric Mazumdar , Michael I. Jordan

Neural Radiance Fields from Sparse input} (NeRF-S) have shown great potential in synthesizing novel views with a limited number of observed viewpoints. However, due to the inherent limitations of sparse inputs and the gap between…

计算机视觉与模式识别 · 计算机科学 2023-08-08 Yanqi Bao , Yuxin Li , Jing Huo , Tianyu Ding , Xinyue Liang , Wenbin Li , Yang Gao

The problem of learning functions over spaces of probabilities - or distribution regression - is gaining significant interest in the machine learning community. A key challenge behind this problem is to identify a suitable representation…

机器学习 · 统计学 2022-06-20 Dimitri Meunier , Massimiliano Pontil , Carlo Ciliberto

We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconductor manufacturing, we propose a three-stage synthesis pipeline incorporating structured…

人工智能 · 计算机科学 2026-05-12 Ke Xu , Zhongyuan Lian

Universal Multimodal Retrieval (UMR) seeks any-to-any search across text and vision, yet modern embedding models remain brittle when queries require latent reasoning (e.g., resolving underspecified references or matching compositional…

信息检索 · 计算机科学 2026-02-10 Jianrui Zhang , Anirudh Sundara Rajan , Brandon Han , Soochahn Lee , Sukanta Ganguly , Yong Jae Lee

We introduce caption-guided face recognition (CGFR) as a new framework to improve the performance of commercial-off-the-shelf (COTS) face recognition (FR) systems. In contrast to combining soft biometrics (eg., facial marks, gender, and…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Md Mahedi Hasan , Nasser Nasrabadi

Spectral subtraction, widely used for its simplicity, has been employed to address the Robot Ego Speech Filtering (RESF) problem for detecting speech contents of human interruption from robot's single-channel microphone recordings when it…

机器人学 · 计算机科学 2024-09-11 Yue Li , Koen V. Hindriks , Florian A. Kunneman

Traditional framework of discriminative correlation filters (DCF) is often subject to undesired boundary effects. Several approaches to enlarge search regions have been already proposed in the past years to make up for this shortcoming.…

计算机视觉与模式识别 · 计算机科学 2019-08-08 Ziyuan Huang , Changhong Fu , Yiming Li , Fuling Lin , Peng Lu

Selecting informative keyframes is critical for efficient video understanding, yet existing approaches often rely on heuristics, ignore semantics, or produce redundant frames. We propose KeyScore, a caption-aware frame scoring method that…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Shih-Yao Lin , Sibendu Paul , Caren Chen

Diabetic retinopathy grading is inherently ordinal and long-tailed, with minority stages being scarce, heterogeneous, and clinically critical to detect accurately. Conventional methods often rely on isotropic Gaussian priors and symmetric…

图像与视频处理 · 电气工程与系统科学 2025-10-01 Nagur Shareef Shaik , Teja Krishna Cherukuri , Adnan Masood , Ehsan Adeli , Dong Hye Ye

The performance of automatic speech recognition (ASR) has improved tremendously due to the application of deep neural networks (DNNs). Despite this progress, building a new ASR system remains a challenging task, requiring various resources,…

计算与语言 · 计算机科学 2015-10-20 Yajie Miao , Mohammad Gowayyed , Florian Metze

Self-supervised learning is one of the most promising approaches to acquiring knowledge from limited labeled data. Despite the substantial advancements made in recent years, self-supervised models have posed a challenge to practitioners, as…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Franciskus Xaverius Erick , Mina Rezaei , Johanna Paula Müller , Bernhard Kainz

We present the first approach to automated audio captioning. We employ an encoder-decoder scheme with an alignment model in between. The input to the encoder is a sequence of log mel-band energies calculated from an audio file, while the…

声音 · 计算机科学 2017-10-25 Konstantinos Drossos , Sharath Adavanne , Tuomas Virtanen
‹ 上一页 1 8 9 10 下一页 ›