English
Related papers

Related papers: APEX: Assumption-free Projection-based Embedding e…

200 papers

Speech embeddings are fixed-size acoustic representations of variable-length speech sequences. They are increasingly used for a variety of tasks ranging from information retrieval to unsupervised term discovery and speech segmentation.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-09 Robin Algayres , Mohamed Salah Zaiem , Benoit Sagot , Emmanuel Dupoux

The increasing availability of video recordings made by multiple cameras has offered new means for mitigating occlusion and depth ambiguities in pose and motion reconstruction methods. Yet, multi-view algorithms strongly depend on camera…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Brian Gordon , Sigal Raab , Guy Azov , Raja Giryes , Daniel Cohen-Or

Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Chunyang Cheng , Tianyang Xu , Xiao-Jun Wu , Tao Zhou , Hui Li , Zhangyong Tang , Josef Kittler

Multi-objective alignment for text-to-image generation is commonly implemented via static linear scalarization, but fixed weights often fail under heterogeneous rewards, leading to optimization imbalance where models overfit high-variance,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Dongliang Chen , Xinlin Zhuang , Junjie Xu , Luojian Xie , Zehui Wang , Jiaxi Zhuang , Haolin Yang , Liang Dou , Xiao He , Xingjiao Wu , Ying Qian

Edge-preserving image smoothing is an important step for many low-level vision problems. Though many algorithms have been proposed, there are several difficulties hindering its further development. First, most existing algorithms cannot…

Computer Vision and Pattern Recognition · Computer Science 2019-06-26 Feida Zhu , Zhetong Liang , Xixi Jia , Lei Zhang , Yizhou Yu

We present a novel method for extracting neural embeddings that model the background acoustics of a speech signal. The extracted embeddings are used to estimate specific parameters related to the background acoustic properties of the signal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Sri Harsha Dumpala , Dushyant Sharma , Chandramouli Shama Sastri , Stanislav Kruchinin , James Fosburgh , Patrick A. Naylor

Recent advances in unsupervised learning for object detection, segmentation, and tracking hold significant promise for applications in robotics. A common approach is to frame these tasks as inference in probabilistic latent-variable models.…

Robotics · Computer Science 2021-09-14 Yizhe Wu , Oiwi Parker Jones , Martin Engelcke , Ingmar Posner

State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of…

Developing artificial intelligence (AI) and machine learning (ML) models for medical imaging typically involves extensive training and testing on large datasets, consuming significant computational time, energy, and resources. There is a…

Image and Video Processing · Electrical Eng. & Systems 2024-12-13 Raj Hansini Khoiwal , Alan B. McMillan

Effectively aligning with human judgment when evaluating machine-generated image captions represents a complex yet intriguing challenge. Existing evaluation metrics like CIDEr or CLIP-Score fall short in this regard as they do not take into…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Yang Liu , Mengyuan Liu , Shudong Huang , Jiancheng Lv

Empirical fixation densities, spatial distributions estimated from human eye-tracking data, are foundational to saliency benchmarking. They directly shape benchmark conclusions, leaderboard rankings, failure case analyses, and scientific…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Susmit Agrawal , Jannis Hollman , Matthias Kümmerer

Visual prompting has emerged as a powerful method for adapting pre-trained models to new domains without updating model parameters. However, existing prompting methods typically optimize a single prompt per domain and apply it uniformly to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Evren Çetinkaya , Sangmin Lee , Jung Uk Kim , Hong Joo Lee , Nassir Navab

We show how perceptual embeddings of the visual system can be constructed at inference-time with no training data or deep neural network features. Our perceptual embeddings are solutions to a weighted least squares (WLS) problem, defined at…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Daniel Severo , Lucas Theis , Johannes Ballé

Although remarkable progress has been made, existing methods for enhancing underexposed photos tend to produce visually unpleasing results due to the existence of visual artifacts (e.g., color distortion, loss of details and uneven…

Computer Vision and Pattern Recognition · Computer Science 2020-07-09 Qing Zhang , Yongwei Nie , Lei Zhu , Chunxia Xiao , Wei-Shi Zheng

Accurate face recognition systems are increasingly important in sensitive applications like border control or migration management. Therefore, it becomes crucial to quantify the quality of facial images to ensure that low-quality images are…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Marcel Grimmer , Christian Rathgeb , Raymond Veldhuis , Christoph Busch

In this paper, we highlight a problem of evaluation metrics adopted in the open-vocabulary segmentation. That is, the evaluation process still heavily relies on closed-set metrics on zero-shot or cross-dataset pipelines without considering…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Hao Zhou , Tiancheng Shen , Xu Yang , Hai Huang , Xiangtai Li , Lu Qi , Ming-Hsuan Yang

Several variants of deep neural networks have been successfully employed for building parametric models that project variable-duration spoken word segments onto fixed-size vector representations, or acoustic word embeddings (AWEs). However,…

Computation and Language · Computer Science 2021-06-17 Badr M. Abdullah , Marius Mosbach , Iuliia Zaitova , Bernd Möbius , Dietrich Klakow

Conventional 2D pose estimation models are constrained by their design to specific object categories. This limits their applicability to predefined objects. To overcome these limitations, category-agnostic pose estimation (CAPE) emerged as…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Matan Rusanovsky , Or Hirschorn , Shai Avidan

Recent advances in large language models and vision-language models have led to growing interest in explainable evaluation metrics for image captioning. However, these metrics generate explanations without standardized criteria, and the…

Computation and Language · Computer Science 2025-07-01 Hyunjong Kim , Sangyeop Kim , Jongheon Jeong , Yeongjae Cho , Sungzoon Cho
‹ Prev 1 2 3 10 Next ›