English
Related papers

Related papers: TaxaBind: A Unified Embedding Space for Ecological…

200 papers

Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the two tasks and preventing fully end-to-end optimization from…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zhiheng Liu , Weiming Ren , Xiaoke Huang , Shoufa Chen , Tianhong Li , Mengzhao Chen , Yatai Ji , Sen He , Jonas Schult , Belinda Zeng , Tao Xiang , Wenhu Chen , Ping Luo , Luke Zettlemoyer , Yuren Cong

We introduce XYZ-IBD, a bin-picking dataset for 6D pose estimation that captures real-world industrial complexity, including challenging object geometries, reflective materials, severe occlusions, and dense clutter. The dataset reflects…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junwen Huang , Jizhong Liang , Jiaqi Hu , Martin Sundermeyer , Peter KT Yu , Nassir Navab , Benjamin Busam

Learning embeddings that are invariant to the pose of the object is crucial in visual image retrieval and re-identification. The existing approaches for person, vehicle, or animal re-identification tasks suffer from high intra-class…

Computer Vision and Pattern Recognition · Computer Science 2020-08-27 Olga Moskvyak , Frederic Maire , Feras Dayoub , Mahsa Baktashmotlagh

Camera traps are important tools in animal ecology for biodiversity monitoring and conservation. However, their practical application is limited by issues such as poor generalization to new and unseen locations. Images are typically…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Vardaan Pahuja , Weidi Luo , Yu Gu , Cheng-Hao Tu , Hong-You Chen , Tanya Berger-Wolf , Charles Stewart , Song Gao , Wei-Lun Chao , Yu Su

This work focuses on reliable detection of bird sound emissions as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term recordings for…

Sound · Computer Science 2016-09-28 Ilyas Potamitis

Fine-grained and instance-level recognition methods are commonly trained and evaluated on specific domains, in a model per domain scenario. Such an approach, however, is impractical in real large-scale applications. In this work, we address…

Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field has achieved remarkable progress in recent years, further…

Face recognition is a crucial task in various multimedia applications such as security check, credential access and motion sensing games. However, the task is challenging when an input face is noisy (e.g. poor-condition RGB image) or lacks…

Computer Vision and Pattern Recognition · Computer Science 2021-12-09 Wenbin Teng , Chongyang Bai

Explaining why the species lives at a particular location is important for understanding ecological systems and conserving biodiversity. However, existing ecological workflows are fragmented and often inaccessible to non-specialists. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Yutong Zhou , Masahiro Ryo

Patent text embeddings enable prior art search, technology landscaping, and patent analysis, yet existing benchmarks inadequately capture patent-specific challenges. We introduce PatenTEB, a comprehensive benchmark comprising 15 tasks…

Computation and Language · Computer Science 2025-10-28 Iliass Ayaou , Denis Cavallucci

We focus on multi-modal fusion for egocentric action recognition, and propose a novel architecture for multi-modal temporal-binding, i.e. the combination of modalities within a range of temporal offsets. We train the architecture with three…

Computer Vision and Pattern Recognition · Computer Science 2019-08-23 Evangelos Kazakos , Arsha Nagrani , Andrew Zisserman , Dima Damen

Heterogeneous Information Network (HIN) embedding refers to the low-dimensional projections of the HIN nodes that preserve the HIN structure and semantics. HIN embedding has emerged as a promising research field for network analysis as it…

Machine Learning · Computer Science 2021-08-10 Rayyan Ahmad Khan , Martin Kleinsteuber

Cross-modal hashing is usually regarded as an effective technique for large-scale textual-visual cross retrieval, where data from different modalities are mapped into a shared Hamming space for matching. Most of the traditional…

Computer Vision and Pattern Recognition · Computer Science 2017-08-09 Yuming Shen , Li Liu , Ling Shao , Jingkuan Song

Predicting future trajectories of traffic agents in highly interactive environments is an essential and challenging problem for the safe operation of autonomous driving systems. On the basis of the fact that self-driving vehicles are…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Chiho Choi , Joon Hee Choi , Jiachen Li , Srikanth Malla

Predicting future trajectories of traffic agents in highly interactive environments is an essential and challenging problem for the safe operation of autonomous driving systems. On the basis of the fact that self-driving vehicles are…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Chiho Choi , Joon Hee Choi , Srikanth Malla , Jiachen Li

In this paper, we propose Emotionally paired Music and Image Dataset (EMID), a novel dataset designed for the emotional matching of music and images, to facilitate auditory-visual cross-modal tasks such as generation and retrieval. Unlike…

Multimedia · Computer Science 2024-08-12 Jialing Zou , Jiahao Mei , Guangze Ye , Tianyu Huai , Qiwei Shen , Daoguo Dong

Generalised zero-shot learning (GZSL) methods aim to classify previously seen and unseen visual classes by leveraging the semantic information of those classes. In the context of GZSL, semantic information is non-visual data such as a text…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Rafael Felix , Ben Harwood , Michele Sasdelli , Gustavo Carneiro

The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Bing Cao , Quanhao Lu , Jiekang Feng , Qilong Wang , Qinghua Hu , Pengfei Zhu

Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the remarkable progress in text-to-image synthesis, existing models…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Mridul Khurana , Amin Karimi Monsefi , Justin Lee , Medha Sawhney , David Carlyn , Julia Chae , Jianyang Gu , Rajiv Ramnath , Sara Beery , Wei-Lun Chao , Anuj Karpatne , Cheng Zhang

Multimodal learning typically relies on the assumption that all modalities are fully available during both the training and inference phases. However, in real-world scenarios, consistently acquiring complete multimodal data presents…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Donggeun Kim , Taesup Kim
‹ Prev 1 8 9 10 Next ›