中文
相关论文

相关论文: Google Landmark Recognition 2020 Competition Third…

200 篇论文

Visual localization has traditionally been formulated as a pair-wise pose regression problem. Existing approaches mainly estimate relative poses between two images and employ a late-fusion strategy to obtain absolute pose estimates.…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Tianchen Deng , Wenhua Wu , Kunzhen Wu , Guangming Wang , Siting Zhu , Shenghai Yuan , Xun Chen , Guole Shen , Zhe Liu , Hesheng Wang

3D global relocalization is one of the key capabilities for mobile robots in practical applications. However, in large scale spaces, existing methods often suffer from prolonged online relocalization time due to factors such as the massive…

机器人学 · 计算机科学 2026-05-11 Jiahua Ren , Kai Shen , Muhua Zhang , Lei Ma

Speech emotion recognition is a challenging classification task with natural emotional speech, especially when the distribution of emotion types is imbalanced in the training and test data. In this case, it is more difficult for a model to…

音频与语音处理 · 电气工程与系统科学 2024-05-31 Mingjie Chen , Hezhao Zhang , Yuanchao Li , Jiachen Luo , Wen Wu , Ziyang Ma , Peter Bell , Catherine Lai , Joshua Reiss , Lin Wang , Philip C. Woodland , Xie Chen , Huy Phan , Thomas Hain

This report to our stage 2 submission to the NeurIPS 2019 disentanglement challenge presents a simple image preprocessing method for learning disentangled latent factors. We propose to train a variational autoencoder on regionally…

机器学习 · 计算机科学 2020-11-18 Maximilian Seitzer , Andreas Foltyn , Felix P. Kemeth

Deep Convolutional Neural Networks (DCNNs) and their variants have been widely used in large scale face recognition(FR) recently. Existing methods have achieved good performance on many FR benchmarks. However, most of them suffer from two…

计算机视觉与模式识别 · 计算机科学 2021-06-28 Jing Xu , Tszhang Guo , Yong Xu , Zenglin Xu , Kun Bai

This report describes our participation in the cDiscount 2015 challenge where the goal was to classify product items in a predefined taxonomy of products. Our best submission yielded an accuracy score of 64.20\% in the private part of the…

机器学习 · 计算机科学 2016-06-10 Ioannis Partalas , Georgios Balikas

This technical report outlines the top-ranking solution for RoboSense 2025: Track 3, achieving state-of-the-art performance on 3D object detection under various sensor placements. Our submission utilizes GBlobs, a local point cloud feature…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Dušan Malić , Christian Fruhwirth-Reisinger , Alexander Prutsch , Wei Lin , Samuel Schulter , Horst Possegger

Accurate localization of cephalometric landmarks holds great importance in the fields of orthodontics and orthognathics due to its potential for automating key point labeling. In the context of landmark detection, particularly in…

计算机视觉与模式识别 · 计算机科学 2023-10-02 Qian Wu , Si Yong Yeo , Yufei Chen , Jun Liu

We devise a graph attention network-based approach for learning a scene triangle mesh representation in order to estimate an image camera position in a dynamic environment. Previous approaches built a scene-dependent model that explicitly…

计算机视觉与模式识别 · 计算机科学 2022-10-03 Mohamed Amine Ouali , Mohamed Bouguessa , Riadh Ksantini

WebFG 2020 is an international challenge hosted by Nanjing University of Science and Technology, University of Edinburgh, Nanjing University, The University of Adelaide, Waseda University, etc. This challenge mainly pays attention to the…

How important is it for training and evaluation sets to not have class overlap in image retrieval? We revisit Google Landmarks v2 clean, the most popular training set, by identifying and removing class overlap with Revisited Oxford and…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Chull Hwan Song , Jooyoung Yoon , Taebaek Hwang , Shunghyun Choi , Yeong Hyeon Gu , Yannis Avrithis

We present an efficient neural network method for locating anatomical landmarks in 3D medical CT scans, using atlas location autocontext in order to learn long-range spatial context. Location predictions are made by regression to Gaussian…

We present two techniques to improve landmark localization in images from partially annotated datasets. Our primary goal is to leverage the common situation where precise landmark locations are only provided for a small data subset, but…

计算机视觉与模式识别 · 计算机科学 2018-10-30 Sina Honari , Pavlo Molchanov , Stephen Tyree , Pascal Vincent , Christopher Pal , Jan Kautz

Many studies have been performed on metric learning, which has become a key ingredient in top-performing methods of instance-level image retrieval. Meanwhile, less attention has been paid to pre-processing and post-processing tricks that…

计算机视觉与模式识别 · 计算机科学 2020-04-24 Byungsoo Ko , Minchul Shin , Geonmo Gu , HeeJae Jun , Tae Kwan Lee , Youngjoon Kim

The Generic Event Boundary Detection (GEBD) task aims to build a model for segmenting videos into segments by detecting general event boundaries applicable to various classes. In this paper, based on last year's MAE-GEBD method, we have…

计算机视觉与模式识别 · 计算机科学 2023-06-29 Yuanxi Sun , Rui He , Youzeng Li , Zuwei Huang , Feng Hu , Xu Cheng , Jie Tang

In this paper, we introduce 3rd place solution for PVUW2023 VSS track. Semantic segmentation is a fundamental task in computer vision with numerous real-world applications. We have explored various image-level visual backbones and…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Shijie Chang , Zeqi Hao , Ben Kang , Xiaoqi Zhao , Jiawen Zhu , Zhenyu Chen , Lihe Zhang , Lu Zhang , Huchuan Lu

This report describes our model for VATEX Captioning Challenge 2020. First, to gather information from multiple domains, we extract motion, appearance, semantic and audio features. Then we design a feature attention module to attend on…

计算机视觉与模式识别 · 计算机科学 2020-06-08 Ke Lin , Zhuoxin Gan , Liwei Wang

The main goal of point cloud registration in Multi-View Partial (MVP) Challenge 2021 is to estimate a rigid transformation to align a point cloud pair. The pairs in this competition have the characteristics of low overlap, non-uniform…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Lifa Zhu , Changwei Lin , Dongrui Liu , Xin Li , Francisco Gómez-Fernández

This paper presents the 2nd place solution to the Facebook AI Image Similarity Challenge : Matching Track on DrivenData. The solution is based on self-supervised learning, and Vision Transformer(ViT). The main breaktrough comes from…

计算机视觉与模式识别 · 计算机科学 2021-11-18 SeungKee Jeon

This paper presents a novel Transformer-based facial landmark localization network named Localization Transformer (LOTR). The proposed framework is a direct coordinate regression approach leveraging a Transformer network to better utilize…

‹ 上一页 1 8 9 10 下一页 ›