English
Related papers

Related papers: 3rd Place Solution for Google Universal Image Embe…

200 papers

We propose a viewpoint invariant model for 3D human pose estimation from a single depth image. To achieve this, our discriminative model embeds local regions into a learned viewpoint invariant feature space. Formulated as a multi-task…

Computer Vision and Pattern Recognition · Computer Science 2016-07-27 Albert Haque , Boya Peng , Zelun Luo , Alexandre Alahi , Serena Yeung , Li Fei-Fei

Image compression is one of the most fundamental techniques and commonly used applications in the image and video processing field. Earlier methods built a well-designed pipeline, and efforts were made to improve all modules of the pipeline…

Image and Video Processing · Electrical Eng. & Systems 2021-03-29 Yueyu Hu , Wenhan Yang , Zhan Ma , Jiaying Liu

This article describes the final solution of team monkeytyping, who finished in second place in the YouTube-8M video understanding challenge. The dataset used in this challenge is a large-scale benchmark for multi-label video…

Computer Vision and Pattern Recognition · Computer Science 2017-06-19 He-Da Wang , Teng Zhang , Ji Wu

In this paper, we present our champion solution to the Global Artificial Intelligence Technology Innovation Competition Track 1: Medical Imaging Diagnosis Report Generation. We select CPT-BASE as our base model for the text generation task.…

Computation and Language · Computer Science 2024-07-08 Xiangyu Wu , Hailiang Zhang , Yang Yang , Jianfeng Lu

Most recent semi-supervised video object segmentation (VOS) methods rely on fine-tuning deep convolutional neural networks online using the given mask of the first frame or predicted masks of subsequent frames. However, the online…

Computer Vision and Pattern Recognition · Computer Science 2020-02-18 Yingjie Yin , De Xu , Xingang Wang , Lei Zhang

The rapid development of multi-view 3D human pose estimation (HPE) is attributed to the maturation of monocular 2D HPE and the geometry of 3D reconstruction. However, 2D detection outliers in occluded views due to neglect of view…

Computer Vision and Pattern Recognition · Computer Science 2023-02-24 Xiaoyue Wan , Zhuo Chen , Xu Zhao

General image completion and extrapolation methods often fail on portrait images where parts of the human body need to be recovered - a task that requires accurate human body structure and appearance synthesis. We present a two-stage deep…

Graphics · Computer Science 2019-12-06 Xian Wu , Rui-Long Li , Fang-Lue Zhang , Jian-Cheng Liu , Jue Wang , Ariel Shamir , Shi-Min Hu

Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on…

Autoencoders are commonly trained using element-wise loss. However, element-wise loss disregards high-level structures in the image which can lead to embeddings that disregard them as well. A recent improvement to autoencoders that helps…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Gustav Grund Pihlgren , Fredrik Sandin , Marcus Liwicki

Existing 2D-to-3D pose lifting networks suffer from poor performance in cross-dataset benchmarks. Although the use of 2D keypoints joined by "stick-figure" limbs has shown promise as an intermediate step, stick-figures do not account for…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Saad Manzur , Wayne Hayes

Hand pose estimation from monocular depth images has been an important and challenging problem in the Computer Vision community. In this paper, we present a novel approach to estimate 3D hand joint locations from 2D depth images. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2020-02-21 Rohan Lekhwani , Bhupendra Singh

Real-time, high-quality, 3D scanning of large-scale scenes is key to mixed reality and robotic applications. However, scalability brings challenges of drift in pose estimation, introducing significant errors in the accumulated model.…

Graphics · Computer Science 2017-02-09 Angela Dai , Matthias Nießner , Michael Zollhöfer , Shahram Izadi , Christian Theobalt

We introduce a comprehensive benchmark for local features and robust estimation algorithms, focusing on the downstream task -- the accuracy of the reconstructed camera pose -- as our primary metric. Our pipeline's modular structure allows…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Yuhe Jin , Dmytro Mishkin , Anastasiia Mishchuk , Jiri Matas , Pascal Fua , Kwang Moo Yi , Eduard Trulls

This technical report describes our first-place solution to the pose estimation challenge at ECCV 2022 Visual Perception for Navigation in Human Environments Workshop. In this challenge, we aim to estimate human poses from in-the-wild…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Jiajun Fu , Yonghao Dang , Ruoqi Yin , Shaojie Zhang , Feng Zhou , Wending Zhao , Jianqin Yin

This paper presents the Axon AI's solution to the 2nd YouTube-8M Video Understanding Challenge, achieving the final global average precision (GAP) of 88.733% on the private test set (ranked 3rd among 394 teams, not considering the model…

Computer Vision and Pattern Recognition · Computer Science 2018-09-24 Choongyeun Cho , Benjamin Antin , Sanchit Arora , Shwan Ashrafi , Peilin Duan , Dang The Huynh , Lee James , Hang Tuan Nguyen , Mojtaba Solgi , Cuong Van Than

Modern image retrieval methods typically rely on fine-tuning pre-trained encoders to extract image-level descriptors. However, the most widely used models are pre-trained on ImageNet-1K with limited classes. The pre-trained feature…

Computer Vision and Pattern Recognition · Computer Science 2023-04-13 Xiang An , Jiankang Deng , Kaicheng Yang , Jaiwei Li , Ziyong Feng , Jia Guo , Jing Yang , Tongliang Liu

Augmented reality (AR) displays become more and more popular recently, because of its high intuitiveness for humans and high-quality head-mounted display have rapidly developed. To achieve such displays with augmented information, highly…

Computer Vision and Pattern Recognition · Computer Science 2015-06-22 Kuan-Wen Chen , Chun-Hsin Wang , Xiao Wei , Qiao Liang , Ming-Hsuan Yang , Chu-Song Chen , Yi-Ping Hung

Head pose estimation is a challenging task that aims to solve problems related to predicting three dimensions vector, that serves for many applications in human-robot interaction or customer behavior. Previous researches have proposed some…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Linh Nguyen Viet , Tuan Nguyen Dinh , Hoang Nguyen Viet , Duc Tran Minh , Long Tran Quoc

This article introduces the solutions of the two champion teams, `MMfruit' for the detection track and `MMfruitSeg' for the segmentation track, in OpenImage Challenge 2019. It is commonly known that for an object detector, the shared…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Yu Liu , Guanglu Song , Yuhang Zang , Yan Gao , Enze Xie , Junjie Yan , Chen Change Loy , Xiaogang Wang

The objective of this paper is to design an embedding method that maps local features describing an image (e.g. SIFT) to a higher dimensional representation useful for the image retrieval problem. First, motivated by the relationship…

Computer Vision and Pattern Recognition · Computer Science 2017-04-05 Thanh-Toan Do , Ngai-Man Cheung
‹ Prev 1 3 4 5 6 7 10 Next ›