English
Related papers

Related papers: Urban-ImageNet: A Large-Scale Multi-Modal Dataset …

200 papers

Large-scale image databases such as ImageNet have significantly advanced image classification and other visual recognition tasks. However much of these datasets are constructed only for single-label and coarse object-level classification.…

Computer Vision and Pattern Recognition · Computer Science 2019-06-17 Sheng Guo , Weilin Huang , Xiao Zhang , Prasanna Srikhanta , Yin Cui , Yuan Li , Matthew R. Scott , Hartwig Adam , Serge Belongie

Despite the promising performance of existing visual models on public benchmarks, the critical assessment of their robustness for real-world applications remains an ongoing challenge. To bridge this gap, we propose an explainable visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Qiang Li , Dan Zhang , Shengzhao Lei , Xun Zhao , Porawit Kamnoedboon , WeiWei Li , Junhao Dong , Shuyan Li

Lightweight vision classification models such as MobileNet, ShuffleNet, and EfficientNet are increasingly deployed in mobile and embedded systems, yet their performance has been predominantly benchmarked on ImageNet. This raises critical…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Weidong Zhang , Pak Lun Kevin Ding , Huan Liu

In this paper, we introduce the ShopSign dataset, which is a newly developed natural scene text dataset of Chinese shop signs in street views. Although a few scene text datasets are already publicly available (e.g. ICDAR2015, COCO-Text),…

Computer Vision and Pattern Recognition · Computer Science 2019-03-26 Chongsheng Zhang , Guowen Peng , Yuefeng Tao , Feifei Fu , Wei Jiang , George Almpanidis , Ke Chen

Many existing 3D semantic segmentation methods, deep learning in computer vision notably, claimed to achieve desired results on urban point clouds. Thus, it is significant to assess these methods quantitatively in diversified real-world…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Maosu Li , Yijie Wu , Anthony G. O. Yeh , Fan Xue

We present ClothesNet: a large-scale dataset of 3D clothes objects with information-rich annotations. Our dataset consists of around 4400 models covering 11 categories annotated with clothes features, boundary lines, and keypoints.…

Federated learning is a new machine learning paradigm which allows data parties to build machine learning models collaboratively while keeping their data secure and private. While research efforts on federated learning have been growing…

Computer Vision and Pattern Recognition · Computer Science 2021-01-06 Jiahuan Luo , Xueyang Wu , Yun Luo , Anbu Huang , Yunfeng Huang , Yang Liu , Qiang Yang

Advances in object recognition flourished in part because of the availability of high-quality datasets and associated benchmarks. However, these benchmarks---such as ILSVRC---are relatively task-specific, focusing predominately on…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Brett D. Roads , Bradley C. Love

This paper provides a first milestone in measuring the floorspace of buildings (that is, building footprint and height) for 40 major Chinese cities. The intent is to maximize city coverage and, eventually provide longitudinal data. Doing so…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Peter Egger , Susie Xi Rao , Sebastiano Papini

In recent years, deep convolutional neural network (DCNN) has seen a breakthrough progress in natural image recognition because of three points: universal approximation ability via DCNN, large-scale database (such as ImageNet), and…

Computer Vision and Pattern Recognition · Computer Science 2020-01-13 Haifeng Li , Xin Dou , Chao Tao , Zhixiang Hou , Jie Chen , Jian Peng , Min Deng , Ling Zhao

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

In this paper, we present a three-step methodological framework, including location identification, bias modification, and out-of-sample validation, so as to promote human mobility analysis with social media data. More specifically, we…

Social and Information Networks · Computer Science 2018-08-15 Yilan Cui , Xing Xie , Yi Liu

Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vision-language models, such as CLIP and large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Sara Ghazanfari , Siddharth Garg , Nicolas Flammarion , Prashanth Krishnamurthy , Farshad Khorrami , Francesco Croce

Fine-grained semantic segmentation of transmission-corridor point clouds is fundamental for intelligent power-line inspection. However, current progress is limited by realistic data scarcity and the difficulty of modeling global corridor…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xu Cui , Xinyan Liu , Chen Yang , Zhaobo Qi , Beichen Zang , Weigang Zhang , Antoni B. Chan

Urban development has been a defining force in human history, shaping cities for centuries. However, past studies mostly analyze such development as predictive tasks, failing to reflect its generative nature. Therefore, this study designs a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Kailai Sun , Yuebing Liang , Mingyi He , Yunhan Zheng , Alok Prakash , Shenhao Wang , Jinhua Zhao , Alex "Sandy'' Pentland

Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily emphasize indoor or street settings, overlooking the unique…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Haotian Xu , Yue Hu , Zhengqiu Zhu , Chen Gao , Ziyou Wang , Junreng Rao , Wenhao Lu , Weishi Li , Quanjun Yin , Yong Li

Urban planners need up-to-date, global, and consistent street network models and indicators to measure resilience and performance, model accessibility, and target local quality-of-life interventions. This article presents up-to-date street…

Physics and Society · Physics 2026-05-04 Geoff Boeing

This study develops FusionTransNet, a framework designed for Origin-Destination (OD) flow predictions within smart and multimodal urban transportation systems. Urban transportation complexity arises from the spatiotemporal interactions…

Machine Learning · Computer Science 2024-05-10 Binwu Wang , Yan Leng , Guang Wang , Yang Wang

Designing socially active streets has long been a goal of urban planning, yet existing quantitative research largely measures pedestrian volume rather than the quality of social interactions. We hypothesize that street view imagery -- an…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Kieran Elrod , Katherine Flanigan , Mario Bergés

Mapping and localization is a critical module of autonomous driving, and significant achievements have been reached in this field. Beyond Global Navigation Satellite System (GNSS), research in point cloud registration, visual feature…