中文
相关论文

相关论文: Urban-ImageNet: A Large-Scale Multi-Modal Dataset …

200 篇论文

Place is an important element in visual understanding. Given a photo of a building, people can often tell its functionality, e.g. a restaurant or a shop, its cultural style, e.g. Asian or European, as well as its economic type, e.g.…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Huaiyi Huang , Yuqi Zhang , Qingqiu Huang , Zhengkui Guo , Ziwei Liu , Dahua Lin

Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model urban perception directly from street view images, but…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Lin Che , Xi Wang , Marc Pollefeys , Konrad Schindler , Martin Raubal , Peter Kiefer

Building footprints provide a useful proxy for a great many humanitarian applications. For example, building footprints are useful for high fidelity population estimates, and quantifying population statistics is fundamental to ~1/4 of the…

计算机视觉与模式识别 · 计算机科学 2021-05-24 Adam Van Etten , Daniel Hogan

Measuring urban safety perception is an important and complex task that traditionally relies heavily on human resources. This process often involves extensive field surveys, manual data collection, and subjective assessments, which can be…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Jiaxin Zhang , Yunqin Li , Tomohiro Fukuda , Bowen Wang

Determining the location of an image anywhere on Earth is a complex visual task, which makes it particularly relevant for evaluating computer vision algorithms. Yet, the absence of standard, large-scale, open-access datasets with reliably…

Learning transferable multimodal embeddings for urban environments is challenging because urban understanding is inherently spatial, yet existing datasets and benchmarks lack explicit alignment between street-view images and urban…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Jie Zhang , Xingtong Yu , Yuan Fang , Rudi Stouffs , Zdravko Trivic

Cities are living systems where urban infrastructures and their functions are defined and evolved due to population behaviors. Profiling the cities and functional regions has been an important topic in urban design and planning. This paper…

社会与信息网络 · 计算机科学 2017-07-14 Lei Shi , Tao Jiang , Ye Zhao , Xiatian Zhang , Yao Lu

Understanding clothes from a single image has strong commercial and cultural impacts on modern societies. However, this task remains a challenging computer vision problem due to wide variations in the appearance, style, brand and layering…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Shuai Zheng , Fan Yang , M. Hadi Kiapour , Robinson Piramuthu

We introduce a novel deep learning-based framework to interpret 3D urban scenes represented as textured meshes. Based on the observation that object boundaries typically align with the boundaries of planar regions, our framework achieves…

计算机视觉与模式识别 · 计算机科学 2022-12-27 Weixiao Gao , Liangliang Nan , Bas Boom , Hugo Ledoux

Modern cities are increasingly reliant on data-driven insights to support decision making in areas such as transportation, public safety and environmental impact. However, city-level data often exists in heterogeneous formats, collected…

机器学习 · 计算机科学 2025-12-15 Takuya Kurihana , Xiaojian Zhang , Wing Yee Au , Hon Yung Wong

Despite the numerous developments in object tracking, further development of current tracking algorithms is limited by small and mostly saturated datasets. As a matter of fact, data-hungry trackers based on deep-learning currently rely on…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Matthias Müller , Adel Bibi , Silvio Giancola , Salman Al-Subaihi , Bernard Ghanem

Urban land use on a building instance level is crucial geo-information for many applications, yet difficult to obtain. An intuitive approach to close this gap is predicting building functions from ground level imagery. Social media image…

计算机视觉与模式识别 · 计算机科学 2022-02-16 Eike Jens Hoffmann , Karam Abdulahhad , Xiao Xiang Zhu

Accurate fine-grained geospatial scene classification using remote sensing imagery is essential for a wide range of applications. However, existing approaches often rely on manually zooming remote sensing images at different scales to…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Yansheng Li , Yuning Wu , Gong Cheng , Chao Tao , Bo Dang , Yu Wang , Jiahao Zhang , Chuge Zhang , Yiting Liu , Xu Tang , Jiayi Ma , Yongjun Zhang

Studies of human mobility increasingly rely on digital sensing, the large-scale recording of human activity facilitated by digital technologies. Questions of variability and population representativity, however, in patterns seen from these…

物理与社会 · 物理学 2018-09-06 Enwei Zhu , Maham Khan , Philipp Kats , Shreya Santosh Bamne , Stanislav Sobolevsky

In this paper, we focus on training and evaluating effective word embeddings with both text and visual information. More specifically, we introduce a large-scale dataset with 300 million sentences describing over 40 million images crawled…

机器学习 · 计算机科学 2016-11-28 Junhua Mao , Jiajing Xu , Yushi Jing , Alan Yuille

We introduce Chinese Text in the Wild, a very large dataset of Chinese text in street view images. While optical character recognition (OCR) in document images is well studied and many commercial tools are available, detection and…

计算机视觉与模式识别 · 计算机科学 2018-03-02 Tai-Ling Yuan , Zhe Zhu , Kun Xu , Cheng-Jun Li , Shi-Min Hu

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Jiaxi Gu , Xiaojun Meng , Guansong Lu , Lu Hou , Minzhe Niu , Xiaodan Liang , Lewei Yao , Runhui Huang , Wei Zhang , Xin Jiang , Chunjing Xu , Hang Xu

Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data. Current methods primarily utilize task-specific models, while recent foundation models for…

机器学习 · 计算机科学 2025-10-16 Dominik J. Mühlematter , Lin Che , Ye Hong , Martin Raubal , Nina Wiedemann

We present PartNet: a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information. Our dataset consists of 573,585 part instances over 26,671 3D models covering 24 object…

计算机视觉与模式识别 · 计算机科学 2018-12-07 Kaichun Mo , Shilin Zhu , Angel X. Chang , Li Yi , Subarna Tripathi , Leonidas J. Guibas , Hao Su

We develop a novel visual model which can recognize protesters, describe their activities by visual attributes and estimate the level of perceived violence in an image. Studies of social media and protests use natural language processing to…

多媒体 · 计算机科学 2017-09-20 Donghyeon Won , Zachary C. Steinert-Threlkeld , Jungseock Joo