中文
相关论文

相关论文: WalkCLIP: Multimodal Learning for Urban Walkabilit…

200 篇论文

Advances in multimodal text-image models have enabled effective text-based querying in extensive image collections. While these models show convincing performance for everyday life scenes, querying in highly homogeneous, specialized domains…

多媒体 · 计算机科学 2025-06-10 Bastian Jäckl , Vojtěch Kloda , Daniel A. Keim , Jakub Lokoč

Importance: Following a century of increase, life expectancy in the United States has stagnated and begun to decline in recent decades. Using satellite images and street view images prior work has demonstrated associations of the built…

计算机视觉与模式识别 · 计算机科学 2020-03-20 Joshua J. Levy , Rebecca M. Lebeaux , Anne G. Hoen , Brock C. Christensen , Louis J. Vaickus , Todd A. MacKenzie

The Internet of Things (IoT) sensors have been widely employed to capture human locomotions to enable applications such as activity recognition, human pose estimation, and fall detection. Motion capture (MoCap) systems are frequently used…

计算机与社会 · 计算机科学 2025-11-18 Yunkai Yu , Yingying Wang , Rong Zheng

Wayfinding behavior and pedestrian movement pattern research relies on objective spatial configuration representation and analysis, such as space syntax, to quantify and control for the difficulty of wayfinding in multi-level buildings and…

计算机与社会 · 计算机科学 2020-12-29 Lingzhu Zhang , Alain J F Chiaradia

Walkability has many health, environmental, and economic benefits. That is why web and mobile services have been offering ways of computing walkability scores of individual street segments. Those scores are generally computed from survey…

社会与信息网络 · 计算机科学 2015-03-11 Daniele Quercia , Luca Maria Aiello , Rossano Schifanella , Adam Davies

Video-based gait analysis has become a promising approach for assessing motor impairment in children with cerebral palsy (CP). However, existing methods usually rely on either pose sequences or handcrafted gait features alone, making it…

图像与视频处理 · 电气工程与系统科学 2026-03-25 Kaiyuan Yang , Xupeng Chen , Jiangpeng He

The emerging ``Floor plan from human trails (PfH)" technique has great potential for improving indoor robot navigation by predicting the traversability of occluded floors. This study presents an innovative approach that replaces…

机器人学 · 计算机科学 2023-10-03 Jonathan Tay Yu Liang , Kanji Tanaka

This paper presents a metric global localization in the urban environment only with a monocular camera and the Google Street View database. We fully leverage the abundant sources from the Street View and benefits from its topo-metric…

机器人学 · 计算机科学 2016-06-17 Li Yu , Cyril Joly , Guillaume Bresson , Fabien Moutarde

Visual perceptual tasks aim to predict human judgment of images (e.g., emotions invoked by images, image quality assessment). Unlike objective tasks such as object/scene recognition, perceptual tasks rely on subjective human assessments,…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Amit Zalcher , Navve Wasserman , Roman Beliy , Oliver Heinimann , Michal Irani

Modeling the dynamics of people walking is a problem of long-standing interest in computer vision. Many previous works involving pedestrian trajectory prediction define a particular set of individual actions to implicitly model group…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Inhwan Bae , Jin-Hwi Park , Hae-Gon Jeon

Passive and non-obtrusive health monitoring using wearables can potentially bring new insights into the user's health status throughout the day and may support clinical diagnosis and treatment. However, identifying segments of free-living…

信号处理 · 电气工程与系统科学 2019-01-31 Yordan P. Raykov , Luc J. W. Evers , Reham Badawy , Marjan J. Faber , Bastiaan R. Bloem , Kasper Claes , Max A. Little

Worldwide geo-localization involves determining the exact geographic location of images captured globally, typically guided by geographic cues such as climate, landmarks, and architectural styles. Despite advancements in geo-localization…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Furong Jia , Lanxin Liu , Ce Hou , Fan Zhang , Xinyan Liu , Yu Liu

A long-standing goal in computer vision is to capture, model, and realistically synthesize human behavior. Specifically, by learning from data, our goal is to enable virtual humans to navigate within cluttered indoor scenes and naturally…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Mohamed Hassan , Duygu Ceylan , Ruben Villegas , Jun Saito , Jimei Yang , Yi Zhou , Michael Black

Foundation models are becoming increasingly effective in the medical domain, offering pre-trained models on large datasets that can be readily adapted for downstream tasks. Despite progress, fetal ultrasound images remain a challenging…

Human motion prediction is crucial for human-centric multimedia understanding and interacting. Current methods typically rely on ground truth human poses as observed input, which is not practical for real-world scenarios where only raw…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Xiao Han , Yiming Ren , Yichen Yao , Yujing Sun , Yuexin Ma

Accurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understanding that textual information can provide. To address this…

计算工程、金融与科学 · 计算机科学 2025-12-11 Xi Xiao , Yunbei Zhang , Janet Wang , Lin Zhao , Yuxiang Wei , Hengjia Li , Yanshu Li , Xinyuan Song , Xiao Wang , Swalpa Kumar Roy , Hao Xu , Tianyang Wang

Training specific deep learning models for particular tasks is common across various domains within seismology. However, this approach encounters two limitations: inadequate labeled data for certain tasks and limited generalization across…

地球物理 · 物理学 2023-09-06 Xu Si , Xinming Wu , Hanlin Sheng , Jun Zhu , Zefeng Li

Life quality in cities is deeply related to the mobility options, and how easily one can access different services and attractions. The pedestrian infrastructure network provides the backbone for social life in cities. While there are many…

物理与社会 · 物理学 2020-06-05 Luis Natera , Dávid Deritei , Anna Vancsó , Orsolya Vásárhelyi

Contrastive Language-Image Pre-training (CLIP)~\citep{radford2021learning} has emerged as a pivotal model in computer vision and multimodal learning, achieving state-of-the-art performance at aligning visual and textual representations…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Shaoan Xie , Lingjing Kong , Yujia Zheng , Yu Yao , Zeyu Tang , Eric P. Xing , Guangyi Chen , Kun Zhang

Medical image segmentation is a cornerstone of computer-assisted diagnosis and treatment planning. While recent multimodal vision-language models have shown promise in enhancing semantic understanding through textual descriptions, their…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Saivan Talaei , Fatemeh Daneshfar , Abdulhady Abas Abdullah , Mustaqeem Khan