中文
相关论文

相关论文: WalkCLIP: Multimodal Learning for Urban Walkabilit…

200 篇论文

This research aims to quantify human walking patterns through depth cameras to (1) detect walking pattern changes of a person with and without a motion-restricting device or a walking aid, and to (2) identify distinct walking patterns from…

人机交互 · 计算机科学 2019-03-22 Behnam Malmir

Pretrained vision-language models (VLMs) such as CLIP excel in general multimodal comprehension but often struggle to capture nuanced, context-dependent visual cues. This makes it difficult to distinguish between similar-looking concepts…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Yuchen Huang , Zhiyuan Fan , Zhitao He , Sandeep Polisetty , Wenyan Li , Yi R. Fung

Cities are inherently dynamic. Interesting patterns of behavior typically manifest at several key areas of a city over multiple temporal resolutions. Studying these patterns can greatly help a variety of experts ranging from city planners…

计算机与社会 · 计算机科学 2018-01-01 Fabio Miranda , Harish Doraiswamy , Marcos Lage , Kai Zhao , Bruno Gonçalves , Luc Wilson , Mondrian Hsieh , Cláudio T. Silva

Markerless human motion capture (mocap) from multiple RGB cameras is a widely studied problem. Existing methods either need calibrated cameras or calibrate them relative to a static camera, which acts as the reference frame for the mocap…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Nitin Saini , Chun-hao P. Huang , Michael J. Black , Aamir Ahmad

We present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space. SignCLIP is an…

计算与语言 · 计算机科学 2024-10-08 Zifan Jiang , Gerard Sant , Amit Moryossef , Mathias Müller , Rico Sennrich , Sarah Ebling

Large multi-modal models (LMMs) hold the potential to usher in a new era of automated visual assistance for people who are blind or low vision (BLV). Yet, these models have not been systematically evaluated on data captured by BLV users. We…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Daniela Massiceti , Camilla Longden , Agnieszka Słowik , Samuel Wills , Martin Grayson , Cecily Morrison

Human locomotion involves continuously variable activities including walking, running, and stair climbing over a range of speeds and inclinations as well as sit-stand, walk-run, and walk-stairs transitions. Understanding the kinematics and…

机器人学 · 计算机科学 2021-10-29 Emma Reznick , Kyle R. Embry , Ross Neuman , Edgar Bolívar-Nieto , Nicholas P. Fey , Robert D. Gregg

People who are blind perceive the world differently than those who are sighted, which can result in distinct motion characteristics. For instance, when crossing at an intersection, blind individuals may have different patterns of movement,…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Hee Jae Kim , Kathakoli Sengupta , Masaki Kuribayashi , Hernisa Kacorri , Eshed Ohn-Bar

We propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a variety of audio tasks including classification, retrieval, and…

声音 · 计算机科学 2022-02-16 Ho-Hsiang Wu , Prem Seetharaman , Kundan Kumar , Juan Pablo Bello

Medical image segmentation remains challenging due to limited annotations for training, ambiguous anatomical features, and domain shifts. While vision-language models such as CLIP offer strong cross-modal representations, their potential…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Taha Koleilat , Hojat Asgariandehkordi , Omid Nejati Manzari , Berardino Barile , Yiming Xiao , Hassan Rivaz

Street-view image attribute classification is a vital downstream task of image classification, enabling applications such as autonomous driving, urban analytics, and high-definition map construction. It remains computationally demanding…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Qi You , Yitai Cheng , Zichao Zeng , James Haworth

We present a new framework to generate human-like lower-limb trajectories in periodic and non-periodic walking conditions. In our method, walking dynamics is encoded in 3LP, a linear simplified model composed of three pendulums to model…

机器人学 · 计算机科学 2018-03-28 Salman Faraji , Auke Jan Ijspeert

We investigate research challenges and opportunities for visualization in motion during outdoor physical activities via an initial corpus of real-world recordings that pair egocentric video, biometrics, and think-aloud observations. With…

人机交互 · 计算机科学 2024-09-11 Ahmed Elshabasi , Lijie Yao , Petra Isenberg , Charles Perin , Wesley Willett

The beauty of synchronized dancing lies in the synchronization of body movements among multiple dancers. While dancers utilize camera recordings for their practice, standard video interfaces do not efficiently support their activities of…

人机交互 · 计算机科学 2022-09-15 Zhongyi Zhou , Anran Xu , Koji Yatani

With the increasing availability of aerial and satellite imagery, deep learning presents significant potential for transportation asset management, safety analysis, and urban planning. This study introduces CrosswalkNet, a robust and…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Zubin Bhuyan , Yuanchang Xie , AngkeaReach Rith , Xintong Yan , Nasko Apostolov , Jimi Oke , Chengbo Ai

Trajectory prediction is a fundamental and challenging task for numerous applications, such as autonomous driving and intelligent robots. Currently, most of existing work treat the pedestrian trajectory as a series of fixed two-dimensional…

计算机视觉与模式识别 · 计算机科学 2021-03-17 Pei Lv , Hui Wei , Tianxin Gu , Yuzhen Zhang , Xiaoheng Jiang , Bing Zhou , Mingliang Xu

Video-based gait recognition has achieved impressive results in constrained scenarios. However, visual cameras neglect human 3D structure information, which limits the feasibility of gait recognition in the 3D wild world. Instead of…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Chuanfu Shen , Chao Fan , Wei Wu , Rui Wang , George Q. Huang , Shiqi Yu

Our goal in this paper is the adaptation of image-text models for long video retrieval. Recent works have demonstrated state-of-the-art performance in video retrieval by adopting CLIP, effectively hitchhiking on the image-text…

计算机视觉与模式识别 · 计算机科学 2022-05-18 Max Bain , Arsha Nagrani , Gül Varol , Andrew Zisserman

In this paper, we present a data-driven approach for safely predicting the future state sets of pedestrians. Previous approaches to predicting the future state sets of pedestrians either do not provide safety guarantees or are overly…

系统与控制 · 电气工程与系统科学 2023-08-22 August Söderlund , Frank J. Jiang , Vandana Narri , Amr Alanwar , Karl H. Johansson
‹ 上一页 1 8 9 10 下一页 ›